From labeled drawings to a cloud endpoint.
An Azure setup guide for fine-tuning YOLO26 on raster HVAC symbols.
Start with one GPU. Prepare a versioned dataset, train a small model, compare validation results, then publish the selected model as a service. This page documents the workflow; it does not submit jobs or deploy resources.
Set up the Azure workspace
Create these resources once, then reuse them for each experiment.
| Resource | Recommended setup | Purpose / what to check |
|---|---|---|
| Azure ML workspace | Create in your chosen subscription, resource group, and region. | Home for jobs, data assets, experiments, model versions, and endpoints. |
| Blob storage + data asset | Upload hvac-symbols-v1; register a versioned folder data asset. | Keep original drawings, labels, split manifest, and processed tiles. Use a new version when annotations change. |
| CPU compute | Small CPU cluster; scale to zero when idle. | Check annotations, generate tiles, and assemble reports without occupying a GPU. |
| GPU compute | One GPU per node; target 24 GB VRAM; minimum 0 nodes, maximum 1. | Check Azure ML supported sizes and regional quota first. 24 GB is a sizing target, not an Azure SKU. A 16 GB GPU may work with a smaller batch; use a larger supported GPU if required. |
| Training environment | Versioned CUDA-compatible PyTorch container with pinned Ultralytics, MLflow, and azureml-mlflow. | Record package versions and verify YOLO26 and GPU availability in a short smoke run. |
| Job identity + outputs | Grant the job identity access to input data and persistent output storage. | Save checkpoints and reports outside temporary compute. Set a job timeout and idle scale-down interval. |
| Model + endpoint | Register the selected weights; create a managed online endpoint when ready. | Serving compute is separate from training compute. Choose it after measuring inference latency and memory. |
Azure references: Compute clusters · Command jobs
Prepare YOLO training data
Each raster image has a matching text file containing its symbol boxes.
Folder structure
hvac-symbols-v1/
├── data.yaml
├── split-manifest.csv
├── images/
│ ├── train/ drawing01_tile001.png
│ ├── val/ drawing08_tile001.png
│ └── test/ drawing12_tile001.png
└── labels/
├── train/ drawing01_tile001.txt
├── val/ drawing08_tile001.txt
└── test/ drawing12_tile001.txtdata.yaml · example classes
path: /data/hvac-symbols-v1 train: images/train val: images/val test: images/test names: 0: supply_diffuser 1: return_grille 2: vav_box
Resolve path to the mounted or downloaded dataset location inside the Azure job.
One object per line
class_id x_center y_center width height0 0.500000 0.250000 0.125000 0.125000
This is class 0: supply diffuser. On a 1,024 × 1,024 tile, its center is (512, 256) and its box is 128 × 128 pixels. Normalize coordinates to 0–1 relative to the tile, not the full sheet. Class IDs start at zero.
- Annotate every target symbol; keep reviewed background tiles with empty label files.
- Start with 1,024 px tiles and 256 px overlap (768 px stride). Remap boxes into tile coordinates and use a consistent clipping/filtering policy for partial objects.
- Store the source image, project ID, tile offset, and split in the manifest so full-sheet predictions can be reconstructed.
- The app’s 80 reference crops are a symbol library, not a labeled detection dataset. Add real drawings, boxes, and representative style variations.
- Leader paths, OCR text, and label-to-symbol links need separate annotations and evaluation; YOLO box labels do not encode those relationships.
How the pipeline runs in Azure
Submit a run once. Each step consumes the previous step’s saved output.
Submit dataset version + experiment settings
- CPUPrepareValidate labels
Split & tile drawings - 1 GPU · 24 GB targetTrainFine-tune YOLO26
Validate each epoch - GPU / CPUEvaluateCompare validation runs
Test selected model - ARTIFACTSReport & registerSave metrics & plots
Version chosen weights
Versioned drawings, annotations, split manifest
Parameters, epoch metrics, precision / recall
Checkpoints, report files, annotated test sheets
Run training, step by step
Start with a small verified run before spending time on full experiments.
Prepare and review labeled data
Draw boxes, agree on class names, split by project, and generate tiles. Check missing labels, invalid coordinates, duplicates, and class counts. Upload and version the data asset.
OUTPUT · versioned dataset + manifestDefine the Azure pipeline
Create Prepare, Train, and Evaluate/Report command components using Azure ML SDK v2 or CLI v2. Declare data/model folders as inputs and outputs, connect them, and assign CPU or GPU compute to each step. Log parameters and metrics to MLflow from the scripts.
OUTPUT · reusable pipeline definition + versioned environmentSet the first experiment’s hyperparameters
Begin with the settings below. Save the dataset version, seed, package versions, and all arguments with the run. Keep the split fixed while comparing models.
Parameter Starting value How to adjust Pretrained weights yolo26s.ptTry Medium after establishing a Small baseline. Image size 1024Increase only if small details remain unresolved and memory allows. Batch size -1(auto)Log the resolved size; reduce it if memory is exhausted. Epochs / patience 100 / 20Maximum epochs / early stopping patience. Inspect validation trends. Optimizer / learning rate Baseline: optimizer=autoRecord resolved settings. To test a specific learning rate, select an explicit optimizer; auto may override manual learning-rate settings. Augmentation Review rotations, flips, scaling Use only transformations that preserve symbol meaning; disable invalid defaults. Compute One GPU; 24 GB VRAM target Verify actual SKU memory and availability. This is not a measured requirement. Run a 2–5 epoch smoke test
Confirm CUDA is active, labels load correctly, checkpoints reach persistent storage, and MLflow receives metrics. Measure seconds per epoch and peak memory before estimating the full run’s duration and cost.
OUTPUT · verified environment + runtime estimateTrain the full baseline on the GPU
Submit the job and inspect Azure ML Studio → Jobs → the run. Log training losses and validation metrics each epoch. Save
OUTPUT · checkpoints + learning curvesbest.pt,last.pt, configuration, and plots. Use the best checkpoint rather than assuming the final epoch is best.Compare fine-tuned candidates
Change one major setting per experiment, such as Small versus Medium or image resolution. Compare full-sheet validation results after tile merging. Choose confidence thresholds on validation data and inspect missed rare symbols, not only the overall average.
OUTPUT · chosen run + frozen inference settingsTest once and register the selected model
Run the selected configuration on untouched test projects. Save the final report, failure examples, and latency measurements. Register the weights with class mapping, preprocessing, tile overlap, merge settings, thresholds, and environment version.
OUTPUT · final test report + registered model version
Decide which model is better
Compare useful operating points on the same projects, not just one headline score.
Precision
TP / (TP + FP)Of the symbols we detected, how many were correct?
Recall
TP / (TP + FN)Of the real symbols, how many did we find?
mAP & per-class results
Use mAP50 and mAP50–95 as summaries, alongside rare-class recall, count errors, and full-sheet latency.
Illustrative validation comparison Invented numbers · not project results
Example decision rule: require at least 95% precision, then prefer higher recall within an agreed latency budget. Tune each run’s threshold on validation data only.
| Run | Model / input | Confidence | Precision | Recall | Seconds / sheet | Decision |
|---|---|---|---|---|---|---|
| A | Small / 1,024 | 0.40 | 96% | 88% | 8 | Fast baseline |
| B | Medium / 1,024 | 0.45 | 96% | 93% | 13 | Prefer if 13 s fits the budget |
| C | Small / 1,280 | 0.30 | 92% | 95% | 12 | Fails precision target at this threshold |
Run B meets the example precision target and finds more symbols than A. Run C’s higher recall does not satisfy the precision target at the shown threshold; inspect its precision–recall curve before rejecting the model entirely.
What to publish in the cloud report
- Azure ML metrics: validation loss, precision, recall, mAP, dataset version, and epoch number.
- Downloadable artifacts: precision–recall curves, confusion matrix, per-class CSV, and annotated missed/false detections.
- Final full-sheet report: deduplicated precision/recall, count errors, class support, and latency on fixed test projects.
- Separate semantic report later: OCR accuracy and label-to-symbol association accuracy. Detector metrics do not measure leader understanding.
Publish the model as an endpoint service
Deployment is a separate action after the selected model passes evaluation.
- Register the model package. Include weights, class names, preprocessing and inference settings, and the environment version.
- Write a scoring adapter. In
score.py, load the model once ininit(). Inrun(), validate the image input, tile it, infer, merge duplicates, and return class names, scores, and full-image boxes. This adapter is future implementation, not included in this guide. - Create a managed online endpoint and deployment. Configure authentication, the registered model, scoring script, environment, and serving instance type/count. Training’s 24 GB recommendation does not determine serving requirements.
- Test before routing traffic. Check health, response schema, known images, latency, invalid inputs, and memory. Assign traffic to the tested deployment; retain a previous version for rollback.
- Connect through an application backend. The backend authenticates to Azure; do not embed endpoint credentials in browser JavaScript. Monitor errors, latency, and new drawing styles.