Team Ai
Modelpublic

Intel/object-classification

sourceHugging Facemitupdated 1mo agoView on Hugging Face
2likes41downloads
README.md389 linesDownload Raw Back to root
1---2license: mit3license_link: LICENSE4library_name: openvino5pipeline_tag: object-detection6tags:7  - openvino8  - intel9  - yolo10  - yolo2611  - object-classification12  - classification13  - smart-city14  - situational-awareness15  - traffic16  - edge-ai17  - metro18  - dlstreamer19language:20  - en21---22 23# Object Classification24 25| Property | Value |26|---|---|27| **Category** | Object Classification (Traffic Categorization: People / Vehicles) |28| **Base Model** | [YOLO26](https://docs.ultralytics.com/models/yolo26/) (Ultralytics) |29| **Source Framework** | PyTorch (Ultralytics) |30| **Supported Precisions** | FP32, FP16, INT8 (mixed-precision) |31| **Inference Engine** | OpenVINO |32| **Hardware** | CPU, GPU, NPU |33| **Detected Class(es)** | `person` and vehicle classes grouped into `People` and `Vehicles` |34 35---36 37## Overview38 39Object Classification is a Metro Analytics use case that detects objects with [YOLO26](https://docs.ultralytics.com/models/yolo26/) and then categorizes each detection into higher-level city-operations groups.40It is built on the state-of-the-art YOLO26 real-time detector, quantized to INT8 for efficient inference on Intel hardware.41Where the general object-detection use case reports every one of the 80 COCO classes individually, this use case rolls the traffic-relevant classes up into two semantic categories, `People` and `Vehicles`, so operators get an at-a-glance picture of a scene.42 43The traffic categories are:44 45- **People** -- the COCO `person` class.46- **Vehicles** -- the COCO `bicycle`, `car`, `motorcycle`, `bus`, `train`, and `truck` classes.47 48Objects outside these categories are ignored to keep the output focused on traffic situational awareness.49 50Typical Metro deployments include:51 52- **Situational Awareness** -- summarize each camera feed as live People and Vehicles counts.53- **Automated City Operations** -- feed category counts into signal timing, congestion, and dispatch logic.54- **Intersection and Roundabout Monitoring** -- track the mix of pedestrians and vehicles at busy junctions.55- **Trend Analytics** -- aggregate category counts over time to understand traffic patterns.56 57Available variants: `yolo26n`, `yolo26s`, `yolo26m`, `yolo26l`, `yolo26x`.58Smaller variants (`yolo26n`, `yolo26s`) are recommended for high-FPS edge deployment; larger variants improve recall for small objects.59 60---61 62## Prerequisites63 64- Python 3.11+65- [Install OpenVINO](https://docs.openvino.ai/2026/get-started/install-openvino.html) (latest version)66- [Install Intel DLStreamer](https://docs.openedgeplatform.intel.com/2026.0/edge-ai-libraries/dlstreamer/get_started/install/install_guide_ubuntu.html) (latest version)67 68Create and activate a Python virtual environment before running the scripts:69 70```bash71python3 -m venv .venv --system-site-packages72source .venv/bin/activate73```74 75> **Note:** The `--system-site-packages` flag is required so the virtual76> environment can access the system-installed OpenVINO and DLStreamer Python77> packages.78 79---80 81## Getting Started82 83### Download and Quantize Model84 85Run the provided script to download, export to OpenVINO IR, and optionally quantize:86 87```bash88chmod +x export_and_quantize.sh89./export_and_quantize.sh90```91 92This exports the default **yolo26n** model in **FP16** precision.93 94#### Optional: Select a Different Variant or Precision95 96```bash97./export_and_quantize.sh yolo26n FP32   # full-precision98./export_and_quantize.sh yolo26n INT8   # quantized99./export_and_quantize.sh yolo26s        # larger variant, default FP16100```101 102Replace `yolo26n` with any variant (`yolo26s`, `yolo26m`, `yolo26l`, `yolo26x`).103The second argument selects the precision (`FP32`, `FP16`, `INT8`); the default is **FP16**.104 105The script performs the following steps:106 1071. Installs dependencies (`openvino`, `ultralytics`; adds `nncf` for INT8).1082. Downloads a sample traffic video (`test_video.mp4`) of an urban roundabout at low resolution (640x360).1093. Downloads the PyTorch weights and exports to OpenVINO IR.1104. *(INT8 only)* Quantizes the model using NNCF post-training quantization.111 112Output files:113 114- `yolo26n_openvino_model/` -- FP32 or FP16 OpenVINO IR model directory.115- `yolo26n_objcls_int8.xml` / `yolo26n_objcls_int8.bin` -- INT8 quantized model *(only when `INT8` is selected)*.116 117#### Precision / Device Compatibility118 119| Precision | CPU | GPU | NPU |120|---|---|---|---|121| FP32 | Yes | Yes | No |122| FP16 | Yes | Yes | Yes |123| INT8 | Yes | Yes | Yes |124 125> **Note:** The INT8 calibration uses a frame from the bundled sample video.126> For production accuracy, replace it with a representative set of frames from127> the target deployment site.128 129### OpenVINO Sample130 131The sample below runs YOLO26 inference on the sample traffic video, maps each132detection into the `People` or `Vehicles` category, draws boxes colored per133category, overlays live category counts, and writes the annotated result to134`output_openvino.mp4`.135YOLO26 is end-to-end (NMS-free), so no manual non-maximum suppression is needed.136Change the `device` string to run on CPU, GPU, or NPU.137 138```python139import cv2140import numpy as np141import openvino as ov142 143CONF_THRESHOLD = 0.4144INPUT_SIZE = 640145 146# Map the traffic-relevant COCO class ids into higher-level city categories.147# People and Vehicles are the two categories tracked for situational awareness.148CATEGORY_BY_CLASS_ID = {149    0: "People",      # person150    1: "Vehicles",    # bicycle151    2: "Vehicles",    # car152    3: "Vehicles",    # motorcycle153    5: "Vehicles",    # bus154    6: "Vehicles",    # train155    7: "Vehicles",    # truck156}157# BGR overlay colors for each category.158CATEGORY_COLORS = {159    "People": (0, 200, 0),160    "Vehicles": (255, 128, 0),161}162 163core = ov.Core()164model = core.read_model("yolo26n_openvino_model/yolo26n.xml")165 166# YOLO26 embeds the 80 COCO class names in rt_info. Ultralytics separates167# multi-word names with underscores (e.g. "traffic_light"), so restore spaces.168COCO_NAMES = [169    name.replace("_", " ")170    for name in model.get_rt_info()["model_info"]["labels"].value.split()171]172 173# Change device to "GPU" or "NPU" to run on integrated GPU or NPU.174compiled = core.compile_model(model, "CPU")175output_port = compiled.output(0)176 177cap = cv2.VideoCapture("test_video.mp4")178fps = cap.get(cv2.CAP_PROP_FPS) or 30.0179width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))180height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))181writer = cv2.VideoWriter(182    "output_openvino.mp4", cv2.VideoWriter_fourcc(*"mp4v"), fps, (width, height)183)184 185totals = {"People": 0, "Vehicles": 0}186frame_idx = 0187while True:188    ok, frame = cap.read()189    if not ok:190        break191    frame_idx += 1192 193    blob = cv2.resize(frame, (INPUT_SIZE, INPUT_SIZE))194    blob = cv2.cvtColor(blob, cv2.COLOR_BGR2RGB).astype(np.float32) / 255.0195    blob = blob.transpose(2, 0, 1)[np.newaxis, ...]  # NCHW196 197    # YOLO26 end-to-end output: [1, 300, 6] = [x1, y1, x2, y2, confidence, class_id].198    output = compiled([blob])[output_port][0]199 200    sx, sy = width / INPUT_SIZE, height / INPUT_SIZE201    counts = {"People": 0, "Vehicles": 0}202    for x1, y1, x2, y2, conf, class_id in output:203        if conf < CONF_THRESHOLD:204            continue205        category = CATEGORY_BY_CLASS_ID.get(int(class_id))206        if category is None:207            continue  # not a traffic-relevant object208        counts[category] += 1209        totals[category] += 1210        color = CATEGORY_COLORS[category]211        px1, py1 = int(x1 * sx), int(y1 * sy)212        px2, py2 = int(x2 * sx), int(y2 * sy)213        label = f"{category}: {COCO_NAMES[int(class_id)]} {conf:.2f}"214        cv2.rectangle(frame, (px1, py1), (px2, py2), color, 2)215        cv2.putText(frame, label, (px1, py1 - 5),216                    cv2.FONT_HERSHEY_SIMPLEX, 2.0, color, 2)217 218    # Overlay the per-category counts for this frame.219    banner = f"People: {counts['People']}  Vehicles: {counts['Vehicles']}"220    cv2.rectangle(frame, (0, 0), (width, 60), (0, 0, 0), -1)221    cv2.putText(frame, banner, (15, 45),222                cv2.FONT_HERSHEY_SIMPLEX, 2.0, (255, 255, 255), 2)223 224    if frame_idx % 30 == 0:225        print(f"frame {frame_idx}: {banner}", flush=True)226 227    writer.write(frame)228 229cap.release()230writer.release()231print(f"Summary: People={totals['People']} Vehicles={totals['Vehicles']}")232print("Saved: output_openvino.mp4")233```234 235**Device targets:**236 237- `"CPU"` -- default, works on all Intel platforms.238- `"GPU"` -- Intel integrated or discrete GPU.239- `"NPU"` -- Intel NPU (validate with `benchmark_app -d NPU`).240 241### Try It on a Sample Video242 243The `export_and_quantize.sh` script downloads `test_video.mp4` automatically.244Re-run the OpenVINO sample above.245The script reads `test_video.mp4`, prints the running People and Vehicles counts to the console, and writes the annotated video to `output_openvino.mp4`.246 247Expected console output (representative):248 249```text250frame 30: People: 4  Vehicles: 6251frame 60: People: 3  Vehicles: 7252frame 90: People: 5  Vehicles: 5253Summary: People=372 Vehicles=548254Saved: output_openvino.mp4255```256 257#### Expected Output258 259![OpenVINO expected output](expected_output_openvino.gif)260 261### DLStreamer Sample262 263The pipeline below runs the FP16 YOLO26 detector on the sample video via264`gvadetect`, overlays bounding boxes with `gvawatermark` for the traffic-relevant265classes only (non-traffic detections such as `handbag` are filtered out via266`show-roi`), saves the annotated result to `output_dlstreamer.mp4`, and prints the267`People` and `Vehicles` category counts per frame from the detection metadata.268 269> **Notes on running this sample:**270>271> - Use the FP16 IR (`yolo26n_openvino_model/yolo26n.xml`). Class names are272>   read automatically from the model's embedded `metadata.yaml` by273>   DLStreamer 2026.0+ -- no external `labels-file` is required.274> - Export `PYTHONPATH` so the DLStreamer Python module is importable:275>276>   ```bash277>   source /opt/intel/openvino_2026/setupvars.sh278>   source /opt/intel/dlstreamer/scripts/setup_dls_env.sh279>   export PYTHONPATH=/opt/intel/dlstreamer/python:\280>   /opt/intel/dlstreamer/gstreamer/lib/python3/dist-packages:${PYTHONPATH:-}281>   ```282 283```python284import gi285 286gi.require_version("Gst", "1.0")287gi.require_version("GstAnalytics", "1.0")288from gi.repository import Gst, GLib, GstAnalytics289 290Gst.init([])291 292INPUT_VIDEO = "test_video.mp4"293 294# Traffic-relevant COCO labels grouped into higher-level city categories.295CATEGORY_BY_LABEL = {296    "person": "People",297    "bicycle": "Vehicles",298    "car": "Vehicles",299    "motorcycle": "Vehicles",300    "bus": "Vehicles",301    "train": "Vehicles",302    "truck": "Vehicles",303}304 305# For CPU: change device=GPU to device=CPU.306# For NPU: change device=GPU to device=NPU (batch-size=1, nireq=4 recommended).307# gvawatermark displ-cfg:308#   show-roi=... draws only the traffic-relevant classes (person + vehicles),309#     so non-traffic detections such as handbag/backpack are not boxed.310#   font-scale=1.5 enlarges the label text for better visualization.311pipeline_str = (312    f"filesrc location={INPUT_VIDEO} ! decodebin3 ! "313    "videoconvert ! "314    "gvadetect model=yolo26n_openvino_model/yolo26n.xml "315    "device=GPU "316    "threshold=0.4 ! queue ! "317    "gvawatermark "318    "displ-cfg=show-roi=person:bicycle:car:motorcycle:bus:train:truck,font-scale=2.5 ! "319    "videoconvert ! video/x-raw,format=I420 ! "320    "openh264enc ! h264parse ! "321    "mp4mux ! filesink name=sink location=output_dlstreamer.mp4"322)323pipeline = Gst.parse_launch(pipeline_str)324 325totals = {"People": 0, "Vehicles": 0}326 327 328def on_buffer(pad, info):329    buf = info.get_buffer()330    rmeta = GstAnalytics.buffer_get_analytics_relation_meta(buf)331    if rmeta is None:332        return Gst.PadProbeReturn.OK333    counts = {"People": 0, "Vehicles": 0}334    idx = 1335    while True:336        ok, od = rmeta.get_od_mtd(idx)337        if not ok:338            break339        label = GLib.quark_to_string(od.get_obj_type())340        category = CATEGORY_BY_LABEL.get(label)341        if category is not None:342            counts[category] += 1343            totals[category] += 1344        idx += 1345    if counts["People"] or counts["Vehicles"]:346        print(f"frame: People={counts['People']} Vehicles={counts['Vehicles']}",347              flush=True)348    return Gst.PadProbeReturn.OK349 350 351sink = pipeline.get_by_name("sink")352sink_pad = sink.get_static_pad("sink")353sink_pad.add_probe(Gst.PadProbeType.BUFFER, on_buffer)354 355pipeline.set_state(Gst.State.PLAYING)356bus = pipeline.get_bus()357bus.timed_pop_filtered(358    Gst.CLOCK_TIME_NONE,359    Gst.MessageType.EOS | Gst.MessageType.ERROR,360)361pipeline.set_state(Gst.State.NULL)362print(f"Summary: People={totals['People']} Vehicles={totals['Vehicles']}")363```364 365#### Expected Output366 367![DLStreamer expected output](expected_output_dlstreamer.gif)368 369**Device targets:**370 371- `device=GPU` -- default in the sample code.372- `device=CPU` -- change `device=GPU` to `device=CPU`.373- `device=NPU` -- change `device=GPU` to `device=NPU`; use `batch-size=1` and `nireq=4` for best NPU utilization.374 375---376 377## License378 379Licensed under the MIT License. See [LICENSE](LICENSE) for details.380 381## References382 383- [YOLO26 Documentation](https://docs.ultralytics.com/models/yolo26/)384- [OpenVINO YOLO26 Notebook](https://github.com/openvinotoolkit/openvino_notebooks/blob/latest/notebooks/yolov26-optimization/yolov26-object-detection.ipynb)385- [Sample video: Urban roundabout with cars and pedestrian (Pexels)](https://www.pexels.com/video/urban-roundabout-with-cars-and-pedestrian-30119018/)386- [OpenVINO Documentation](https://docs.openvino.ai/)387- [NNCF Post-Training Quantization](https://docs.openvino.ai/latest/nncf_ptq_introduction.html)388- [Intel DLStreamer](https://docs.openedgeplatform.intel.com/2026.0/edge-ai-libraries/dlstreamer/index.html)389