What is object detection? YOLO explained, detection vs classification vs segmentation, training on your own data, and running it in real time on a budget GPU

Object detection finds and labels every object in an image with a box. How YOLO-style detectors work in one pass, how detection differs from classification and segmentation, how mAP and IoU are measured, how to train a detector on your own images, how to run it in real time on cheap hardware or in the browser, and the mistakes that produce a demo that fails on real cameras.

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Key takeaways

  • Classification says what is in an image; detection says what and where, with a box per object; segmentation says which pixels.
  • YOLO-style detectors predict boxes and classes in a single pass, which is why they run in real time on modest hardware.
  • Accuracy is measured by mAP at an IoU threshold. Learn what those mean before comparing models.
  • Training on your own data is mostly labelling: a few hundred well-labelled images per class beats thousands of noisy ones.
  • Real cameras differ from datasets in lighting, angle and resolution. Test on the deployment camera, then quantize and export for the device.

What is object detection?

Object detection is the computer vision task of finding every object of interest in an image and returning, for each, a bounding box, a class label and a confidence score. Where image classification answers "what is in this picture", detection answers "what is here and exactly where", and it does so for many objects at once. It powers counting, tracking, quality inspection, safety monitoring, retail analytics and the perception stacks of robots and vehicles.

Also asked as: what is object detection · object detection meaning · object detection in computer vision · object detection explained · what is object detection in ai · object detection vs image recognition · how does object detection work · object detection examples · object detection applications

I have built vision systems that run in browsers and on modest servers, including FaceVision, a face detection and recognition pipeline that runs entirely in the browser, and the lessons about cameras, lighting and evaluation carry directly to general object detection. My notes on budget hardware, camera problems and evaluation are on the portfolio [11][12][13].

The model that scored highest on the dataset was not the one that worked on the factory camera. The camera is part of the model. Test on it. Pranjul Rathour

Classification vs detection vs segmentation: what is the difference?

Classification assigns one label to the whole image. Detection assigns a box and a label to each object. Semantic segmentation assigns a class to every pixel. Instance segmentation combines detection and segmentation: a pixel mask per object. Each step up gives more information and costs more compute and, above all, more labelling effort, because a box takes seconds to draw and a mask takes minutes. Choose the least detailed output that answers your question.

Also asked as: object detection vs image classification · object detection vs segmentation · classification vs detection vs segmentation · semantic segmentation vs instance segmentation · difference between object detection and object recognition · object localization vs detection

How does YOLO work?

YOLO, "you only look once", treats detection as a single regression problem: one network pass over the whole image predicts boxes, class probabilities and confidences for a grid of locations at once, instead of proposing regions first and classifying each, as two-stage detectors such as Faster R-CNN do [1][2]. Non-maximum suppression then removes duplicate boxes for the same object. The single pass is why YOLO-family models run in real time on a laptop GPU or even a phone, at some cost in accuracy on very small or crowded objects.

Also asked as: how does yolo work · what is yolo in object detection · yolo algorithm explained · yolo architecture · yolo vs faster rcnn · one stage vs two stage detector · yolo versions · what does yolo stand for · yolo object detection tutorial · latest yolo model

DETR later showed a transformer could detect end to end without anchors or suppression [3]. For most student and production projects in 2026, a current YOLO-family model from the Ultralytics library is the practical default: documented, exportable and fast [6].

How is object detection accuracy measured?

With intersection over union and mean average precision. IoU is the overlap between a predicted box and the ground-truth box divided by their union; a prediction counts as correct if IoU exceeds a threshold, commonly 0.5. Average precision summarises the precision-recall curve for one class across confidence thresholds; mAP averages it across classes. COCO reports mAP averaged over IoU thresholds from 0.5 to 0.95, which is stricter than the older PASCAL VOC 0.5 [4][5]. Compare models on the same metric and the same dataset, or the comparison is meaningless.

Also asked as: what is map in object detection · mean average precision explained · iou object detection · what is intersection over union · map50 vs map50-95 · how to evaluate object detection · object detection metrics · precision recall object detection · confidence threshold object detection

How do I train an object detector on my own data?

Collect images from the camera and conditions you will deploy on. Label boxes with a tool such as Label Studio or Roboflow, consistently, with a written labelling guide [7][8]. Split by scene, not by frame, so near-identical frames do not leak into validation. Start from a pretrained YOLO-family model and fine-tune on your set; a few hundred well-labelled images per class is often enough for a first version [6]. Evaluate mAP on the held-out set, then look at the false positives and false negatives by eye, which tells you what to label next.

Also asked as: how to train yolo on custom dataset · custom object detection · train object detection model on own data · object detection dataset labeling · how much data to train yolo · transfer learning object detection · annotate images for object detection · yolo custom training tutorial · roboflow yolo

Labelling quality dominates. Two hundred images with tight, consistent boxes beat two thousand with sloppy ones, and every hour spent on a labelling guide saves three hours of debugging the model.

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Why does my detector fail on real cameras?

Because the deployment camera differs from the training images: lower resolution, motion blur, different mounting angle, backlight, dust on the lens, compression from a video stream, objects smaller in the frame than in the dataset. Fix the input before touching the model: more light, a better angle, higher resolution, a cleaner stream. Then add training images from that camera. Then lower the confidence threshold for recall or raise it for precision, depending on which error costs more.

Also asked as: object detection not working on real video · yolo low accuracy real world · object detection small objects · object detection in low light · camera quality object detection · improve yolo accuracy · false positives object detection · object detection domain shift

My write-up of the camera problems that broke vision models I shipped, and the fixes, is on the portfolio [12]. The one-line summary: the camera is part of the model.

How do I run object detection in real time on a budget?

Pick a small model variant, export it to ONNX or TensorRT, quantize to 8-bit, and run at a reduced input resolution and frame rate that still catches what you need. A laptop GPU or a recent CPU handles small YOLO variants at video frame rates; a phone or browser handles them at lower rates through ONNX Runtime Web or a mobile runtime [9]. Skip frames when nothing moves, and track objects between detections instead of detecting on every frame. Measure latency on the target device, not the training machine.

Also asked as: real time object detection · object detection on cpu · yolo on raspberry pi · object detection on mobile · object detection in browser · yolo onnx · fastest object detection model · object detection low end gpu · edge object detection · yolo webcam python

My notes on real-time detection on a budget GPU, with the measured frame rates, are written up separately [11].

What are common object detection use cases?

Counting people or vehicles. Quality inspection on production lines. Safety compliance, helmets and vests on sites. Shelf and stock monitoring in retail. Crop and pest detection in agriculture. Document layout detection, finding tables and signatures on pages. Wildlife and traffic monitoring. Robotics perception. In every case the demo is easy and the deployment is the work: cameras, lighting, edge hardware and a definition of what counts as correct.

Also asked as: object detection use cases · object detection applications in real life · object detection in agriculture · object detection in manufacturing · object detection in retail · object detection project ideas · computer vision projects for students · yolo project ideas

Which tools and libraries should I use?

For training and inference: the Ultralytics YOLO library, with export to ONNX and other formats built in [6]. For labelling: Label Studio if you want open source and control, Roboflow if you want a hosted workflow with augmentation and export [7][8]. For image handling and camera capture: OpenCV [10]. For running in the browser: ONNX Runtime Web [9]. Learn the concepts once and the tools become interchangeable.

Also asked as: best object detection library · yolo vs detectron2 · ultralytics vs yolov5 · best labeling tool for object detection · label studio vs roboflow · object detection python libraries · opencv object detection · object detection frameworks

Object detection interview questions

Explain detection versus classification versus segmentation. Explain how a one-stage detector like YOLO differs from a two-stage detector. Define IoU and mAP and the difference between mAP at 0.5 and at 0.5 to 0.95. Describe how you would build a dataset and avoid leakage. Explain how you would get a detector running on a phone. Describe a failure on a real camera and the fix. Each is answered above; a detector you trained and deployed is the strongest evidence.

Also asked as: object detection interview questions · computer vision interview questions · yolo interview questions · map iou interview · deep learning interview questions vision

Where should I start?

Take fifty photos of one object type with your phone in the room where you would deploy, label them, fine-tune a small YOLO variant, measure mAP on ten held-out photos, and run the model on your webcam. Then export it to ONNX and run it in a browser tab. That weekend covers the entire page. For a hands-on computer vision session at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Redmon et al., You Only Look Once: Unified, Real-Time Object Detection (2016)arxiv.org
  2. Ren et al., Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks (2015)arxiv.org
  3. Carion et al., End-to-End Object Detection with Transformers (DETR, 2020)arxiv.org
  4. Lin et al., Microsoft COCO: Common Objects in Context (2014)arxiv.org
  5. Everingham et al., The PASCAL Visual Object Classes Challenge (2010)host.robots.ox.ac.uk
  6. Ultralytics YOLO documentationdocs.ultralytics.com
  7. Roboflow: dataset labelling and managementroboflow.com
  8. Label Studio: open-source data labellinglabelstud.io
  9. ONNX Runtime Web documentationonnxruntime.ai
  10. OpenCV documentationdocs.opencv.org
  11. Real-time object detection on a budget GPU, Pranjul Rathourpranjulrathour.scult.in
  12. Camera quality issues that break vision models, Pranjul Rathourpranjulrathour.scult.in
  13. Vision model evaluation beyond accuracy, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Vision, speech & OCR

All Vision, speech & OCR guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur