Signal archive

Signal

Issue 11 · September 22, 2026 · 6 min read

Object Detection Is Moving to the Edge. Learn It Now.

Object detection explained: IoU, NMS, mAP, YOLO vs Faster R-CNN, and why over half of new vision deployments run on edge devices. Build one this week.

"To take pictures is not the same as to see." - Fei-Fei Li (Stanford, TED 2015)

A camera that classifies "forklift" is a demo. A camera that says where the forklift is, 30 times a second, on a $200 board bolted to a wall, is a product.

This week we're talking about object detection, the computer vision task that pays because it runs where the cameras are, and yes, this can get you a very comfortable job.

Here's the idea: object detection is the computer vision task of finding every instance of a class in an image and returning a bounding box, a label, and a confidence for each one. Image classification outputs one label per image. A detector outputs a variable-length list. That difference is why detection needs its own matching rules, losses, and metrics, and why most people who "know CNNs" cannot ship one.

The metric everything hangs on is intersection over union: A predicted box counts as correct only if its IoU with a ground-truth box clears a threshold, usually 0.5. COCO's headline number, mAP@[.5:.95], averages precision across IoU thresholds from 0.5 to 0.95 in steps of 0.05, so sloppy boxes that would pass at 0.5 get punished at 0.75. Before any of that, non-maximum suppression (NMS) has to collapse the dozens of overlapping boxes a detector fires on one object down to a single survivor. IoU, NMS, and mAP are the three things you will be asked to implement on a whiteboard.

The market signal is about where this code runs. Datature's 2026 Enterprise Vision AI Adoption Report puts more than 50% of new enterprise computer vision deployments on edge devices, up from roughly 30% in 2023. That is the shift we covered in the edge computing for robotics issue, and detection is its workhorse. The YOLO family owns that niche because it is a single forward pass: YOLOv5 launched in 2020 with a 140 frames-per-second claim, and current YOLO models run on Jetson-class boards at video rate. Frigate, an open-source network video recorder with real-time detection, was one of the most upvoted computer vision stories ever on Hacker News. Industry skills lists for computer vision engineers in 2026 name YOLO, edge AI deployment, and CUDA optimization next to PyTorch. The job is not "train a detector." It is "train a detector, quantize it, and hit 30 FPS at under 10 watts."

How does object detection work?

A modern detector has three parts. A backbone, usually a convolutional network or a vision transformer, turns the image into feature maps at several scales. A neck fuses those scales so small and large objects both get features. A head predicts, at every location on the grid, a box offset, an objectness score, and class probabilities. Two-stage detectors (Faster R-CNN) first propose regions and then classify them. One-stage detectors (YOLO, RetinaNet, DETR-style transformers) predict boxes directly, which is why they are faster. Training matches each ground-truth box to predicted boxes by IoU, applies a regression loss to the box coordinates and a classification loss to the labels, and NMS cleans up at inference.

What is the difference between object detection and image classification?

Classification answers "what is in this image" with one label. Detection answers "what is where" with a box per object. Segmentation goes further and labels every pixel.

Image classificationObject detectionInstance segmentation
OutputOne label per imageBox + label per objectPixel mask + label per object
Output lengthFixedVariableVariable
Core lossCross-entropyBox regression + classificationBox + classification + mask
Headline metricTop-1 accuracymAP over IoU thresholdsMask mAP
Typical modelResNet, ViTYOLO, Faster R-CNN, DETRMask R-CNN, SAM-based
Labeling cost per imageSecondsMinutesTens of minutes

Now I'm going to explain what this means for you:

The Skill Employers EXPECT

Learn the detection primitives by hand. Computer vision teams assume you can:

  • Implement IoU for a batch of boxes in NumPy with correct handling of non-overlapping pairs
  • Implement NMS and explain what the IoU threshold does to crowded scenes
  • Compute mAP@0.5 and mAP@[.5:.95] yourself and match a reference implementation
  • Read and convert COCO and YOLO annotation formats (corner vs center-width-height, absolute vs normalized)
  • Apply flips, crops, and mosaic augmentation without corrupting the boxes
  • Split a video dataset by scene, not by frame, so adjacent frames do not leak across train and test

If you cannot explain why your mAP dropped when you raised the NMS threshold, you are guessing.

The Skill That Separates You

Learn to deploy a detector under a latency and power budget.

Anyone can fine-tune YOLO in a notebook. Teams pay for the person who can put it on a device. You should know how to:

  • Export to ONNX or TensorRT and verify outputs match PyTorch within tolerance
  • Quantize to INT8 and measure the mAP cost against the latency gain
  • Profile latency at batch size 1 on the target CPU or Jetson, not on a datacenter GPU
  • Handle small objects: input resolution, feature pyramid choice, and anchor or stride settings
  • Add tracking-by-detection so IDs persist across frames and the downstream system gets counts, not boxes

The Project To Learn These Skills This Week

Build the detection metrics stack from scratch, fine-tune a small detector, and put it on a CPU budget. Dataset: Penn-Fudan Pedestrian (170 images, boxes and masks, used in the torchvision tutorial) or any small COCO-format dataset with 2 to 5 classes.

Requirements:

  • Implement iou, nms, and mAP (at 0.5 and at [.5:.95]) in NumPy. Match torchvision.ops and torchmetrics within 1e-6 for IoU and NMS and within 0.01 mAP.
  • Fine-tune a lightweight detector (Faster R-CNN MobileNetV3 or a YOLO-nano variant) to at least 0.6 mAP@0.5 on a held-out split.
  • Ablation: sweep the NMS IoU threshold over {0.3, 0.5, 0.7}. Plot precision and recall at each and explain the crowded-scene trade-off.
  • Export to ONNX. Measure median latency at batch size 1 on your laptop CPU. Then quantize to INT8 and report latency and mAP again.
  • Break it on purpose: re-split the data randomly by frame if your dataset comes from video, or by near-duplicate images if it does not. Show the inflated mAP, then fix the split by scene.
  • Write a half-page deployment memo for a Jetson-class edge box: model, input resolution, expected FPS, mAP you would promise, and what fails when the camera moves or the lighting changes.

This project teaches a practical lesson for production ML:

A detector is judged on where it runs and what it misses, not on the notebook where it trained.

If you want to sharpen your machine learning skills even more, I also selected a challenge problem for you this week:

If you learned something from this newsletter, make sure to forward it to a friend.

Mean IoU from a Confusion Matrix

Easy · ~12 min

Concept: illumination robust segmentation

Ready to practice?

Turn weekly insights into hands-on ML skills on GRADuateML.