Object Detection
Object Detection
Definition: Object detection is a computer vision task that simultaneously identifies what objects appear in an image and where each one is located, typically expressed as an axis-aligned bounding box plus a class label and confidence score. It sits between image classification (which only answers “what is in this image, overall”) and segmentation (which answers “which exact pixels belong to each object”). A single image can produce zero, one, or hundreds of detections, each independently classified and localized. Modern systems process this end-to-end through a single neural network trained to output boxes, classes, and scores in one forward pass or in two coordinated stages.
How It Works
The Two Sub-Problems: Classification + Localization
Every object detector has to solve two coupled problems at once:
- Classification — for a given region of the image, decide which class (car, person, dog, background) it most likely contains.
- Localization — predict the precise coordinates of a bounding box that tightly encloses the object, usually parameterized as
(x_center, y_center, width, height)or(x_min, y_min, x_max, y_max).
These are trained jointly with a multi-task loss: a classification loss (cross-entropy or focal loss) plus a regression loss (L1, smooth-L1, or IoU-based loss) for the box coordinates.
The network shares a feature extractor between both heads because the same spatial features — edges, textures, object parts — are useful for both deciding “what” and refining “where.” Splitting the two heads too early, or giving them entirely separate backbones, tends to waste compute and hurt accuracy compared to a shared trunk with two lightweight output heads.
The Detection Pipeline
Nearly every modern detector, one-stage or two-stage, follows the same conceptual pipeline. The stages differ in how proposals are generated and how many passes are used, but the flow is consistent:
The backbone is a convolutional network (or, increasingly, a transformer) pretrained on a large image dataset; it converts raw pixels into a stack of feature maps at progressively lower spatial resolution and higher semantic abstraction. The neck, when present, merges features from different depths of the backbone so the detector can see both fine detail (good for small objects) and coarse context (good for large objects) at the same time.
The head turns each candidate region into a class prediction and a box refinement. Non-max suppression is the cleanup step that collapses redundant, overlapping boxes into one final detection per object, and it runs after the network’s forward pass as a post-processing step rather than being learned by gradient descent in most classic architectures.
A Brief History: From R-CNN to Real-Time Detectors
The field moved from “accurate but slow” toward “accurate and fast” in roughly a decade, with each generation solving the previous one’s clearest bottleneck:
- 2014 — R-CNN: runs a classical algorithm (selective search) to propose ~2000 regions, then a separate CNN feature extractor and SVM classifier per region; accurate but painfully slow (tens of seconds per image).
- 2015 — Fast R-CNN: folds the classifier into a single network and introduces ROI Pooling, so features are computed once per image instead of once per region.
- 2015 — Faster R-CNN: replaces selective search with a learned Region Proposal Network, making the whole pipeline trainable end-to-end and much faster.
- 2016 — YOLO (v1): reframes detection as one regression problem over a coarse grid, trading some accuracy for dramatic real-time speed.
- 2016 — SSD: predicts from multiple feature map resolutions simultaneously, improving multi-scale performance without a separate proposal stage.
- 2017 — RetinaNet: introduces focal loss, closing much of the accuracy gap that had separated one-stage from two-stage detectors.
- 2017 — Mask R-CNN: adds a pixel-mask prediction branch onto Faster R-CNN, extending detection into instance segmentation.
- 2020 — DETR: reframes detection as transformer-based set prediction, removing anchors and NMS entirely.
- 2020s — YOLOv5 through YOLOv10+ and successors: iterative refinement of the one-stage design, now the dominant choice for production real-time detection.
One-Stage vs Two-Stage Detectors
The single biggest architectural fork in object detection is how candidate regions are generated and refined.
Two-stage detectors (the R-CNN family: R-CNN, Fast R-CNN, Faster R-CNN, Mask R-CNN) explicitly separate “where might an object be” from “what is it.” Stage one, a Region Proposal Network (RPN), scans the feature map and proposes a few thousand candidate boxes likely to contain something. Stage two takes each proposal, pools its features to a fixed size (ROI Pooling / ROI Align), and runs a classifier plus box regressor on it.
This division of labor makes two-stage detectors more accurate, particularly on small or crowded objects, because the second stage gets to refine a much smaller, pre-filtered candidate set with dedicated capacity. The cost is speed: two full passes of network computation per image, rather than one.
One-stage detectors (YOLO, SSD, RetinaNet) skip the explicit proposal step and instead predict classes and box offsets directly from a dense grid of anchors (or, in anchor-free variants, directly from feature map locations) in a single forward pass. This trades a small amount of accuracy — especially on small objects — for a large speed advantage, since there is no separate proposal-then-refine loop. YOLO in particular treats detection as one regression problem: divide the image into a grid, and have each grid cell predict boxes and classes directly.
| Aspect | One-Stage | Two-Stage |
|---|---|---|
| Forward passes per image | One | Two (proposal, then refinement) |
| Raw speed | Faster, often real-time | Slower |
| Accuracy on small/crowded objects | Weaker without extra tuning | Stronger by default |
| Architectural complexity | Simpler, fewer moving parts | More complex (separate RPN + head) |
| Typical starting point today | Real-time products, edge deployment | Accuracy-critical offline or research pipelines |
Anchors and Anchor-Free Approaches
Most early and mid-generation detectors rely on anchor boxes: a fixed set of reference boxes at multiple scales and aspect ratios, tiled densely across the image. The network doesn’t predict raw coordinates from scratch — it predicts offsets relative to the nearest anchor, which is a far easier regression target and stabilizes training.
The downside is that anchor scales and ratios become hyperparameters that must be tuned to the dataset’s object shapes; a detector tuned for street scenes with cars and pedestrians will underperform on a dataset of tiny aerial-view objects unless anchors are re-tuned.
Anchor-free detectors (CenterNet, FCOS, and the query-based DETR family) remove this hyperparameter entirely. CenterNet predicts an object’s center point directly and regresses size from that point; FCOS predicts boxes per-pixel using distance-to-edges; DETR reframes detection as a set-prediction problem solved with a transformer and learned object queries, eliminating both anchors and NMS by using bipartite matching during training. Anchor-free approaches simplify the design space and have become the dominant direction in newer architectures.
Training Signal: Assigning Labels to Predictions
A subtlety that trips up newcomers: during training, the network doesn’t know in advance which anchor or query corresponds to which ground-truth object. A label assignment step decides this — commonly by IoU thresholding (an anchor with IoU > 0.5 against a ground-truth box is a positive example, IoU < 0.4 is background, the rest are ignored) or, in newer methods, by learned or optimal-transport-based matching.
Get this assignment strategy wrong and the loss signal becomes noisy: too many negative anchors relative to positives creates severe class imbalance, which is precisely the problem focal loss (introduced with RetinaNet) was designed to fix by down-weighting easy, correctly-classified background examples.
Data Augmentation and Multi-Scale Training
Detectors are notoriously data-hungry because every training image contains multiple objects at multiple scales, so augmentation strategy matters more here than in plain classification. Random crops, flips, color jitter, and mosaic augmentation (stitching four training images into one, popularized by YOLO) all force the network to generalize across object position, scale, and context rather than memorizing fixed layouts.
Multi-scale training — feeding the network images resized to different resolutions across training steps — directly improves robustness to the huge range of object sizes a real deployment will encounter, from a distant pedestrian a few pixels wide to a car filling half the frame. Test-time augmentation (running inference on multiple scaled/flipped versions of the same image and merging results) can further boost accuracy at the cost of extra inference time, which is why it’s common in benchmark leaderboards but rare in latency-sensitive production systems.
Open-Vocabulary and Multi-Modal Detection
Classic detectors are closed-set: they can only output classes that appeared in their training labels, fixed at training time. Adding a new class means collecting labeled data and retraining the head, which doesn’t scale to the long tail of everything a general-purpose system might need to recognize.
Open-vocabulary detection breaks this constraint by pairing a detector with a language-aligned embedding space, typically derived from a vision-language model like CLIP. Instead of predicting a fixed class index, the model predicts a region embedding and matches it against text embeddings of arbitrary class names supplied at inference time — so a user can ask the model to find “a red backpack” or “a cracked ceramic mug” without that exact phrase ever appearing in training labels. Grounding models like GLIP and Grounding DINO extend this further, accepting a free-text prompt and returning boxes for whatever the prompt describes, blurring the line between object detection and a general-purpose Natural Language Processing (NLP)-driven visual search.
From Static Detection to Video and Tracking
Everything above describes detection on a single, static image, but most real deployments run on video: a continuous stream of frames arriving 15 to 60 times per second. Running a full detector independently on every frame works, but it throws away useful information — an object detected in frame 10 is almost certainly the same object, in roughly the same place, in frame 11.
Multi-object tracking builds on top of frame-by-frame detection by adding an identity-association step: given this frame’s detections and last frame’s tracked objects, decide which boxes correspond to which existing tracks, which are new objects entering the scene, and which tracks have left. Classic approaches like SORT and DeepSORT combine a motion model (predicting where an existing track should appear next) with appearance features (to disambiguate similar-looking objects) to solve this association problem efficiently enough to run in real time alongside the detector itself.
Glossary of Core Building Blocks
A handful of terms recur across nearly every architecture description above; it helps to have them in one place:
- Backbone — the shared convolutional or transformer network that turns pixels into feature maps.
- Neck (FPN/PANet) — the layer that fuses feature maps from multiple backbone depths into a multi-scale representation.
- Anchor — a predefined reference box at a fixed scale/ratio, tiled across the image, used as a regression baseline.
- RPN (Region Proposal Network) — the sub-network in two-stage detectors that scores anchors and proposes candidate regions.
- ROI Pooling / ROI Align — operations that crop and resize a variable-sized region’s features into a fixed-size feature map for the head.
- NMS (Non-Max Suppression) — the post-processing step that removes duplicate boxes covering the same object.
- Objectness score — a prediction of “does this region contain any object at all,” used before or alongside class prediction.
- Focal Loss — a modified cross-entropy loss that down-weights easy background examples to fight class imbalance.
- IoU (Intersection over Union) — the overlap ratio used both to evaluate accuracy and to drive NMS and label assignment.
- mAP (mean Average Precision) — the standard benchmark metric combining classification and localization accuracy.
Why It Matters
- Foundational to embodied AI — any system that has to physically interact with the world (robots, drones, autonomous vehicles) needs to know not just what’s present but exactly where, in pixel and then real-world coordinates.
- Enables downstream tracking — multi-object tracking, the basis of surveillance systems and sports analytics, runs detection on every frame and links boxes across time; detection quality is the ceiling on tracking quality.
- Drove the modern deep learning benchmark culture — datasets like PASCAL VOC, COCO, and Open Images turned detection into a measurable, competitive research problem, accelerating architectural progress from R-CNN (2014) through YOLO and beyond.
- Underpins industrial automation — defect detection on manufacturing lines, produce sorting, and package inspection all rely on detectors trained on narrow, domain-specific object sets.
- Core to retail and inventory tech — shelf-scanning robots, cashier-less checkout (Amazon Go-style systems), and warehouse robotics all depend on real-time detection of products and obstacles.
- A privacy and surveillance flashpoint — face and person detection at scale raises real ethical and regulatory questions distinct from the underlying technology, which is why object detection research increasingly intersects with AI Bias and Fairness discussions.
- Medical imaging applications — tumor and lesion detection in radiology scans is directly an object detection problem, where missed detections (false negatives) carry outsized real-world cost compared to typical benchmark settings.
- Feeds the AR/VR and creative tooling stack — augmented reality apps that overlay information on real-world objects, and photo/video editing tools that auto-select subjects, both depend on fast, reliable detection.
- A proving ground for efficiency research — because detection must often run in real time on edge hardware (phones, cameras, drones), it has pushed model compression, quantization, and architecture search harder than many other vision tasks.
- A recurring input to multimodal AI systems — vision-language models and robotic planning agents increasingly use detector outputs (boxes plus labels) as a structured, symbolic bridge between raw pixels and reasoning components like an Intelligent Agent.
Evaluation: IoU and mAP
Object detection needs evaluation metrics that account for both correct classification and accurate localization — a metric like plain classification accuracy is meaningless here, since a “detection” is only useful if the box actually covers the object.
Intersection over Union (IoU)
IoU measures how well a predicted box overlaps a ground-truth box. It’s the ratio of the overlapping area to the total area covered by both boxes combined:
IoU ranges from 0 (no overlap) to 1 (perfect overlap). A predicted box is typically counted as a true positive only if its IoU with a ground-truth box exceeds a chosen threshold (0.5 is the classic PASCAL VOC standard; COCO evaluates across thresholds from 0.5 to 0.95).
IoU is also used inside NMS itself: two boxes with IoU above a suppression threshold are treated as duplicates of the same object, and the lower-confidence one is discarded. The same formula therefore shows up in three different roles across a single pipeline — training-time label assignment, inference-time duplicate removal, and evaluation-time scoring.
Mean Average Precision (mAP)
mAP is the standard headline metric for detector quality. Its construction, step by step:
- For a single class, rank all predictions by confidence score.
- Sweep down that ranked list, marking each prediction as a true positive (IoU ≥ threshold against an unmatched ground-truth box of that class) or a false positive.
- Compute precision and recall at each point in the ranked list, producing a precision-recall curve.
- Average Precision (AP) is the area under that precision-recall curve for one class.
- mAP is the mean of AP across all classes.
COCO’s mAP additionally averages AP over ten IoU thresholds (0.5 to 0.95 in steps of 0.05), which rewards detectors for tight, well-localized boxes rather than just “roughly in the right place.” This is why a model can score high mAP@0.5 but noticeably lower mAP@0.5:0.95 — it’s finding objects but not boxing them tightly.
Reporting AP separately for small, medium, and large objects (as COCO does) also surfaces a common failure mode: many detectors do well on large, centered objects and poorly on small or occluded ones, a gap that a single overall mAP number would hide.
Loss Functions in Detail
A detector’s total training loss is a weighted sum of a classification term and a localization term, computed per candidate region and summed across the image. Getting the balance and formulation of these terms right matters as much as the architecture itself — the same backbone trained with a naive loss can underperform one trained with a carefully tuned loss by a wide margin.
Classification Loss
The classification head is trained with standard cross-entropy loss over the candidate classes (plus a “background” class), which penalizes the model in proportion to how confidently wrong its predicted class distribution is. In the vast majority of anchors this loss is trivially easy — most anchors are obviously background — which is exactly the imbalance problem focal loss addresses below.
Localization Loss
Box regression is typically trained with smooth L1 loss (quadratic near zero, linear for larger errors, which keeps outlier predictions from dominating the gradient) applied to the four box offset values relative to the matched anchor. Newer detectors increasingly replace this with IoU-based losses (IoU loss, GIoU, DIoU, CIoU) that directly optimize the overlap metric the model will actually be evaluated on, rather than optimizing coordinate offsets as a proxy for it — a case where aligning the training objective with the evaluation metric measurably improves results.
Focal Loss for Class Imbalance
Focal loss modifies standard cross-entropy to down-weight well-classified examples, forcing training to focus on hard, misclassified, or ambiguous regions:
Here is the model’s estimated probability for the correct class, (commonly 2) controls how aggressively easy examples are down-weighted, and balances the relative importance of positive and negative examples. When is close to 1 (an easy, confidently-correct background prediction), the term shrinks toward zero and that example contributes almost nothing to the gradient — exactly the effect needed to stop background anchors from drowning out the sparse handful of true object anchors.
Deployment: Speed vs. Accuracy Tradeoffs
Unlike many benchmark-driven vision tasks, object detection is disproportionately deployed in real-time, resource-constrained settings — a camera on a drone, a chip inside a car, a phone’s AR app — which makes the speed/accuracy tradeoff a first-class design decision rather than an implementation detail.
| Detector Family | Typical Speed Profile | Typical Accuracy Profile | Common Deployment Target |
|---|---|---|---|
| Two-stage (Faster R-CNN, Mask R-CNN) | Slower — multiple network passes per image | Highest, especially on small or crowded objects | Server-side batch analysis, offline processing, research |
| One-stage anchor-based (YOLO, SSD, RetinaNet) | Fast, often real-time on a single GPU | Slightly below two-stage, narrowed by focal loss | Real-time video, robotics, edge devices with a GPU/NPU |
| Anchor-free / transformer-based (FCOS, DETR) | Varies; DETR-style models historically need longer training and more compute | Competitive with or exceeding CNN detectors given enough data | Research, and increasingly production as tooling matures |
Common techniques for shrinking a detector to fit a latency or power budget:
- Quantization — running the network in INT8 instead of FP32, trading a small accuracy hit for large speed and memory gains.
- Pruning — removing redundant channels, filters, or entire layers that contribute little to accuracy.
- Knowledge distillation — training a small “student” detector to mimic a larger, more accurate “teacher” model’s outputs.
- Export to specialized runtimes — converting a trained model to ONNX, TensorRT, or TensorFlow Lite so it can run on hardware-specific inference engines rather than a general research framework.
Throughput requirements vary enormously by use case, and the required frame rate should be decided before an architecture is chosen, not after:
- Offline batch analysis (e.g., tagging a photo archive) — no strict latency requirement; two-stage detectors are often the right default.
- Recorded video review (e.g., security footage search) — moderate throughput, often acceptable at well under real-time speed.
- Live video analytics (e.g., retail or crowd monitoring) — needs to sustain roughly 15-30 frames per second continuously.
- Interactive robotics and AR — needs 30-60 FPS with low, consistent per-frame latency, since jitter is as damaging as raw slowness.
- Autonomous vehicle perception — needs sustained high frame rate with strict worst-case latency bounds, because a single missed or late detection has safety consequences.
Choosing a two-stage detector for the last two categories, however accurate on paper, is a design error that only shows up once the system is under real-time load.
Notable Datasets and Benchmarks
Progress in object detection has always been benchmark-driven — architectures are typically introduced alongside a specific dataset result, and the datasets themselves have shaped what “good” detection means:
- PASCAL VOC — an early, relatively small benchmark (20 classes) that established the IoU@0.5 mAP evaluation convention still referenced today.
- COCO (Common Objects in Context) — the dominant modern benchmark; 80 classes, images with many objects per scene, heavy occlusion, and the stricter mAP@0.5:0.95 metric.
- Open Images — a much larger-scale dataset (millions of images, thousands of classes) used to test detection breadth rather than just accuracy on a narrow class set.
- LVIS — designed specifically to test performance on rare, long-tail categories rather than the common objects that dominate COCO.
- KITTI — an autonomous-driving-focused benchmark with car, pedestrian, and cyclist classes captured from a vehicle-mounted camera.
- BDD100K — a large-scale driving dataset covering diverse weather, time-of-day, and geographic conditions, used to stress-test domain robustness.
- Cityscapes — urban street-scene imagery, more commonly used for segmentation but also widely used for detection in autonomous driving research.
- Custom domain-specific datasets — in practice, most production detectors end up fine-tuned on a narrow, in-house labeled dataset rather than evaluated purely on these general-purpose benchmarks.
Comparison
| Task | Output | Question Answered |
|---|---|---|
| Image Classification | A single label (optionally with confidence) for the whole image | “What is the main subject of this image?” |
| Object Detection | A set of bounding boxes, each with a class label and score | “What objects are present, and roughly where is each one?” |
| Semantic Segmentation | A class label for every pixel, with no distinction between instances | “Which pixels belong to which class, ignoring individual objects?” |
| Instance Segmentation | A pixel-precise mask per object instance, each with a class label | “Which exact pixels belong to this specific car versus that car?” |
| Panoptic Segmentation | Every pixel labeled, combining per-instance masks for objects with per-class masks for background “stuff” | “What is every pixel, and which specific instance does it belong to when that’s meaningful?” |
Detection is the middle ground: coarser than segmentation (a box, not a mask) but far more localized and multi-object-aware than plain classification. Instance segmentation is often built as a detection model with an added mask-prediction branch (Mask R-CNN is the canonical example) — proof that these tasks form a natural hierarchy rather than four unrelated problems.
Real-World Use Cases
- Autonomous driving perception stacks — detecting vehicles, pedestrians, cyclists, and traffic signs from camera (and often fused LiDAR) input, frame by frame, as a safety-critical real-time system.
- Retail cashier-less checkout — camera arrays detecting which products a shopper picks up and puts back, replacing manual scanning.
- Warehouse and logistics robotics — bin-picking robots detecting individual items in cluttered bins, and mobile robots detecting obstacles and pallets.
- Video surveillance and security — person and vehicle detection for perimeter monitoring, restricted-area alerts, and after-the-fact forensic search through footage.
- Manufacturing quality control — detecting scratches, dents, missing components, or misaligned parts on a production line at line-speed.
- Agricultural tech — drone or ground-robot detection of ripe fruit for automated harvesting, or detection of crop disease and weeds for targeted treatment.
- Medical imaging — detecting tumors, nodules, or fractures in X-rays, CT scans, and MRIs as a first-pass triage tool for radiologists.
- Wildlife and conservation monitoring — camera-trap footage analyzed automatically to count and identify species without manual review of thousands of hours of video.
- Sports analytics — detecting players and the ball across broadcast video to generate live statistics, heatmaps, and highlight reels.
- Document and receipt processing — detecting tables, signatures, stamps, and line items within scanned documents before downstream text extraction.
- Satellite and aerial imagery analysis — detecting buildings, vehicles, or vegetation change across large geographic areas for urban planning, disaster response, and defense applications.
- Assistive technology — real-time scene description for visually impaired users, where detected objects and their positions are converted to spoken guidance.
Common Pitfalls
- Confusing detection with segmentation — detection gives boxes, not pixel masks; teams sometimes reach for a detector when the actual product need (e.g., precise background removal) requires segmentation instead.
- Class imbalance between foreground and background — the overwhelming majority of anchors or grid cells in any image correspond to background, and without techniques like focal loss or hard-negative mining, the model learns to trivially predict “background” everywhere.
- Poor performance on small objects — small objects occupy few pixels by the time features reach deep layers of the backbone, and get further degraded by downsampling; this is why feature pyramids and multi-scale training exist.
- Anchor mis-tuning — using default anchor scales/ratios from a general-purpose dataset (COCO) on a domain with very different object shapes (e.g., long thin conveyor-belt items) silently caps recall no matter how much training data is added.
- Overlapping and occluded objects collapsing into one detection — aggressive NMS thresholds can suppress a legitimately separate object that happens to heavily overlap another (e.g., a crowd of people); soft-NMS and better-tuned thresholds mitigate but don’t eliminate this.
- Domain shift at deployment — a detector trained on daytime, clear-weather images degrades sharply on night, fog, rain, or a different camera sensor unless that variation was represented in training data.
- Treating mAP as the only metric that matters — a model can have excellent mAP while still failing badly on the specific object sizes, lighting conditions, or classes that matter most for the actual product; per-class and per-size breakdowns matter more than the headline number.
- Label noise and inconsistent annotation — bounding box tightness varies by annotator, and inconsistent labeling (loose boxes vs. tight boxes) directly degrades the localization loss signal during training.
- Ignoring inference latency constraints until too late — a two-stage detector with excellent accuracy can be entirely unusable on an embedded camera with a real-time frame-rate requirement; the accuracy/speed tradeoff has to be a design decision from the start, not an afterthought.
- Data leakage between train and test splits — near-duplicate frames from the same video clip ending up in both training and validation sets inflates reported metrics without reflecting real generalization.
- Skipping calibration of confidence scores — a detector’s raw confidence values aren’t automatically well-calibrated probabilities; downstream systems that threshold on a fixed score cutoff can behave unpredictably across different scenes without explicit calibration.
Related Terms
- Computer Vision
- Convolutional Neural Network (CNN)
- Neural Network
- Loss Function
- Supervised Learning
- Overfitting vs Underfitting
- AI Bias and Fairness
- Intelligent Agent
Example
A logistics company builds an automated package-sorting line: packages move down a conveyor belt under an overhead camera, and a detector has to identify each package’s location and orientation so a robotic arm can pick it up and route it. The team starts with a pretrained YOLO-style one-stage detector, chosen over a two-stage R-CNN model specifically because the conveyor runs at a speed that requires inference well under 50 milliseconds per frame — accuracy that can’t be delivered in that time budget is worthless on this line.
They fine-tune it on a few thousand labeled images of their actual packages, since the pretrained model’s default classes (trained on COCO’s everyday-object categories) don’t include “cardboard box” or “poly mailer” as distinct, tightly-boxed classes. Early results look strong on paper — mAP@0.5 above 90% — but the arm keeps missing badly-oriented, partially overlapping packages, which turn out to be under-represented in the training set relative to how often they occur in production.
The team adds targeted training images of overlapping and rotated packages, switches from standard NMS to soft-NMS to stop legitimately adjacent boxes from being suppressed as duplicates, and re-tunes anchor aspect ratios to better match the mostly-rectangular, often-elongated shape of their packages rather than the general-purpose defaults. After this second pass, per-size and per-orientation breakdowns — not just the single headline mAP number — become the team’s primary tracking metric, since that’s what actually predicts pick success on the floor. The revised model is exported to TensorRT and runs on an edge GPU mounted above the line at the required frame rate, and the arm’s mis-pick rate drops enough to clear the line’s throughput target.
Referenced by