ECCV 2026

O‑VAD: Industrial Video Anomaly Detection through Object‑Centric Tracking and Reasoning

A training‑free agentic framework  ·  ground → track → reason

Mei Yuan1, Qi Long1, Qifeng Wu1, Zhenyang Li2, Yizhou Zhao1, Lei Wang3, Yang Liu1, Min Xu1
1Carnegie Mellon University  ·  2University of Alabama at Birmingham  ·  3Griffith University
Paper Code
O-VAD teaser: a toothpaste tube filmstrip across pressing, deformation and leakage states, with GPT-5 reporting no anomaly while O-VAD produces a grounded leakage report

Figure 1. A toothpaste tube undergoes pressing and deformation before a subtle leakage emerges. Traditional VAD predicts a binary label and VLM prompting misses the event; O‑VAD tracks object‑wise state changes and produces an open‑ended anomaly report with grounded frames and causal analysis.

01 / Abstract

Reasoning over object state evolution, like a human inspector

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality‑control systems. Existing VLM‑based anomaly‑reasoning methods can detect open‑ended anomalies in general domains, but their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints.

To tackle the complexity of such interaction‑intensive detection, we introduce O‑VAD, a training‑free agentic framework free of domain‑specific knowledge that emphasizes object state evolution. It tracks the spatial‑temporal dynamics and underlying transformations of detected objects over time, then reasons over the object‑wise temporal state trajectories to identify abnormal objects in grounded frames — overcoming prior approaches that rely on retraining on normal clips or injecting domain knowledge at test time.

Extensive experiments on three IVAD datasets show O‑VAD outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine‑tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.

Training‑free — no fine‑tuning, no domain knowledge, no predefined taxonomy.

Best video‑level AUROC on every benchmark

0.584
Phys‑AD video AUROC
73%1st/2nd · 16/22 categories
0.692
LiquidAD video AUROC
8 pipette types
0.565
IPAD video AUROC
75%1st/2nd · 12/16 scenes

02 / Method

A ground → track → reason pipeline

Given an industrial video, O‑VAD produces a structured anomaly report — type, severity, affected object, temporal localization, and a natural‑language causal explanation — without any domain‑specific knowledge, predefined taxonomy, or training data. It runs in three agentic stages.

O-VAD framework: (a) Object Grounding via VLM inventory and SAM masks, (b) Object-Centric State Tracking with spatiotemporal partitions and state-change detection, (c) State-Aware Anomaly Reasoning producing the anomaly report
1

Object Grounding

Sample frames spanning the video, query a VLM for a structured object inventory (name, material, initial state), then segment every object with SAM3 to produce initial masks and metadata.

GPT‑5SAM3
2

Object‑Centric State Tracking

Build spatiotemporal tubelets with CropFormer entity segmentation and SAM2 propagation. Recover tracks lost through transformations using spatial‑proximity and CLIP semantic‑consistency priors, then query the VLM for open‑ended state‑change events — type, cause, severity, per object.

CropFormerSAM2CLIP
3

State‑Aware Reasoning

A cascaded six‑step chain‑of‑thought reasons over the accumulated state trajectories, separating expected process actions from failure outcomes.

1Process understanding
2Observation
3Expectation
4Comparison
5Causation
6Classification & severity
+ confidence‑gated visual verification

03 / Quantitative results

State of the art without any training

We foreground the metrics that matter for open-ended IVAD — detection AUROC (video and frame level) and anomaly-type description quality — and drop Acc/P/R/F1. O‑VAD posts the best video-level AUROC on all three datasets and the most faithful type descriptions, then leads or ties across most Phys-AD and IPAD categories.

Comprehensive resultsdetection AUROC + anomaly‑type score · video & frame level
Phys-ADLiquidADIPAD
MethodVid AUROC ↑BERTScore ↑LLM‑judge ↑Vid AUROC ↑Frm AUROC ↑Vid AUROC ↑Frm AUROC ↑
Traditional VAD
MNAD.p0.4950.4530.4870.5220.430
S3R0.5550.6510.6250.5390.495
Direct prompting
Qwen3‑VL0.5130.7980.3720.4890.4990.4980.497
GPT‑50.5020.8780.5800.4650.4840.5190.532
Agentic workflow
URF0.4260.3650.5060.4170.512
VERA0.4560.5340.5000.5190.490
Ours
O‑VAD0.5840.8030.5950.6920.5120.5650.518
Phys-AD · per-category AUROCvideo level ↑ · 22 object categories
CategoryMNAD.pS3RQwen3‑VLGPT‑5URFVERAO‑VAD
Ball0.5890.1550.5280.7440.4610.5480.579
Button0.5260.1080.5060.4560.7170.4690.393
Car0.5990.5270.4920.6170.4970.3830.575
Caster Wheel0.4680.8300.5330.5430.4100.6090.301
Clip0.6280.5370.4420.5180.2750.1800.423
Clock0.5130.5290.5810.7660.4830.5000.669
Fan0.4090.6020.5660.4720.5230.3890.669
Gear0.4180.3580.6440.3970.4510.5120.879
Hinge0.4970.7160.5890.5160.5440.0410.526
Liquid0.4530.7330.8330.5120.6910.8270.776
Lock0.4420.3230.3880.2310.0290.5570.780
Magnet0.5120.7240.5920.3640.5180.8920.576
Rolling Bearing0.4370.0160.4180.4510.3040.5000.650
Rubber Band0.5610.4630.7670.2790.7830.2760.721
Screw0.6380.6200.4500.4410.2980.4000.704
Servo0.4790.4520.4920.3300.3420.3410.495
Slide0.5340.5330.5290.5280.3450.2250.415
Spherical Bearing0.5470.8940.4830.4400.6290.1140.641
Sticky Roller0.5170.7780.5330.5230.7130.4740.811
Toothpaste0.2570.5350.5780.8320.5280.5790.689
U Disk0.5130.5850.5780.4370.4040.3110.606
Zipper0.2620.7910.5170.4770.2730.2240.388
Average0.4950.5550.5130.5030.4260.4560.584
IPAD · per-scenario AUROCvideo level ↑ · 16 scenarios (S synthetic, R real)
ScenarioMNAD.pS3RQwen3‑VLGPT‑5URFVERAO‑VAD
S010.5090.5090.6250.2750.3270.5000.396
S020.3750.7160.5750.4380.4430.4500.688
S030.6000.2670.3250.6000.5110.5250.533
S040.3890.1330.6750.6250.3110.4000.389
S050.5000.2400.5500.3120.4000.3380.570
S060.3780.1780.3920.1730.6670.2870.794
S070.5560.3440.5500.6750.3330.5120.622
S080.5670.4440.7000.6880.2330.5000.639
S090.3520.5230.4750.5000.4770.6250.580
S100.4440.5670.5500.6130.2670.6250.433
S110.5500.0000.5500.5500.6560.6500.467
S120.3670.1440.5000.7620.2000.4500.539
R010.4290.5000.6000.5000.5890.5000.679
R020.1150.5000.5250.6250.3080.5250.712
R030.3000.0670.4790.4900.3670.4900.567
R040.5000.4000.6000.1900.5000.500
Average0.5220.5390.4980.5190.4170.5190.565

Higher is better; per row the best is boxed and the 2nd-best shaded; O‑VAD shown in accent. Phys-AD has no frame-level labels; LiquidAD and IPAD carry no free-form type score.  trained   training-free.

04 / Qualitative results

Grounded reports where frontier VLMs miss the anomaly evidence.

We compare O‑VAD against GPT‑5 prompted with the same six‑step chain‑of‑thought but without object grounding or state tracking. Lacking object‑wise evidence, GPT‑5 defaults to “no anomaly”; O‑VAD localizes the object, frame range, severity, and root cause.

4.1 / CASE COMPARISONS

Case 1
Plastic bottle rotation
GPT‑5

“No anomaly detected.” Reports a secure grip, controlled rotation, and no leakage.

O‑VAD

Loss‑of‑containment leakage · water bottle · frames 60–80 · high severity. Progressive clamp‑induced deformation precedes material release.

Full reasoning traces for the plastic water-bottle case: GPT-5 answers no anomaly detected while O-VAD reports loss-of-containment leakage at frames 60-80 with objects, per-frame observations, expectation, comparison, causation and classification.⤢ Full size
Case 2
Hinge screw fastening
GPT‑5

“No anomaly detected.” Concludes a normal tightening sequence, screw fully seated.

O‑VAD

Three co‑occurring defects — head stripping (80–90), shaft bending (60–100), repeated back‑out — attributed to torque‑control / bit‑mismatch failure · high severity.

Full reasoning traces for the hinge-screw case: GPT-5 answers no anomaly detected while O-VAD reports head stripping, shaft bending and repeated back-out at frames 80-90 with the full structured trace.⤢ Full size
Case 3
Sticky roller press
GPT‑5

“No anomaly detected.” Reads the roller rotating then stopping as normal actuation — no damage.

O‑VAD

Core & roll crushing · cardboard tube + paper roll · frames 60–70, 140–170 · high severity. Repeated crushing/indentation with a surface tear at the roll end.

Full reasoning traces for the sticky-roller case: GPT-5 answers no anomaly detected while O-VAD detects crushing of the cardboard core and paper roll with a surface tear, at frames 60-70 and 140-170.⤢ Full size

4.2 / Failure cases

Where perception‑driven reasoning still falls short.

Without predefined expert knowledge, O‑VAD misses two families of anomaly: specification‑dependent ones whose ground truth is a quantitative threshold invisible to inspection, and perception‑ambiguous ones whose defining evidence never appears in the frame. Four representative cases:

Servo motorSpecification-dependent

A restricted rotation angle is labeled normal by both O‑VAD and the VLM baseline — the small‑angle motion is visually indistinguishable from a normal commanded rotation.

Full reasoning traces for the servo case: both Qwen3-VL-32B and O-VAD answer no anomaly detected for a servo with a restricted rotation angle.⤢ Full size
Degaussed magnetInvisible property

O‑VAD detects the over‑clamping deformation as the magnet drops, but misattributes the cause: a demagnetized magnet looks identical to an ordinary red plastic block, so loss of magnetism is unobservable.

Full reasoning traces for the degaussed-magnet case: the magnet drops from the gripper; O-VAD reports over-clamping deformation but does not identify loss of magnetism as the root cause.⤢ Full size
Unpressable clipNo observable evidence

A clip that cannot be pressed produces no observable state change across frames, so both methods default to a normal prediction.

Full reasoning traces for the clip case: both Qwen3-VL-32B and O-VAD answer no anomaly detected for a spring clip that cannot be pressed and shows no observable state change.⤢ Full size

Challenging categories. O‑VAD’s weakest Phys‑AD AUROC lands on Caster Wheel (0.301), Zipper (0.388), Button (0.393), and Slide (0.414) — subtle rotational resistance, small localized transitions missed at the default sampling interval, and alignment hard to judge from cropped views. These motivate adaptive temporal sampling and multi‑scale spatial reasoning.

4.3 / Human evaluation & LLM-as-judge

Experts and an LLM judge agree: the reports are usable

Beyond automated metrics, five domain experts rated report quality on 10 Phys-AD videos (5-point Likert), and a GPT-4o judge mirroring the same questionnaire scaled the study to the full test set. O‑VAD leads every dimension of both, against direct prompting (Qwen3-VL-32B) and the agentic URF-ZS-HVAA.

Part I · Detection
Q1 Binary · Q2 Type · Q3 Object · Q4 Temporal
Part II · Explanation
Q5 Faithfulness · Q6 Completeness · Q7 Causal coherence
Parts III–V · Utility
Q8 Actionability · Q9 Overall · Q10 Preference
Human evaluation
5 domain experts · 10 Phys-AD videos · 5-point Likert
DimensionO‑VADQwen3‑VL‑32BURF‑ZS‑HVAA
Part I — Detection correctness
Binary detection (Q1)
4.98
2.701.24
Anomaly type (Q2)
4.26
2.601.08
Temporal localization (Q4)
3.96
2.321.08
Detection average
4.42
2.891.14
Part II — Explanation quality
Completeness (Q6)
4.58
3.361.00
Causal coherence (Q7)
4.30
3.021.00
Parts III–V — Utility
Actionability (Q8)
4.16
2.841.00
Ranked #1 (Q10)
82%
LLM-as-judge (GPT-4o)
full test set · mirrors the human questionnaire
DimensionO‑VADQwen3‑VL‑32BURF‑ZS‑HVAA
Part I — Detection correctness
Binary detection (Q1)
4.60
2.601.40
Detection average
3.20
2.331.13
Part II — Explanation quality
Faithfulness (Q5)
3.90
4.90
Completeness (Q6)
3.90
4.00
Causal coherence (Q7)
3.60
3.70
Parts III–V — Utility
Diagnostic usefulness (Q8)
3.40
2.301.10
Overall quality (Q9)
3.50
2.701.20
Ranked #1 (Q10)
70%
30%10%

Mean ratings (1–5, higher better) unless marked %; O‑VAD in accent with a fill bar; “–” = URF-ZS-HVAA near-floor value not itemized in the paper. Q3 and Q5 human per-method scores are reported only as part-level aggregates.

05 / Ablations

Every component earns its place.

Removing one module at a time on four representative Phys‑AD subsets isolates each contribution.

State tracking is indispensable

Removing it collapses precision, recall, and F1 to zero on three of four subsets — the system predicts everything “normal.”

Captioning is a global prior

Scene semantics guide holistic anomalies; dropping captions costs up to 0.25 AUC on sticky‑roller and rubber‑band.

Structured CoT improves calibration

The six‑step chain lifts AUC by 0.04–0.09 over a generic “think step by step” prompt.

Post‑verification is a safeguard

Multiplicative confidence gating sharpens borderline cases, adding up to 0.18 AUC on sticky roller.

Ablation on Phys-AD subsets video-level AUROC ↑
VariantSticky rollerLiquidScrewRubber band
O-VAD (full)0.8110.7760.7040.721
w/o Caption0.5640.7420.7310.590
w/o State Tracking0.4840.6670.4000.587
w/o CoT Reasoning0.7200.7380.6640.681
w/o Post-verifier0.6330.7180.6220.714

Full model in bold; orange marks a drop from the full model. Removing state tracking additionally collapses F1 to 0.000 on three of four subsets.

06 / DEMO

Method walkthrough

A short overview of the ground → track → reason pipeline and representative anomaly reports.

1Stage 1 · grounding mask
Stage 1 SAM grounding masks on a LiquidAD frame: eight pipettes and two trays each segmented as a distinctly colored instance.
2Stage 2 · state tracking

Grounding → tracking on LiquidAD. Stage 1 assigns one mask per object — eight visually identical pipettes plus the trays. Stage 2 propagates each into a stable tubelet, holding identity through partial occlusion by the dispensing head and during liquid transfer.

3Stage 3 · anomaly reportO‑VAD output · JSON
Scene description

The multichannel pipette starts static above trays of colored liquids, then moves down to the red tray, aspirates, lifts, and translates to a clear well tray to dispense before retracting — a smooth, consistent pipetting cycle.

!Anomaly detected Severityhigh Confidence0.86 Anomalies1
manipulation_failureanomaly_0000 · obj_id "1"
Subtype  single‑channel pipetting failure — aspiration / retention / dispensing

The 4th pipette tip (obj_id 1) repeatedly fails to aspirate, retain, or dispense while the other channels in the same head show normal liquid columns. Evidence: strong within‑type difference during aspiration (frames 10–40) and missing red liquid in that tip through dispensing (frames 130–240) — a persistent channel‑level malfunction.

Show full JSON report
{
  "caption": "Firstly, the multichannel pipette is positioned above a set of trays containing different colored liquids. The pipette is STATIC at this point. Then, the pipette begins to MOVE downward towards the red liquid tray. As it descends, the pipette tips make contact with the liquid, indicating the start of liquid uptake. The pipette remains STATIC momentarily as it draws the liquid. After that, the pipette begins to MOVE upward, separating from the liquid surface. The pipette continues to MOVE horizontally towards a clear tray with wells. Upon reaching the tray, the pipette tips descend again, making contact with the wells. The pipette remains STATIC as it dispenses the liquid into the wells. Following this, the pipette retracts and MOVES back to its original position above the colored trays. Throughout the sequence, the pipette exhibits smooth and consistent movements without any unusual or anomalous behaviors.",
  "anomaly_detected": true,
  "num_anomalies": 1,
  "overall_severity": "high",
  "overall_confidence": 0.86,
  "anomalies": [
    {
      "anomaly_id": "anomaly_0000",
      "anomaly_type": "manipulation_failure",
      "anomaly_subtype": "single_channel_pipetting_failure (aspiration/retention/dispensing)",
      "severity": "high",
      "description": "The 4th pipette tip/channel (obj_id \"1\") repeatedly fails to aspirate/retain and/or dispense liquid while the other pipette tips in the same multichannel head show normal liquid columns and dispensing behavior. Evidence includes severe difference_within_same_type_objects during aspiration (frames 10-20, 20-30, 30-40) and during red-liquid handling/dispense where the 4th tip lacks visible red liquid while others have it (frames 130-140, 140-150, 150-160, 160-170, 180-240). This indicates a persistent channel-level liquid handling malfunction.",
      "affected_objects": [
        "1"
      ]
    }
  ]
}

07 / Citation

BibTeX

@inproceedings{yuan2026ovad,
  title     = {O-VAD: Industrial Video Anomaly Detection
               through Object-Centric Tracking and Reasoning},
  author    = {Yuan, Mei and Long, Qi and Wu, Qifeng and Li, Zhenyang and
               Zhao, Yizhou and Wang, Lei and Liu, Yang and Xu, Min},
  booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
  year      = {2026}
}