NeurIPS 2026  ·  Evaluations & Datasets Track

OmniCapBench A Deep-Structured Evaluation Framework for
Fine-Grained Audio-Visual Captioning

Zhongyu Yang1,♥ Jiale Tao1,♥,† Ruitao Chen1,♥
Zuhao Yang2 Yingfang Yuan3 Xueliang Zhao1 Auden1 Kai Wang1 Shuai Shao1
Biao Wang1,✉ Steve Yves1,✉ Qinglin Lu1
1 Hunyuan, Tencent 2 Nanyang Technological University 3 Northumbria University
♥ Equal Contribution † Project Leader ✉ Corresponding Author
786videos
12.8 hduration
20 / 125categories (L1/L2)
5,818references
11,419shots
39,160subshots
6,537audio events
Comparison of holistic evaluation, probe-based evaluation, and the OmniCapBench deep-structured evaluation framework.
Existing evaluation frameworks vs. the deep-structured evaluation framework. (A) Holistic evaluation and (B) probe-based evaluation both derive scores from unstructured, free-form captions. The former treats the caption as a whole, producing an opaque score that masks local errors; the latter probes local errors with QA but sacrifices global coverage and is unstable across judges. (C) OmniCapBench (Ours) natively represents the caption as atomic, verifiable audio-visual units. Deterministic rules verify structural and temporal relationships, then dispatch aligned units to localized LLMs for bounded semantic checks — preserving coverage and fine-grained error localization while removing the uncertainty of global LLM judges.

Abstract

Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio–visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio–visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio–visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio–visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.

Contributions

Deep-Structured Evaluation Paradigm

We decompose audio–visual captioning into atomic, verifiable units, explicitly isolating semantic content, temporal grounding, identity tracking, and cross-modal association.

The OmniCapBench Benchmark

A rigorous testbed of 786 densely annotated videos. Native Reference, Shot and Event tracks yield 5,818 entities, 6,537 audio events and 11,419 visual shots for fine-grained omnimodal diagnosis.

Empirical Audit of Holistic Scoring

We expose its vulnerability to judge instability and its tendency to mask localized errors — and reward structurally flawed outputs — demonstrating the necessity of deep-structured metrics.

Method

Deep-Structured Evaluation

For a video V, OmniCapBench formalizes content as a native reference system S*(V) = (R, E, H) — three interdependent tracks that separate what happened from when and where it occurred.

Reference track  R

Persistent identities

Unique identifiers for key entities — people, objects and scenes — enabling consistent tracking of each entity throughout the video.

PersonObjectScene
Event track  E

Timestamped audio

Auditory events and spoken dialogue with precise, continuous temporal boundaries (e.g. a dog bark from 10.5 s to 12.0 s).

SoundDialogueTime range
Shot track  H

Visual timeline & grounding

Segments the visual timeline into shots and subshots, and acts as the unifying structure that anchors references and audio events to specific frames.

ShotSubshotCross-link

Scoring contract: structure first, semantics second

1

Native unit generation

The model receives the video and directly emits the three prediction tracks as a structured JSON object, rather than a free-form paragraph — removing the ambiguity of monolithic text.

2

Structural verification & matching

Deterministic rules reject undefined IDs, malformed time spans and impossible event–shot links. Surviving units are matched by tIoU, with a bounded LLM matcher for references and subshots.

3

Localized semantic scoring

Only structurally aligned unit pairs reach the LLM judge, which is restricted to local semantic equivalence on isolated fields — bidirectionally, so omissions hurt recall and hallucinations hurt precision.

Visual

  • Ref Subject / Scene F1 — entity detection
  • RefUse — correct persistent ID cited
  • CCC — cross-shot coreference consistency
  • Shot / Subshot F1, tIoU — temporal grounding

Audio

  • Event F1 — sound event recovery
  • Event tIoU — boundary alignment
  • Dialogue / Non-dialogue semantics

Audio-Visual

  • EVSA F1 — event attached to concurrent shot
  • Speaker F1 — dialogue assigned to right speaker
  • SGC — schema compliance gate
Data

Benchmark Construction

Three-stage OmniCapBench data construction pipeline.
OmniCapBench data construction pipeline. Stage I filters raw videos to isolate inputs with rich audio–visual complexity. Stage II generates atomic units (References → Events → Shots) through an iterative multi-model loop (generate → evaluate & validate → local refine → select), enforcing structural integrity before semantic refinement. Stage III applies traceable human audits to finalize the verified audio–visual reference system.

Statistics

StatisticAvg. / Pct.Total
Videos (< 1 min)61.32%482
Videos (1–3 min)29.01%228
Videos (3–5 min)9.67%76
Duration (hours)—12.8
Categories (Level-1)—20
Categories (Level-2)—125
References7.405,818
Shots14.5311,419
Subshots49.8239,160
Events8.326,537
— Dialogue6.835,370

Key statistics of OmniCapBench, comprising 786 videos. “Avg. / Pct.” denotes either the average per video or the percentage of total videos.

Category distribution of OmniCapBench: inner ring level-1 categories, outer ring level-2 subcategories.
Category distribution. Inner and outer rings represent level-1 categories and level-2 subcategories; segment area is proportional to video count.

Evaluation paradigm comparison

Does the protocol natively support the structural dimensions of omnimodal diagnosis? ✓ natively supported as a verifiable unit △ implicitly judged ✗ not supported

BenchmarkEvaluated UnitScoring Operator Identity
Tracking
Temporal
Grounding
Audio-Visual
Association
Diagnostic
Traceability
Whole-Caption Paradigm
AuroraCapCaptionGlobal LLM✗✗✗✗
VCapsBenchCaptionGlobal LLM✗✗✗✗
UGC-VideoCapCaptionGlobal LLM✗✗✗✗
video-SALMONN 2CaptionGlobal LLM✗✗✗✗
Probe-Based Paradigm
Omni-ClozeCloze QALLM Scorer△△△△
Parse-then-Score Paradigm
LongVALEEvent textParser + LLM✗✓✗✗
TimeChatTimed scriptParser + LLM✗✓✗△
OmniScriptScript textParser + LLM△✓△△
Native Atomic Paradigm
OmniCapBench (Ours)Atomic unitsRule + Local LLM✓✓✓✓
Results

Leaderboard

Nine omnimodal MLLMs, macro-averaged across videos. Best and second-best per column are highlighted. Evaluation is decoupled: rule-based structural compliance first, bounded semantic checks only on structurally aligned units.

Model SGC Visual Audio Audio-Visual
RefUseCCCRef Subj
F1
Ref Scene
F1
Shot
F1
Shot
tIoU
Subshot
F1
Subshot
tIoU
Event
F1
Event
tIoU
Speaker
F1
EVSA
F1
Proprietary Models
Gemini 3.1-Pro 97.3670.3934.4378.6082.68 76.2470.7465.2748.76 63.3271.9291.1851.46
Gemini 2.5-Pro 96.7370.7537.8180.7083.65 70.9372.5866.6949.18 58.5269.3589.7243.74
Qwen3.5-Omni-Plus 96.0966.9229.5573.5779.97 71.2572.3565.6948.29 53.9267.7791.2040.20
Qwen3.5-Omni-Flash 95.5863.1024.3172.2766.07 65.6568.9958.9446.52 50.6869.3489.9135.58
Seed2.0 87.6864.6931.8865.2975.13 72.6272.5668.8247.94 4.5462.7891.803.52
MiMo-2.5 96.5861.9923.9161.1969.27 66.5865.5360.3746.51 43.8261.8286.6231.69
Open-source Omnimodal Models
Qwen3-Omni-Instruct 89.0240.969.8554.8464.55 49.9867.9229.5147.95 41.7040.9988.3513.33
Qwen3-Omni-Captioner 90.7744.9611.5154.3657.08 52.2970.8039.5149.69 41.3152.6088.6817.42
MiniCPM-o-2.6 72.1321.994.2924.363.38 17.2858.459.9042.99 6.4127.4480.072.46

Rule-based structural evaluation. Metrics quantify structural compliance across Visual, Audio and Audio-Visual dimensions. SGC (schema compliance) is an independent prerequisite gate.

Model Visual Audio
Shot SubjectSceneSubshot DialogueNon-
dialogue
RecallPrec. RecallPrec. RecallPrec.
Proprietary Models
Gemini 3.1-Pro 59.0543.7867.0751.7875.62 53.2970.9291.1841.49
Gemini 2.5-Pro 55.5851.5357.0856.8065.12 53.0571.2189.7237.38
Qwen3.5-Omni-Plus 54.1745.4162.8450.7569.80 50.2270.6691.2036.82
Qwen3.5-Omni-Flash 47.6643.8765.5944.4171.26 44.0867.2789.9128.12
Seed2.0 54.7742.0563.6550.5971.42 51.1572.8891.8030.20
MiMo-2.5 48.3246.7058.6346.7265.47 44.9367.2786.6230.85
Open-source Omnimodal Models
Qwen3-Omni-Instruct 34.6340.3162.0236.0770.73 29.4257.8488.3520.51
Qwen3-Omni-Captioner 36.0842.0050.9638.0754.53 34.0146.0588.6819.65
MiniCPM-o-2.6 28.5325.9263.8720.3976.32 27.6255.8880.074.84

Localized semantic evaluation. Semantic fidelity is assessed exclusively on structurally aligned units. Bidirectional scoring isolates distinct failure modes: recall penalizes factual omissions, precision penalizes hallucinations.

Want your model on the leaderboard? Get the data on Hugging Face and the evaluation toolkit on GitHub.
Diagnosis

Key Findings

01

Identity does not survive a cut

Frontier models show strong basic perception (Gemini 2.5-Pro: 84.80% Ref Subject F1 on sub-minute videos), but tracking those identities across shots collapses — Cross-Shot Coreference Consistency reaches only 37.81% for Gemini 2.5-Pro and falls below 12% for every open-source model (MiniCPM-o-2.6: 4.29%). Long-term visual object permanence remains unsolved, and traditional metrics hide it by only checking whether an object was mentioned once.

02

Audio and vision are still two unaligned pathways

Models excel at speech transcription (dialogue semantics ≥ 80%), yet environmental-sound reasoning stays weak, and binding sounds to their concurrent visual source is harder still: Gemini 3.1-Pro falls from 63.32% Event F1 to 51.46% EVSA F1, and no open-source model exceeds 20%. Seed2.0 fails outright — it emits almost no audio-event units (4.54% Event F1, 3.52% EVSA F1) despite top-ranked speaker attribution (91.80%).

03

A precision–recall trade-off masked by fluent prose

Scoring semantics only on structurally valid units exposes divergent strategies: Gemini 3.1-Pro is conservative (67.07% precision vs. 43.78% recall on subjects), omitting details to avoid errors. MiniCPM-o-2.6 is the extreme case — it tops Scene precision (76.32%) while recalling only 20.39% of scene content, a gap a single holistic score would never reveal.

04

Schema is solved; deep structure is not

Format errors are negligible (1.3% of Gemini 3.1-Pro failures). Its remaining errors concentrate in action hallucination (28.5%), audio-visual misalignment (23.7%) and audio hallucination (17.9%), while open-source models degrade uniformly across basic perception and cross-modal binding.

The illusion of global text scores

ModelText Metric
(Prose)
Native
Constraints
Human
Audit
Pair 1: Qwen3.5-Omni-Flash vs. Qwen3-Omni
Weaker (Qwen3-Omni)52.237.521.6
Stronger (Qwen3.5-Omni-Flash)47.862.578.4
Pair 2: Qwen3.5-Omni-Plus vs. Qwen3.5-Omni-Flash
Weaker (Qwen3.5-Omni-Flash)41.721.429.9
Stronger (Qwen3.5-Omni-Plus)58.378.670.1
Average (stronger models)53.070.674.2
Average (weaker models)47.029.425.8

On two hard-to-distinguish pairs, global text metrics falsely penalize the stronger model and overestimate the weaker one, creating an illusion of parity. Native structural constraints resolve the gap and align with the human audit.

Judge sensitivity

JudgeHolisticQA probesOmniCapBench
Weak71.2461.1756.82
Mid83.0853.4958.24
Strong93.9256.5757.60
Spread (max − min)22.687.681.42

Swapping the LLM judge moves holistic scores by more than 22 points and QA probes by nearly 8, while bounded local scoring under OmniCapBench varies by only 1.4 — reliability comes from restricting the judge, not from a stronger judge.

Failure composition of evaluated models, normalized per model.
Failure composition. Each rule-based structural metric is converted into an absolute capability deficit (1 − score) and normalized to 100% of a model's failure distribution. Frontier models concentrate failures in fine-grained hallucination and audio-visual alignment; open-source models degrade broadly. Composition is shown for six representative models.

Error taxonomy

Every failure mode is tied to a rule-based metric, so a capability deficit is independently traceable rather than absorbed into one aggregate score. Counts and rates are from evaluation logs for Gemini 3.1-Pro.

Error familyPrimary metricWhat triggers itCountRate
Schema Failure GateSGCOutput cannot be parsed into references, events and shots, or misses required JSON fields313.8%
Reference Recovery VisualRef F1A required ground-truth reference is missed, or a hallucinated reference cannot be matched3,77831.7%
Reference-Use VisualRefUseA matched shot or subshot cites the wrong persistent reference ID8,64934.1%
Identity Drift VisualCCCA persistent reference is split, merged or inconsistently tracked across shots2,65665.2%
Shot Recovery / Alignment VisualShot F1 / tIoUA visual micro-action or boundary is missed, hallucinated or severely misaligned in time5,15626.2%
Event Recovery / Alignment AudioEvent F1 / tIoUA required audio event is missed, hallucinated or aligned to the wrong time segment4,21836.2%
Event-Shot Link A-VEVSAA predicted audio event is attached to the wrong visual shot in the timeline1,65444.5%
Dialogue Error A-VSpeaker F1Matched dialogue is mis-transcribed or assigned to the wrong visual speaker2607.0%

Qualitative case study

Qualitative case study of deep-structured diagnostics.
Deep-structured diagnostics in action. Holistic evaluation over-rewards fluent prose and misses errors such as hallucinated “protective eyewear” or temporally unsupported actions. By running bounded semantic checks on structurally valid units and assigning localized recall and precision, OmniCapBench decouples descriptive fidelity from structural noise and yields a high-resolution diagnostic trace.

BibTeX

@inproceedings{yang2026omnicapbench,
  title     = {OmniCapBench: A Deep-Structured Evaluation Framework for
               Fine-Grained Audio-Visual Captioning},
  author    = {Yang, Zhongyu and Tao, Jiale and Chen, Ruitao and Yang, Zuhao and
               Yuan, Yingfang and Zhao, Xueliang and Auden and Wang, Kai and
               Shao, Shuai and Wang, Biao and Yves, Steve and Lu, Qinglin},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS),
               Evaluations and Datasets Track},
  year      = {2026},
  url       = {https://openreview.net/forum?id=WgEmdr50mC},
  eprint    = {2610.12458},
  archivePrefix = {arXiv}
}