VGI-Bench

Xuan He1 †§, Cong Wei3,6 †, Yuhao Cheng1 †, Linrui Ma2,4 †, Yuxuan Zhang5,6,10 †
Zuojun Li2, Yuhao Wen2, Jize Jiang1, Zeyi Liu1, Yuren Hao1, Songcheng Cai3,6, Keming Wu2
Penghui Du10, Kai Zou9, Rui Yang1, Chenkai Sun8, Ke Yang1,7, Ping Nie3
Kelsey R. Allen5,6, Chenglong Wang7, Michel Galley7, Jianfeng Gao7, ChengXiang Zhai1

1UIUC, 2THU, 3UWaterloo, 4MIT, 5UBC
6Vector Institute, 7MSR, 8Independent, 9NetMind.ai, 10Etude AI

† Main Contributor    § Project Lead

EMNLP'2026 Main

Abstract

Recent studies suggest that video generative models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible.

To this end, we introduce VGI-Bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generative models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria.

Our analysis further explores the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors.

Data Construction

Comparison with Related Works

Comparison with related video benchmarks. Each pie encodes the fraction of tasks that satisfy a desideratum versus not.

Benchmark Reasoning
Demand
Realistic
Appearance
Process-
Sensitive
Difficulty
Control
PhysGenBench Low / Mid ⍻
WorldSimBench Low / Mid ✓
TiVi-Bench† High ⍻
V-ReasonBench High ✗
VBVR-Bench High ✗
VGI-Bench (Ours) High ✓
pie = fraction satisfying ✓ Yes ⍻ Partial ✗ No

† Data has not been publicly released; statistics are inferred from the paper. Realistic Appearance: photorealistic-style inputs vs. synthetic (line-art / schematic) inputs.

Construction Pipeline

1

Task Proposal

Reasoning-intensive, process-sensitive tasks expressible as video — success hinges on the trajectory, not just the final frame. 3 difficulty levels, ~10 instances each.

2

Task Material Collection (image)

Input images from web / datasets, or generated by GPT-Image-2 & Nano Banana Pro (direct or schematic → photorealistic). Human-in-the-loop, standardised to 16:9.

3

Task Material Collection (text)

Text prompt (goal, object / action rules, generation controls), a reference solution (image or target state), and per-task evaluation rubrics.

4

Quality Control

Pre-generation: probe 2 easy instances on SOTA models; keep a task only if ≥1 solves and ≥1 fails — challenging yet partly feasible. Manual review: goal fidelity, 5–10s feasibility, prompt clarity.

See the paper for full construction details.

Eval Suite

Every generated video is scored by two complementary VLM-as-judge metrics; their per-instance product is the Final Score reported on the leaderboard.

Completeness global

Captures the global progress toward the task goal. A task-specific tiered standard maps each run to one of three levels — complete / partial / failed — and the VLM-judge, given uniformly sampled frames (2 fps) plus that standard, returns the tier, mapped to {1, 0.5, 0}. complete allows minor visual imperfections; partial means meaningful progress without the correct final state; failed covers little progress or severe drift from the input.

Rubric Score local

Measures local process validity throughout the video against a fine-grained per-task checklist of explicit rules and constraints. A coarse-to-fine adaptive pass (coarse sampling at 4 fps, flagged intervals resampled at 8 fps) with a 10-frame sliding focus window catches transient violations. Each item violated x times scores 1/(x+1) (inverse-decay penalty), averaged over all items.

Final Score aggregation

Final Score = Completeness × Rubric Score

We combine the two metrics multiplicatively rather than by an arithmetic mean, because they are jointly necessary conditions — not interchangeable contributions. A near-static clip preserves most local rules (high Rubric) yet makes little progress toward the goal (low Completeness); conversely, a clip can reach the target state (high Completeness) while violating the rules along the way (low Rubric). An average would let one strong factor mask the other's failure, whereas the product requires the model to both reach the goal and respect the process — exactly when neither aspect alone should be allowed to dominate.

Image-generation outputs have no temporal axis, so they instead use a single binary Success Rate against the ground-truth image.

Leaderboard

Video Generation Models — main benchmark
ModelOverallVisual OrganizationSpatiotemporal DynamicsStructured PuzzlesPhysical Manipulation
EasyMidHardAvg.EasyMidHardAvg.EasyMidHardAvg.EasyMidHardAvg.
Commercial
Seedance 2.051.072.656.453.560.854.647.433.445.346.945.141.844.664.459.244.456.0
Kling 3.044.067.647.543.752.945.235.228.236.545.338.328.937.560.651.245.652.5
Sora 236.755.637.039.143.945.334.624.434.840.525.822.429.547.941.832.640.8
Gen-4.536.660.538.341.946.939.634.629.534.829.220.718.722.952.143.439.545.0
Wan 2.735.737.032.232.533.938.942.628.936.825.623.324.724.555.848.237.147.1
Veo 3.132.050.343.344.045.928.121.216.922.331.821.813.522.449.640.437.342.5
Open Source
MiniMax-H344.442.842.339.841.644.340.735.940.351.952.347.650.658.641.533.044.4
VBVR-Wan 2.240.246.132.434.337.637.021.628.229.057.945.050.051.046.743.137.942.6
Wan 2.221.630.415.318.521.438.732.019.930.217.18.45.610.433.722.117.524.4
HunyuanVideo-1.519.126.617.924.723.134.727.223.928.79.98.87.98.922.516.711.917.0
Image Generation Models — reference (adapted subset, measured by Success Rate %)
Model Avg. Easy Mid Hard
Commercial
Nano-Banana-Pro55.062.550.052.5
Qwen-Image-3-Pro42.751.943.832.5
GPT-Image-241.248.838.836.2
MAI-Image-2.5-Pro36.844.935.130.3
Grok-Imagine34.631.241.231.2
Seedream 4.523.827.521.222.5
Flux.2 Max Edit20.623.820.317.7
Open Source
SenseNova-U1.5-8B-MoT22.532.521.213.8
SenseNova-U116.322.214.412.2
JoyAI-Image10.411.210.010.0
Qwen-Image-Edit7.511.25.06.2
BAGEL0.81.21.20.0
Step1X-Edit0.81.20.01.2

Success Rate (%) against the ground-truth image (different metric from the video branch); best and second-best per row highlighted.

Eval Results

Pick a category, choose a task, then tick the models to compare their generated results. Video models cover all tasks; image models only cover the adapted subset (greyed-out where unavailable).

BibTeX

@misc{he2026vgibenchprobingvisualintelligence,
      title={VGI-Bench: Probing Visual Intelligence in Video Generation Models},
      author={Xuan He and Cong Wei and Yuhao Cheng and Linrui Ma and Yuxuan Zhang and Zuojun Li and Yuhao Wen and Jize Jiang and Zeyi Liu and Yuren Hao and Songcheng Cai and Keming Wu and Penghui Du and Kai Zou and Rui Yang and Chenkai Sun and Ke Yang and Ping Nie and Kelsey R Allen and Chenglong Wang and Michel Galley and Jianfeng Gao and ChengXiang Zhai},
      year={2026},
      eprint={2608.19583},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2608.19583},
}