BVB: Benchmarking Agentic Video Understanding
via Programmatic Reconstruction in Blender

1 University of Rochester2 Sony Group Corporation3 Carnegie Mellon University4 University of Washington

“If an agent truly understands a video, it can reconstruct it programmatically.”

Real videoARKitScenes · 42446049
Blender reconstructionGPT-6 Astra · high

Matched relative video times. Each full clip is shown over a 30-second preview; scores are for this scene. Examples are drawn from the paper.

The paper

Abstract

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introduce BVB, Blender-VideoBench, a benchmark that tests this ability by asking agents to reconstruct real-world videos as animated Blender scenes. To ensure fair comparison, each agent programs the reconstruction through a lightweight harness, Mini-BVB, in an identical sandbox under a shared cost limit. The benchmark renders each reconstruction from its animated camera and evaluates it on two axes: (1) Dual VQA measures how many spatiotemporal facts the reconstruction preserves. (2) Latent Similarity measures how closely the reconstruction matches the source video perceptually. Our overall score, a square-root mean, favors balanced performance. We evaluate 51 configurations from 10 model families and analyze semantic retention, perceptual similarity, reasoning effort, and cost. The best model reaches 88.6 Latent Similarity but retains only 53.7% of the source-correct spatiotemporal answers. Additional reasoning improves visual similarity but does not close this gap in factual accuracy. In a blind study with 15 raters and five configurations, Latent Similarity correlates strongly with human preference. These results show that programmatic reconstruction is a viable test of agentic video understanding, and that semantic retention remains the main challenge.

01 / Full results

51 configurations.
Two ways to understand.

We evaluate frontier multimodal models as zero-shot reconstruction agents, without fine-tuning. Both axes are reported beside Overall.

51 configurations · ranked by OverallDownload results
All 51 BVB configurations. Scores on a 0–100 scale. Overall is the square-root mean of Dual VQA and Latent Similarity. Cost and runtime are per-scene means.
Rank Model & Reasoning Effort Harness
1GPT-6 Astrahighmini-BVB70.0753.788.687.390.01.25814.00
2GPT-5.6 Solxhighmini-BVB67.4951.785.484.086.70.7784.96
3Grok-4.6xhighmini-BVB67.1752.184.282.985.51.68235.51
4Qwen3.8-Maxhighmini-BVB66.2452.781.379.982.80.59912.89
5Claude Opus 5highmini-BVB66.2152.681.479.882.92.1577.97
6GPT-5.6 Solhighmini-BVB65.7349.883.982.585.30.4192.21
7GPT-5.6 Terraxhighmini-BVB64.7948.982.981.684.30.2354.76
8GPT-5.6 Terrahighmini-BVB64.7649.282.481.083.90.1223.37
9GPT-5.5highmini-BVB64.5349.381.880.483.20.9566.77
10GPT-5.6 Solmini-BVB64.0949.580.679.182.00.2121.20
11GPT-5.6 Solmediummini-BVB64.0148.581.680.183.10.2711.60
12GLM-5.3-Flashxhighopen-weightmini-BVB63.9650.878.777.080.30.0246.41
13GPT-5.6 Lunahighmini-BVB63.5848.281.179.682.70.0801.02
14Muse Spark 1.3xhighmini-BVB63.4048.979.778.381.20.2296.65
15GPT-5.6 Sollowmini-BVB63.2349.079.377.880.80.2001.04
16Grok-4.5highmini-BVB62.8148.479.077.680.40.3002.16
17GPT-5.5lowmini-BVB62.8049.677.576.079.00.3201.71
18GPT-5.5mediummini-BVB62.5646.980.579.181.90.5624.14
19GPT-5.6 Terramediummini-BVB62.4348.278.577.080.10.1498.42
20Claude Sonnet 4.6mini-BVB62.3750.675.373.876.90.4322.92
21Gemini 3.1 Prohighmini-BVB62.2351.174.572.976.10.6676.18
22Gemini 3.8 Flashhighmini-BVB62.0848.477.476.078.80.62810.63
23GPT-5.6 Lunamediummini-BVB61.6246.578.877.280.50.0470.57
24Claude Sonnet 5mini-BVB61.2449.274.672.976.30.6443.15
25GPT-5.2highmini-BVB61.0847.975.874.477.20.5787.45
26GPT-5.6 Terralowmini-BVB61.0747.176.875.278.40.1465.82
27Claude Sonnet 5highmini-BVB60.6348.374.472.776.10.4213.08
28Claude Sonnet 4.6highmini-BVB60.5347.774.973.476.50.74722.86
29GPT-5.6 Lunalowmini-BVB60.3346.875.673.977.30.0270.35
30Claude Opus 4.7mini-BVB60.3249.272.670.974.30.3691.53
31Claude Opus 4.6mini-BVB60.2347.574.573.075.90.6722.33
32Muse Spark 1.2xhighmini-BVB59.9046.275.474.176.70.2356.65
33Claude Opus 4.7highmini-BVB59.1748.371.169.372.90.3511.52
34Claude Opus 4.8mini-BVB59.0747.272.270.474.10.3251.06
35Gemini 3 Flashhighmini-BVB58.9149.669.067.370.70.0831.96
36Claude Opus 4.8highmini-BVB58.4646.471.969.973.90.4601.52
37GPT-5.2mini-BVB57.9546.770.468.772.20.1201.25
38GPT-5.5mini-BVB57.0342.673.671.875.40.1630.75
39GPT-5.4mini-BVB56.8046.668.066.369.60.2031.07
40GPT-5.4 Minimini-BVB56.5146.567.565.769.40.0340.53
41GLM-5V-Turbomini-BVB56.4746.867.065.468.70.2766.16
42Kimi-K2.5open-weightmini-BVB55.9144.069.267.770.80.1506.24
43MiniMax-M3open-weightmini-BVB55.2945.665.964.367.60.0583.06
44Grok-4.3highmini-BVB54.0844.764.362.865.90.1011.69
45Qwen3.5-122Bopen-weightmini-BVB53.5044.163.862.065.60.2255.16
46Qwen3.5-397Bopen-weightmini-BVB52.7543.063.561.965.20.17913.99
47Seed-2.0-Minimini-BVB52.2443.461.959.864.00.0284.25
48Qwen3.6-Plusmini-BVB51.2841.462.260.863.70.17011.73
49Grok-4.3mini-BVB50.5944.457.255.459.00.0470.54
50Seed-2.0-Litemini-BVB49.4741.358.456.260.60.0391.73
51Qwen3.5-27Bopen-weightmini-BVB47.5038.657.355.858.90.0736.08

All 288 test scenes are included. Failed reconstructions score zero on both axes. Rank follows the selected metric across the full pool; filtering does not renumber it. Tied displayed scores are ordered using full precision.

Task-level Dual VQA breakdown 8 spatiotemporal tasks

Retention is conditioned on the questions the judge answers correctly on the source. Relative direction combines the easy, medium, and hard subsets.

Dual VQA task scores for each configuration in the current filters and ranking.
Model & Reasoning Effort Harness Object
count
Absolute
distance
Object
size
Room
size
Relative
distance
Relative
direction
Route
planning
Appearance
order
GPT-6 Astrahighmini-BVB48.533.573.733.149.050.054.836.0
GPT-5.6 Solxhighmini-BVB38.039.969.053.143.650.747.316.0
Grok-4.6xhighmini-BVB40.542.266.246.942.752.559.124.0
Qwen3.8-Maxhighmini-BVB42.537.665.845.438.656.172.030.0
Claude Opus 5highmini-BVB42.046.266.443.844.052.065.616.0
GPT-5.6 Solhighmini-BVB34.537.665.040.837.852.263.426.0
GPT-5.6 Terraxhighmini-BVB33.030.665.843.140.250.762.412.0
GPT-5.6 Terrahighmini-BVB37.534.166.245.430.352.563.416.0
GPT-5.5highmini-BVB34.534.765.846.237.850.563.412.0
GPT-5.6 Solmini-BVB31.037.067.137.740.250.761.322.0
GPT-5.6 Solmediummini-BVB35.040.564.333.841.148.061.318.0
GLM-5.3-Flashxhighopen-weightmini-BVB35.038.264.546.246.951.754.828.0
GPT-5.6 Lunahighmini-BVB28.536.465.234.637.354.454.810.0
Muse Spark 1.3xhighmini-BVB35.036.464.135.440.251.262.420.0
GPT-5.6 Sollowmini-BVB28.541.063.035.439.454.462.422.0
Grok-4.5highmini-BVB35.039.364.142.339.847.359.114.0
GPT-5.5lowmini-BVB35.033.561.842.344.854.455.926.0
GPT-5.5mediummini-BVB25.535.863.941.537.850.051.612.0
GPT-5.6 Terramediummini-BVB36.034.162.439.243.247.355.934.0
Claude Sonnet 4.6mini-BVB30.537.063.943.843.257.462.414.0
Gemini 3.1 Prohighmini-BVB30.044.565.243.141.155.163.420.0
Gemini 3.8 Flashhighmini-BVB41.535.865.437.742.744.951.618.0
GPT-5.6 Lunamediummini-BVB23.031.862.235.441.952.553.814.0
Claude Sonnet 5mini-BVB30.039.962.436.240.754.964.518.0
GPT-5.2highmini-BVB39.038.260.946.232.451.259.112.0
GPT-5.6 Terralowmini-BVB27.036.462.438.540.249.057.024.0
Claude Sonnet 5highmini-BVB29.037.065.428.539.854.754.810.0
Claude Sonnet 4.6highmini-BVB35.532.462.834.636.952.958.112.0
GPT-5.6 Lunalowmini-BVB20.535.860.534.643.251.267.718.0
Claude Opus 4.7mini-BVB36.041.061.534.642.354.957.08.0
Claude Opus 4.6mini-BVB28.039.363.730.046.148.050.524.0
Muse Spark 1.2xhighmini-BVB33.532.462.836.934.047.562.410.0
Claude Opus 4.7highmini-BVB33.542.261.535.438.253.755.914.0
Claude Opus 4.8mini-BVB30.537.062.035.441.948.857.018.0
Gemini 3 Flashhighmini-BVB31.538.766.440.839.052.558.118.0
Claude Opus 4.8highmini-BVB29.033.560.930.041.151.754.816.0
GPT-5.2mini-BVB30.041.060.931.540.250.750.512.0
GPT-5.5mini-BVB21.524.359.830.834.946.359.114.0
GPT-5.4mini-BVB23.541.664.843.132.848.550.516.0
GPT-5.4 Minimini-BVB19.545.160.923.840.752.263.414.0
GLM-5V-Turbomini-BVB25.542.261.536.932.852.959.112.0
Kimi-K2.5open-weightmini-BVB24.032.959.233.836.947.558.16.0
MiniMax-M3open-weightmini-BVB20.035.861.530.840.249.562.414.0
Grok-4.3highmini-BVB22.036.457.532.341.151.052.712.0
Qwen3.5-122Bopen-weightmini-BVB16.532.963.925.429.052.957.08.0
Qwen3.5-397Bopen-weightmini-BVB16.032.960.929.238.245.157.010.0
Seed-2.0-Minimini-BVB17.031.861.520.831.152.961.34.0
Qwen3.6-Plusmini-BVB20.531.854.130.029.948.561.312.0
Grok-4.3mini-BVB13.029.559.030.040.252.568.812.0
Seed-2.0-Litemini-BVB10.533.557.930.032.048.552.78.0
Qwen3.5-27Bopen-weightmini-BVB13.527.254.526.230.745.349.54.0

02 / The benchmark

From video to an
editable world.

BVB asks multimodal agents to reconstruct real-world videos as animated Blender scenes. Every model faces the same sandbox, prompt, and cost ceiling.

288real indoor videos
5,130spatiotemporal questions
51agent configurations
10model families
01 · Observe

Real indoor video

Video frames alone. No depth maps, segmentation, or 3D ground truth.

ARKitScenes · ScanNet · ScanNet++
02 · Reconstruct

Mini-BVB Harness

The agent inspects frames, writes Python, and runs Blender in a fresh sandbox.

frames + bash · $3 per scene
03 · Render

Animated Blender scene

Geometry built from primitives, with an animated camera along the source trajectory.

result.blend · Blender 4.2
Controlled comparison

External asset libraries are disallowed. Stage 2 scores the file and never re-enters the agent loop. If an agent fails to produce a renderable file, that scene scores zero on both axes.

DVSemantic retention

Dual VQA

Measures how many spatiotemporal facts the reconstruction preserves.

DV = |Cs ∩ Cr||Cs|

A VLM judge answers the same questions on source and rendered videos. Retention is scored only on questions answered correctly on the source.

Judge: GPT-5.4 mini · 16 frames per video
LSPerceptual similarity

Latent Similarity

Measures how closely the reconstruction matches the source video perceptually.

LS = Layout + Motion2

A frozen V-JEPA encoder compares spatial arrangement and temporal dynamics separately, then averages them.

V-JEPA 2.1 ViT-G · 64 frames per video
Two complementary axes

Strong performance on both.

We combine both axes into a square-root mean. More uneven performance receives a larger penalty.

Overall = (DV + √LS2)2Both axes on a 0–100 scale

03 / What current agents reveal

Looking right is not
the same as being right.

The best model reaches 88.6 Latent Similarity but retains only 53.7% of the source-correct answers, so nearly half of verifiable semantic content is still lost after reconstruction.

Best configuration · GPT-6 Astra high

Visually plausible.
Semantically incomplete.

Current agents therefore build reconstructions that look right but get many facts wrong.

Dual VQA53.7%
Latent Similarity88.6
Overall70.07Rank 1 of 51

What does the frontier cost?

Cost and quality do not scale together.

The paper's cost frontier across 51 configurations, showing Overall score versus mean per-scene cost.

Mean Stage-1 spend per scene · logarithmic cost axis · higher score is better.
Dashed line: best score available at each cost. Select a point to inspect a configuration.

Semantic retention

Whole-scene facts are harder.

Object size and route planning have the highest mean retention, while appearance order and object count are lowest.

Object size63.0%
Route planning58.5%
Relative direction51.0%
Relative distance39.0%
Room size36.5%
Absolute distance36.5%
Object count29.3%
Appearance order16.1%

Mean retention across all 51 configurations.

Full task profiles from the paper

Use BVB

Build on the benchmark.

BVB provides a reproducible evaluation for agentic video understanding, built on real videos, shared budgets, and released editable artifacts.

Citation
@misc{tang2026bvb,
  title = {BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender},
  author = {Yolo Y. Tang and
    Daiki Shimada and
    Jiayue Meng and
    Jing Bi and
    Pinxin Liu and
    Yicheng Wang and
    Yunzhong Xiao and
    Zhangyun Tan and
    Zeliang Zhang and
    Chao Huang and
    Susan Liang and
    Qianxiang Shen and
    Luchuan Song and
    Ali Vosoughi and
    Mingqian Feng and
    Melika Filvantorkaman and
    Chenliang Xu},
  year = {2026},
  eprint = {2609.15478},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url = {https://arxiv.org/abs/2609.15478}
}