Solve Vision with Code

Coding Agents Are Stronger Visual Reasoners Than Video Models

A coding agent is given the first frame and the prompt of a VBVR-Pro-Bench task and writes a program that renders the answer video; the video is scored by the benchmark's official evaluator, so it sits on the same leaderboard as the video generation models.

best coding agent
 
VS
best video model
 
01

Leaderboard

VBVR-Pro-Bench, video setting: 100 tasks (50 In-Domain, 50 Out-of-Domain) × 5 instances = 500 per system. Mean official-evaluator score over all 500 instances; an instance without a video counts 0. Video-model rows are the published VBVR-Pro-Bench numbers. Click a header to sort.

coding agentcoding agent, open weightsvideo model
Rank Model Agent Overall ID OOD Videos
02

Out-of-domain vs in-domain

Each row is one system; the stem runs from its In-Domain score to its Out-of-Domain score. Every coding agent above the noise floor scores higher out of domain than in domain, without training on any family; every video model trained on the In-Domain families falls out of domain. Systems with at least one video.

coding agentopen weightsvideo model○ In-Domain  ● Out-of-Domain
In-Domain to Out-of-Domain score per system
03

Categories & efficiency

Per-category score for the coding agents (five VBVR-Pro-Bench categories) and how much work each agent did per instance: videos produced, tool-use rate, tool calls, timeouts, median seconds and tokens.

04

Per-task heatmap

100 tasks (In-Domain first, then Out-of-Domain, grouped by category) × the three closed-model agents; each cell is the mean score over the task's five instances. Hover for the task and score; click a cell to jump to that task in the gallery.

01
06

Open-weight examples

Videos rendered by programs from the open-weight roster (24 models through OpenCode on Amazon Bedrock). Each card names the model, the task and its official score.

Loading…