Benchmark · Leaderboard

VISTA Leaderboard

Visual Spec-To-App Benchmark — how well do coding agents build real web apps from a design?

10 app categories · 128 annotated pages · 458 visual anchor points · built from Figma designs

Leaderboard

Combined score — condition C4 · latest harness
#ModelHarnessEffortCombined scoreTokens / taskCost / task
1Grok 4.6cursor-agent 2026.08.11High
0.552 ± 0.022438K$2.38
2GPT-5.6-solCAMEL 0.2.90 · SingleHigh
0.538 ± 0.01690K$1.77
3fable-5claude code 2.1.152High
0.533757K$12.04
4Grok 4.5cursor-agent 2026.07.01Med
0.517161K$1.83
5GPT-5.6-solcodex 0.144 · highHigh
0.507354K$4.49
6Opus 4.7claude code 2.1.152High
0.504
7Opus 5claude code 2.1.219High
0.490§
8Opus 4.8claude code 2.1.215High
0.484
9DeepSeek-V4-Flashdsh 0.1.0-rc.6
0.484
10DeepSeek-V4-Prodsh 0.1.0-rc.6
0.462
11Composer 2.5cursor-agent 2026.06.15
0.448120K
12Sonnet 4.6claude code 2.1.152High
0.446§252K$3.85
13Kimi K3kimi-code 0.26.0
0.440
14GLM-5.2claude code 2.1.152 · z.aiMax
0.438304K$3.66
15GPT-5.5codex 0.144.6 · highHigh
0.413
16Muse-Spark-1.2-contributorclaude code 2.1.205 · api.meta.ai
0.409185K$0.03
17GPT-5.6-terracodex 0.144.6 · medMed
0.385
18Muse-Spark-1.1claude code 2.1.205 · api.meta.ai
0.384§3.8M$5.16
19GPT-5.4-minicodex 0.134Med
0.358611K$1.68
20GPT-5.6-lunaCAMEL · WorkforceMed
0.357
21MAI-Code-1-Flashcopilot 1.0.68
0.3267.4M

Agents are ranked by the Combined score S — DOM-grounded localization × behavior on human-annotated UI anchors, scored under the content-based interaction criterion (a click counts only if it produces the corresponding observable content change), averaged over 10 apps (failures count as 0). Each model runs on its latest harness release with a complete 10-app batch (Codex-CLI 0.134–0.144.6 / Claude Code up to 2.1.215 / Cursor / Copilot / CAMEL), free choice of stack. Where repeated complete batches are available, ± SD is the sample standard deviation of the independent 10-app batch means; Grok 4.6 + Cursor and GPT-5.6-sol + CAMEL each use three complete batches (n=3, 30 valid task runs). §Rows marked § are averaged over fewer than 10 apps (Opus 5 over 9/10; Sonnet 4.6 and Muse-Spark over 8/10) — a few apps' containers no longer build (environment drift) and are excluded rather than scored 0. Rows marked † are a single 10-app batch rather than a 3-round mean (Kimi K3 — its per-cycle quota is too small to complete more rounds). ConditionC4 gives the agent the richest spec: the page's rendered Figma image (a screenshot mockup) and its pruned Figma structure (the layout tree as JSON) — but no target framework.

Reasoning effort used

Each harness exposes a different reasoning-effort ladder, so levels aren't directly comparable. The highlighted dot marks the level each model ran at; the row length shows the granularity available. Notably, GLM-5.2 ran at its ceiling (max), while the Claude and Codex models still had headroom above the level used.

High · 3 / 6 ClaudeClaude Code (6 levels) — fable-5, Opus 4.8, Opus 4.7, Sonnet 4.6, Haiku 4.5
High · 3 / 6 OpenAICodex GPT-5.6 (sol & terra: 6 levels — Low·Medium·High·Extra-high·Max·Ultra, where Ultra = Max + auto task delegation; luna: 5, no Ultra) — sol ran High (3/6); terra & luna ran Medium (2/6, 2/5)
High · 3 / 4 OpenAICodex GPT-5.5 (4 levels: Low·Medium·High·Extra-high — no Max/Ultra) — GPT-5.5 (high run)
Medium · 2 / 4 OpenAICodex (4 levels) — GPT-5.5, GPT-5.4, GPT-5.4-mini
Max · 2 / 2 Zhipu GLMz.ai GLM-5.2 (2 levels: high · max) — ran at max
Medium · 2 / 3 Cursorcursor-agent (Grok 4.5: medium · high · xhigh) — Grok 4.5 (ran at high = "Medium")
High Cursorcursor-agent 2026.08.11 — Grok 4.6 (High; 30 valid task runs)
Medium · 2 / 3 AntigravityAntigravity (3 levels: low · medium · high) — Gemini 3.5 Flash

Cursor Composer 2.5 runs at its agent default; its effort ladder isn't published in a directly comparable form.

Cost & speed

Tokens & wall-clock per task

Median over the available task runs — normally one 10-app batch; Grok 4.6 + Cursor and CAMEL 0.2.90 each use all three complete batches (30 runs). The median is outlier-robust, so a single runaway or oversized task doesn't skew a model's value. Bars are colored by token type.

Tokens per task
cache input   input (non-cached)   output
Grok 4.5Cursor 2026.07.01 2.6M GPT-5.6-solCodex 0.144 4.2M fable-5Claude 2.1.152 11.9M GLM-5.2Claude 2.1.152 11.5M Opus 4.8Claude 2.1.152 9.2M GPT-5.5 highCodex 0.134 3.5M Sonnet 4.6Claude 2.1.152 6.7M Opus 4.7Claude 2.1.152 7.4M Muse-Spark-1.1Claude 2.1.205 4.3M Composer 2.5Cursor 2026.06.15 1.9M GPT-5.4-miniCodex 0.134 12.6M Haiku 4.5Claude 2.1.152 6.4M GPT-5.6-solCAMEL 0.2.90 1.45M Grok 4.6Cursor 2026.08.11 4.24M
Median tokens per task, stacked by type: cache-read input, non-cached input, and output. Harness and version are shown beneath each model. Grok 4.6 uses 4.24M total tokens/task (3.786M cache, 358K non-cached input, 77.7K output); CAMEL uses 1.45M total, about 95.5% cache reads.
Time per task
model / harness  ·  agent run only
Grok 4.5Cursor 2026.07.01 9.2m GPT-5.6-solCodex 0.144 18.0m fable-5Claude 2.1.152 43.3m GLM-5.2Claude 2.1.152 29.4m Opus 4.8Claude 2.1.152 26.1m GPT-5.5 highCodex 0.134 14.8m Sonnet 4.6Claude 2.1.152 20.7m Opus 4.7Claude 2.1.152 19.3m Muse-Spark-1.1Claude 2.1.205 10.8m Composer 2.5Cursor 2026.06.15 10.2m GPT-5.4-miniCodex 0.134 23.0m Haiku 4.5Claude 2.1.152 11.6m GPT-5.6-solCAMEL 0.2.90 7.5m Grok 4.6Cursor 2026.08.11 20.3m
Median wall-clock to build each app (agent run only, excludes eval). Harness is shown beneath each model; Grok 4.6 and CAMEL each use the median over all 30 valid task runs.

Trajectory audit: token and time values were recomputed from each run's terminal usage/result event. Unless noted, each row uses its named complete 10-task C4 trajectory group. Grok 4.6 and CAMEL each use 30 valid runs (three batches). Grok 4.6's archived zero-score attempt is excluded and replaced by the valid 0.505 rerun. Composer 2.5 has only 8 parseable terminal result events, so its 1.9M / 10.2m medians are n=8. The chart's Opus 4.8, Opus 4.7 and GPT-5.5 data come from the locally archived Claude Code 2.1.152, Claude Code 2.1.152 and Codex 0.134 trajectories respectively; the harness sublabels make this explicit even where the leaderboard above has a newer scored harness.

How tokens are counted: leaderboard Tokens uses non-cached input + output so cache re-reads do not make multi-turn harnesses look artificially heavy. The stacked chart shows cached input, non-cached input, and output. Grok 4.6's 30-run medians are 3.786M cached input, 358.2K non-cached input, 77.7K output, and 438.4K non-cached input + output. For CAMEL 0.2.90, the corresponding medians are 1.356M, 66.8K, 21.9K, and 90.2K. Claude/Cursor use input + cache-creation + output; Codex/CAMEL use (input − cached) + output. Codex's output_tokens already includes its reasoning_output_tokens; the latter must not be added again. GLM-5.2 totals come from each run's terminal result, since z.ai's endpoint leaves per-message usage empty.

How cost is estimated: median API cost per task = each run's captured billed tokens × provider pricing, then the median across task runs. GPT-5.6-sol uses OpenAI's $5 / $0.50 / $30; GPT-5.4-mini uses $0.75 / $0.075 / $4.50; GLM-5.2 uses z.ai's $1.40 / $0.26 / $4.40 input/cached/output rates. Recomputed medians are Sol/Codex $4.49, GPT-5.4-mini $1.68, GLM-5.2 $3.66, and Muse-Spark $5.16 using the captured Meta rate card. Claude rows use the CLI's own total_cost_usd (fable-5 $12.04; Sonnet 4.6 $3.85). GPT-5.6-sol + CAMEL is an estimate of $1.77/task / $57.01 for 30 runs, including the documented 1.25× cache-write premium; without it, $1.68/task. Grok 4.6 is a theoretical estimate of $2.38/task / $71.04 for 30 runs using captured tokens and Grok 4.5 public rates as a proxy because Cursor reports no USD total; it is not an observed bill. Grok 4.5's $1.83 is from Cursor's dashboard total and cannot be independently reconstructed from trajectory pricing.

Co-evolution

Harness × LLM — co-evolving over releases

An agent is a model inside a harness, and both ship on their own cadence. Tracking the same model across harness releases shows the two co-evolving — sometimes lifting each other, sometimes regressing. Combined score by harness release.

0.300.200.100.00 0.1160.1280.134 Mar 19Apr 30May 26 Codex-CLI version (release date) → GPT-5.4 GPT-5.5 GPT-5.4-mini
Codex-CLI — all three GPT models dip from the Apr-30 (0.128) to May-26 (0.134) build; GPT-5.4 peaks at 0.128.
0.300.200.100.00 2.1.582.1.1262.1.152 Feb 25Apr 30May 26 Claude Code version (release date) → Sonnet 4.6 Haiku 4.5 Opus 4.7 Opus 4.8 fable-5
Claude Code — Sonnet 4.6 rises monotonically across releases; fable-5 (2.1.152) tops the field; Haiku 4.5 trails. Opus 4.7/4.8 shown at 2.1.126.

Mean C4 combined score (n=10, failures scored 0). Codex-CLI and Claude Code use independent version schemes, shown as separate panels.

How it's scored

DOM-grounded, behavior-aware

Each task asks an agent to build and launch a multi-page web app from a visual spec. We bring the app up in Docker and, for every human-annotated UI anchor, match it to a rendered DOM element — scoring localization L (IoU / distance to the mockup position) and behavior B (interaction-specific browser checks). The headline Combined score is S = mean(L · B) over the critical anchors.

Visual spec Agent Live app Anchor → DOM Combined S mockup · Figma · anchors model × harness docker compose up localize L × behavior B mean(L · B)
Pipeline: a visual spec drives the agent (model × harness) to build a runnable app; each annotated UI anchor is matched to a DOM element and scored on localization and behavior, combined into S.