REGISTER / VBENCH

Public Bench

vbench

  • Six model configs
  • One attempt each
  • Failures published

One build brief goes to six model-through-CLI configurations, each gets exactly one reply, and every recorded outcome inside a released round is published, whether or not it worked. It tests a narrow build workflow, not general model intelligence.

OPEN THE PUBLISHED ROUNDS

RECORD / BENCH

Configurations
6 model-through-CLI configurations, reached through two command-line clients on live provider logins.
Attempts
One request per configuration per brief. The first reply is the result. Nothing is retried, repaired or selected.
Released rounds
3 released, of which 2 carry hand scores. Curation happens at round level: an unreleased round produces no page, no route and no total.
Published runs
18 recorded outcomes across those rounds, every one of them shown.
Artifacts recovered
18 of 18. No failure is on public display today, because every released round happens to hold only recovered artifacts. The rule still binds: a failure inside a released round stays in it, at full size, attributed to its source.
What is recorded
Per run: the transport, the full argument list, the isolation string, the client version, the timeout and the output cap. Every claim here is checkable against that file.

03 ROUNDS ON THE SHELF

Published rounds

03 BRIEFS · 18 ARTIFACTS RECOVERED · GRADED ROUNDS IDENTIFIED IN PLACE

Pilot evidence predates the levelled harness. Any released pilot round is shown as-is and excluded from cross-vendor conclusions; protocol v1 is the canonical series.

R03 · MAP V1

World Map From Memory

The brief asked for a world map drawn from memory.

HAND-SCORED VERDICT Fable 5 wins decisively with the only map that genuinely looks like Earth, Sol close behind. Every toggle worked, so the round came down to geographic memory.

06 / 06 ARTIFACTS RECOVERED

VIEW FULL ROUND

R05 · SHOW V1

Fireworks Finale

The brief asked for a 60-second fireworks finale.

HAND-SCORED VERDICT Opus 4.8 and Sol 5.6 tie for the win with the only genuinely layered shows. Fable 5 shipped complete machinery with near-invisible bursts: structure without spectacle.

06 / 06 ARTIFACTS RECOVERED

VIEW FULL ROUND

THE BLIND TEST

Can you tell who built it?

One artifact from a published round. Guess the maker, then see the round and the recorded result.

UNLABELED SPECIMEN

An artifact built by one of six configurations, unlabeled

GUESS THE MAKER

ONE REPLY EACH · NO RETRIES · THE ANSWER IS RECORDED

01 / 03

Same brief, side by side

Each configuration returns one self-contained HTML file from an identical brief. The six results sit side by side on the round page and run in the browser.

The load-bearing claim is not "raw model". It is disclosed provenance. Every model here is reached through a command-line client, the client is part of the configuration under test, and the flags, caps and timeouts that shaped each reply are written down next to the reply.

02 / 03

One shot, and the failures stay

Curation happens once, at round level. A round that is not in the manifest produces no page and no number anywhere on this site. A round that is released shows every outcome it recorded. An empty or broken reply stays in the round at full size, attributed to its source, so a harness cap is never read as a model failing the brief.

That is a rule about what happens when a run fails, not a boast about having published failures. Every released round so far holds six recovered artifacts; the archive's unsuccessful runs are withheld because their whole round is, never because the run was removed.

The full protocol, both transports, their disclosed asymmetries, the rubric and the complete list of limits live on the methodology page.

03 / 03

What it does not measure

Single sample
One run per configuration per brief. No repeats, no variance estimate, no confidence interval. A rerun could reorder the results.
No iteration
One shot means no follow-up turn and no agentic loop. That loop is how these tools are actually used, and none of it is measured here.
One judge
Victor built the test and grades it, by eye, on a subjective rubric. Single evaluator, declared. No blind panel, no inter-rater check.
A configuration, not a model
Every result belongs to one model as reached through one client with one set of flags. The same model through a different client can produce a different artifact.
No global winner
A round result is a result for that brief, that day, that configuration. No ranking of model intelligence is declared anywhere on this site.
SOURCE / CURATION MANIFEST AND RUN METADATA EVERY NUMERAL DERIVED AT BUILD TIME