Bench leaderboard
This table ranks models on markup, not on whether the output works.
No run has been executed and graded yet. see verified examples →
No run has been executed and graded yet. see verified examples →
Auto-scored matrix across 18 runs · 1 model · 4 skill kinds. Scoring rubric: validity + fidelity + structure + depth + cleanliness (max 10/step).
Score over time
Dots = individual steps. Solid lines = rolling-3 average per model. As we swap models + skills the line tells the story of where we got better — or worse.
qwen-image (42)
Axis breakdown over time
Each rubric axis (max 2) plotted independently. Useful for diagnosing — e.g. fidelity climbs when subject-token recall improves, structure climbs when models start emitting proper semantic tags.
validity fidelity structure depth cleanliness
Model × skill (avg of 18 runs)
| Model | code-app | image-gen | slide-deck | website | overall |
|---|---|---|---|---|---|
| qwen-image | 7.3 /10 (3) | 8.0 /10 (24) | 10.0 /10 (2) | 8.6 /10 (13) | 8.2 /10 |
All graded runs
r-1786987902920-3k0cp 11 / 20 qwen-image r-1786964694998-hewm2 26 / 30 qwen-image r-1786964043543-v1r2j 26 / 30 qwen-image r-1786963127399-1hzlf 26 / 30 qwen-image r-1786618138520-2xqw2 9 / 10 qwen-image r-1786618060020-emxew 10 / 10 qwen-image r-1786617981960-gs9yg 10 / 10 qwen-image r-1785771742629-wo16g 18 / 20 qwen-image r-1785165929042-uxxuw 24 / 30 qwen-image r-1784907420621-cvmvv 16 / 20 qwen-image r-1784907129958-wn6uy 24 / 30 qwen-image r-1784906834360-t673h 24 / 30 qwen-image r-1784906560821-ep9yh 16 / 20 qwen-image r-1784906272609-b8sdc 24 / 30 qwen-image r-1784905978328-jmffo 24 / 30 qwen-image r-1784905702820-huqd4 24 / 30 qwen-image r-1784905252746-1otpt 24 / 30 qwen-image r-1784543685543-7tepp 10 / 10 qwen-image