AI design
under human
judgment
We gave the same ten creative briefs to ten foundation and open models. Explore their decisions, inspect the work, and see what people chose.
Explore the resultsModels
Shared creative briefs
Generated designs
Human preference votes
Science Publication
The designs are ordered by how often people chose each model’s work here, 90 votes per model.
Outputs for Science Publication (ranked by human preference)
Rank and win rate are each model’s result on this brief: 90 votes per model, ties split equally, tied models sharing a position. They are keyed by model and brief, not by this generation’s ID.
02 / The decisions · brief 01 of 10
Where the directions converge
In a separate run without tools, each model set out a visual direction for this brief. Typefaces on the left, every palette color on the right. Select a model’s mark or a color to read that direction in the model’s own words.
Model-authored directions from the tool-free arm. Family names are taken as written; descriptions without a name are counted together. Specimens marked “approx.” are licensed faces shown in a stand-in. We did not collect private chain of thought, and the direction was not necessarily supplied to the build run.
03 / Overall rankings
Overall rankings
This table pools every vote across the ten briefs above. The top three models are separated by 4.0 percentage points, and seven different models take first place on at least one brief. The brief range column shows how far each model moves.
| Rank | Model | Preference win rate | BT strength | Brief range | Pilot |
|---|---|---|---|---|---|
| 01 | GPT 6 Astra | 64.3% | 1.78 | 1–8 | 01 |
| 02 | Grok 4.6 | 62.7% | 1.67 | 1–8 | 03 |
| 03 | Muse Spark 1.3 | 60.3% | 1.52 | 1–7 | 02 |
| 04 | DeepSeek V4 Pro | 56.2% | 1.29 | 1–9 | 05 |
| 05 | Kimi K3 | 53.3% | 1.16 | 1–8 | 04 |
| 06 | GLM 5.3 | 53.2% | 1.15 | 1–10 | 06 |
| 07 | Claude Sonnet 5 | 49.1% | 0.98 | 4–9 | 07 |
| 08 | Qwen 3.8 2.4T | 48.9% | 0.98 | 1–10 | 08 |
| 09 | Gemini 3.8 Flash | 32.8% | 0.51 | 4–10 | 09 |
| 10 | Seed 2.0 Code | 19.1% | 0.26 | 7–10 | 10 |
04 / Shared patterns
Different models.
Familiar decisions.
Look for convergence in the models’ stated directions. Shared typefaces, similar palettes, and recurring compositions help describe the design space without turning difference into a quality score.
Distance from the other decisions.
A higher index means a model’s stated font and primary color choices differ more from those of its peers on the same brief.
We average two distances equally: named-font set difference (Jaccard) and normalized RGB distance for the primary color. We then average comparisons within the same brief and multiply by 100. Pairs without recognized fonts are excluded. This index does not measure originality, layout, or human preference.
From the research
More thoughts and observations
Is taste a moat?
How human judgment can become a learning system, and why consensus alone can hide useful differences.
Read the essay notes NIKOLAI LIUBIMOV / METHOD NOTESFrom guidelines to good annotations.
Make the observations behind a judgment explicit. Preserve disagreement when preference is what you want to measure.
Read the method notesHumanSignal / Research & evaluation
Put human judgment
into your next evaluation.
Studying model behavior or comparing creative outputs? Talk to us about designing the rubric and collecting the human feedback your research needs.