AI design
under human
judgment

A STUDY OF AI DESIGN & HUMAN PREFERENCE

We gave the same ten creative briefs to ten foundation and open models. Explore their decisions, inspect the work, and see what people chose.

Explore the results
10

Models

10

Shared creative briefs

100

Generated designs

4,500

Human preference votes

Science Publication

The designs are ordered by how often people chose each model’s work here, 90 votes per model.

Outputs for Science Publication (ranked by human preference)

10 DESIGNS · 90 VOTES PER MODEL Download data

Rank and win rate are each model’s result on this brief: 90 votes per model, ties split equally, tied models sharing a position. They are keyed by model and brief, not by this generation’s ID.

02 / The decisions · brief 01 of 10

Where the directions converge

In a separate run without tools, each model set out a visual direction for this brief. Typefaces on the left, every palette color on the right. Select a model’s mark or a color to read that direction in the model’s own words.

TYPEFACES NAMEDLoading the stated directions…
PALETTE COLORSEvery color in the ten palettes, placed by hue around the ring and by saturation from the center. Larger circles are each model’s accent color. Near-neutral grounds and text colors are not drawn.
Loading palettes…

Model-authored directions from the tool-free arm. Family names are taken as written; descriptions without a name are counted together. Specimens marked “approx.” are licensed faces shown in a stand-in. We did not collect private chain of thought, and the direction was not necessarily supplied to the build run.

03 / Overall rankings

Overall rankings

This table pools every vote across the ten briefs above. The top three models are separated by 4.0 percentage points, and seven different models take first place on at least one brief. The brief range column shows how far each model moves.

Global · 900 votes per modelPooled across all ten briefs. Bradley–Terry is fit on this schedule.
Overall human preference ranking from 4,500 votes, 900 per model
RankModelPreference win rateBT strengthBrief rangePilot
01GPT 6 Astra
64.3%
1.781–801
02Grok 4.6
62.7%
1.671–803
03Muse Spark 1.3
60.3%
1.521–702
04DeepSeek V4 Pro
56.2%
1.291–905
05Kimi K3
53.3%
1.161–804
06GLM 5.3
53.2%
1.151–1006
07Claude Sonnet 5
49.1%
0.984–907
08Qwen 3.8 2.4T
48.9%
0.981–1008
09Gemini 3.8 Flash
32.8%
0.514–1009
10Seed 2.0 Code
19.1%
0.267–1010

HumanSignal / Research & evaluation

Put human judgment
into your next evaluation.

Studying model behavior or comparing creative outputs? Talk to us about designing the rubric and collecting the human feedback your research needs.

Talk to our team