# Human Preference Study — Results (Pilot 001)

Source: Label Studio project 287119, pulled 2026-09-22.

## Executive summary

Across all 450 head-to-head tasks, using the first 5 annotations on each (2,250 votes), **gpt-6-astra** is the most-preferred model and **seed-2.0-code** the least. Preference win rate and a Bradley-Terry strength model produce the identical 1-to-10 order, so the hierarchy is robust to the choice of metric.

The ranking falls into four tiers:

1. **Leaders (~61–64% win rate)** — gpt-6-astra, muse-spark-1.3, grok-4.6. These three are bunched within ~2 points of each other; the gap between #1 and #3 is small.
2. **Upper-middle (~54–56%)** — kimi-k3, deepseek-v4-pro, glm-5.3.
3. **Coin-flip middle (~49%)** — claude-sonnet-5, qwen3.8-2.4t, both essentially even against the field.
4. **Dispreferred** — gemini-3.8-flash (33%) and seed-2.0-code (20%), clearly behind the rest.

**Caveat that matters:** this global order is an average that hides large brief-to-brief swings. The outright winner changes across 7 of the 10 briefs, and gpt-6-astra — #1 overall — wins only one brief outright and drops as low as #7 on another. Read the per-brief breakdown before treating the top-line ranking as a single verdict.

## Methodology

The project is a full balanced round-robin: 10 models compared pairwise on 10 briefs, every model pair meeting once per brief — 45 pairs × 10 briefs = 450 tasks. Each annotation is a single choice of A, B, or Tie.

1. **Select the first 5 votes per task.** For each task, annotations were ordered by creation time, cancelled and empty ones dropped, and the first 5 valid votes kept. Every task had at least 5 valid votes, so all 450 tasks and exactly 2,250 votes were used; annotations beyond the first 5 were discarded.
2. **Per-task preference.** Within each task's 5 votes, a Tie counts 0.5 to each side. The side with the higher score is the task's preferred model:

   ```
   score_A = wins_A + 0.5 · ties
   score_B = wins_B + 0.5 · ties
   ```

3. **Global hierarchy — headline metric.** Each model's preference win rate pools every vote it appears in (450 votes per model), with ties at 0.5:

   ```
   WR_i = (wins_i + 0.5 · ties_i) / games_i
   ```

4. **Corroboration — Bradley-Terry.** A latent-strength model was fit to the pairwise vote counts by MM (Zermelo) iteration, where the probability model prefers i over j is strength-based. Strengths are normalized to a geometric mean of 1, so 1.0 is an average model and higher is stronger:

   ```
   P(i ≻ j) = p_i / (p_i + p_j)
   ```

Because the schedule is balanced, win rate is already unbiased; Bradley-Terry serves as an independent check. The two rankings agree exactly.

## How the ranking algorithm works

In plain terms, the algorithm turns thousands of individual A-vs-B votes into one ordered list of 10 models. The formulas above state it precisely; this is the intuition, step by step.

1. **Collapse each task to a score.** A task shows two model outputs and collects 5 votes. Count votes for each side, splitting any Tie evenly. Example: a task where gpt-6-astra drew 4 votes, seed-2.0-code 0, and there was 1 Tie scores 4.5 to gpt-6-astra and 0.5 to seed-2.0-code — so gpt-6-astra is preferred on that task.
2. **Pool every task a model appears in.** Each model is compared against the other 9, once per brief, so it accumulates 450 votes total. Add up all the points it earned (wins plus half-ties) and divide by 450. That fraction is its win rate — the headline number. gpt-6-astra earned 286 of 450 points, i.e. 63.6%.
3. **Sort by win rate.** Highest win rate is rank 1, lowest is rank 10. Because every model faces exactly the same set of opponents the same number of times, no model gets an easier or harder schedule, so raw win rate is already a fair comparison.
4. **Confirm with Bradley-Terry.** Win rate treats every opponent as equal. Bradley-Terry instead asks: what hidden "strength" for each model best explains all the pairwise results at once? It solves for those strengths by repeatedly adjusting each model's number until the strengths reproduce the observed win counts. If a model beat mostly weak opponents, this would pull its strength down relative to win rate. Here it does not — Bradley-Terry returns the exact same 1-to-10 order, which is the evidence that the ranking is not an artifact of who-played-whom.

The two methods agreeing is the point: a simple average and a model that corrects for opponent strength land on the identical hierarchy.

## Global model hierarchy

Ranked by preference win rate across all 2,250 votes. Bradley-Terry strength (geometric mean = 1; higher is stronger) is shown alongside and produces the same order. The record is votes won / tied / lost out of the 450 votes each model received.

| Rank | Model | Win rate | Bradley-Terry strength | Record (W–T–L) |
|------|-------|----------|------------------------|----------------|
| 1 | gpt-6-astra | 63.6% | 1.72 | 280–12–158 |
| 2 | muse-spark-1.3 | 61.7% | 1.60 | 270–15–165 |
| 3 | grok-4.6 | 61.4% | 1.58 | 266–21–163 |
| 4 | kimi-k3 | 55.9% | 1.27 | 241–21–188 |
| 5 | deepseek-v4-pro | 53.8% | 1.18 | 233–18–199 |
| 6 | glm-5.3 | 53.6% | 1.17 | 232–18–200 |
| 7 | claude-sonnet-5 | 48.8% | 0.97 | 213–13–224 |
| 8 | qwen3.8-2.4t | 48.6% | 0.96 | 210–17–223 |
| 9 | gemini-3.8-flash | 33.1% | 0.52 | 142–14–294 |
| 10 | seed-2.0-code | 19.7% | 0.27 | 82–13–355 |

The strength scores make the tier gaps concrete: the three leaders sit near 1.6–1.7, the middle six cluster between 0.96 and 1.27, and the bottom two (0.52, 0.27) are well below an average model. seed-2.0-code loses roughly four of every five votes it appears in.

## Head-to-head matrix

Each cell is the row model's vote win rate against the column model (%), pooled over the 50 votes in their 10 shared tasks. A value above 50 means the row model was preferred; below 50 means it lost the matchup. Rows and columns are in global-ranking order, so the general gradient runs from strong (top-left) to weak (bottom-right).

| vs → | gpt | muse | grok | kimi | deep | glm | claude | qwen | gem | seed |
|------|-----|------|------|------|------|-----|--------|------|-----|------|
| gpt-6-astra | — | 53 | 54 | 51 | 59 | 55 | 69 | 68 | 76 | 87 |
| muse-spark-1.3 | 47 | — | 50 | 62 | 56 | 55 | 69 | 60 | 69 | 87 |
| grok-4.6 | 46 | 50 | — | 49 | 59 | 58 | 63 | 64 | 78 | 86 |
| kimi-k3 | 49 | 38 | 51 | — | 50 | 49 | 56 | 55 | 75 | 80 |
| deepseek-v4-pro | 41 | 44 | 41 | 50 | — | 58 | 53 | 50 | 69 | 78 |
| glm-5.3 | 45 | 45 | 42 | 51 | 42 | — | 58 | 54 | 69 | 76 |
| claude-sonnet-5 | 31 | 31 | 37 | 44 | 47 | 42 | — | 61 | 66 | 80 |
| qwen3.8-2.4t | 32 | 40 | 36 | 45 | 50 | 46 | 39 | — | 68 | 81 |
| gemini-3.8-flash | 24 | 31 | 22 | 25 | 31 | 31 | 34 | 32 | — | 68 |
| seed-2.0-code | 13 | 13 | 14 | 20 | 22 | 24 | 20 | 19 | 32 | — |

The order is close to transitive but not perfectly so: notably kimi-k3 beats muse-spark-1.3 62–38 despite ranking below it, and the three leaders are near 50-50 against each other. Every model beats gemini and seed decisively.

Column abbreviations: gpt = gpt-6-astra, muse = muse-spark-1.3, grok = grok-4.6, kimi = kimi-k3, deep = deepseek-v4-pro, glm = glm-5.3, claude = claude-sonnet-5, qwen = qwen3.8-2.4t, gem = gemini-3.8-flash, seed = seed-2.0-code.

## Per-brief breakdown

The global order does not hold within individual briefs. Each brief has its own round-robin (45 tasks, 5 votes each), and the winner changes across 7 of the 10 briefs. Six different models take a first place somewhere.

Winner and extremes by brief (win rate in parentheses):

| Brief | 1st | 2nd | Last (10th) |
|-------|-----|-----|-------------|
| boutique-hotel | qwen3.8-2.4t (70%) | gpt-6-astra (68%) | gemini / deepseek (13%) |
| dog-park-bar | kimi-k3 (76%) | deepseek-v4-pro (64%) | seed-2.0-code (14%) |
| independent-film | gpt-6-astra (80%) | muse-spark-1.3 (66%) | seed-2.0-code (7%) |
| italian-soda-presentation | grok-4.6 (88%) | muse-spark-1.3 (74%) | qwen3.8-2.4t (11%) |
| medical-research-conference | grok-4.6 (72%) | qwen3.8-2.4t (69%) | seed-2.0-code (0%) |
| metropolitan-transit-app | glm-5.3 (71%) | gpt-6-astra (69%) | seed-2.0-code (1%) |
| outdoor-adventure-wear | muse-spark-1.3 (72%) | deepseek-v4-pro (71%) | seed-2.0-code (0%) |
| science-publication | kimi-k3 (70%) | qwen3.8-2.4t (64%) | glm-5.3 (26%) |
| skatewear-brand | glm-5.3 (64%) | gpt-6-astra (63%) | qwen3.8-2.4t (22%) |
| telemedicine-app | muse-spark-1.3 (76%) | gpt-6-astra (72%) | seed-2.0-code (0%) |

How much each model's rank moves across briefs. Best/worst are its highest and lowest brief-level rank; a wide spread means the model is brief-dependent, a narrow one means it is consistent.

| Model | Global rank | Best brief rank | Worst brief rank | Mean brief rank |
|-------|-------------|-----------------|------------------|-----------------|
| gpt-6-astra | 1 | 1 | 7 | 3.4 |
| muse-spark-1.3 | 2 | 1 | 8 | 3.8 |
| grok-4.6 | 3 | 1 | 7 | 4.3 |
| kimi-k3 | 4 | 1 | 8 | 5.1 |
| deepseek-v4-pro | 5 | 2 | 9 | 4.6 |
| glm-5.3 | 6 | 1 | 10 | 4.7 |
| claude-sonnet-5 | 7 | 5 | 9 | 6.7 |
| qwen3.8-2.4t | 8 | 1 | 10 | 5.7 |
| gemini-3.8-flash | 9 | 3 | 10 | 7.8 |
| seed-2.0-code | 10 | 4 | 10 | 8.9 |

**Takeaways:** qwen3.8-2.4t and glm-5.3 are the most volatile — each ranges from #1 on one brief to #10 on another, so their mid-pack global position is an average of extremes, not steady performance. claude-sonnet-5 is the most consistent (always #5–#9, never top or bottom). Even the leaders swing: gpt-6-astra wins only independent-film outright and falls to #7 on dog-park-bar. seed-2.0-code is last on 6 briefs but reaches #4 on skatewear-brand — the one brief where it is genuinely competitive.

## Data notes and caveats

- **Coverage.** All 450 tasks had at least 5 valid annotations, so none were excluded and exactly 2,250 votes fed the analysis. 20 annotations (10 cancelled, 10 empty) were dropped before the first-5 selection; annotations beyond the first 5 on a task were not used, even where more existed.
- **Ties.** 81 of the 2,250 used votes (3.6%) were Ties, counted 0.5 to each side. Only 17 of 450 tasks had no clear per-task winner after the 5 votes.
- **What this measures.** These are human preferences between two rendered outputs, not an absolute quality or correctness score. A high win rate means people picked that model's output more often in side-by-side comparison.
- **Position bias looks low.** Across all used votes, A was chosen 1,966 times and B 1,954 — essentially even — so there is no strong systematic pull toward the left/right slot in aggregate. Note that A/B placement is fixed per task, so all 5 annotators on a task saw the same layout; a per-task position-bias check is possible if you want it.
- **Annotator mix.** 25 annotators contributed, with uneven volume. "First 5 by time" can lean toward whoever annotated earliest on each task; an alternative is a random 5 or all-available votes. The balanced round-robin means this is unlikely to change the top-line order, but it can be re-run if you'd prefer a different selection rule.
- **Reproducibility.** Source: Label Studio project 287119, pulled 2026-09-22. Metric order (win rate) and corroborating model (Bradley-Terry) agree exactly, which is the main evidence the hierarchy is stable.
