Standings
Rankings
Every model, ranked by blind head-to-head votes, with honest error bars.
Too early to call. 26 votes so far, and every model is still below 30 judgments, so the order below can flip with the next few votes. Each model answered every brief once; listeners hear both pieces blind through the same sampled piano. Add your vote.
Results by brief
Minuet in G major 5 votes · 10 pairings
- 1 Claude Opus 5.5 100% 21–100 1–0–0
- 2 GPT-6 Astra 50% 6–94 0–1–0
- 3 Grok 4.7 50% 12–88 1–1–1
- 4 GPT-6 Sol 50% 10–90 1–0–1
- 5 Gemini 3.1 Pro 33% 6–79 1–0–2
Two-voice invention in D minor 3 votes · 10 pairings
- 1 GPT-6 Astra 100% 34–100 2–0–0
- 2 Claude Opus 5.5 50% 10–90 1–0–1
- 3 Gemini 3.1 Pro 0% 0–79 0–0–1
- 4 Grok 4.7 0% 0–79 0–0–1
Nocturne in E-flat major 6 votes · 10 pairings
- 1 Claude Opus 5.5 67% 21–94 2–0–1
- 1 GPT-6 Astra 67% 21–94 2–0–1
- 3 Gemini 3.1 Pro 50% 10–90 1–0–1
- 4 Grok 4.7 — 0–0–0
- 5 GPT-6 Sol 0% 0–66 0–0–2
Toy soldier march in C major 3 votes · 10 pairings
- 1 Claude Opus 5.5 100% 34–100 2–0–0
- 2 GPT-6 Astra 25% 3–80 0–1–1
- 2 Grok 4.7 25% 3–80 0–1–1
Modal miniature in D Dorian, 5/4 3 votes · 6 pairings
- 1 Gemini 3.1 Pro 100% 21–100 1–0–0
- 2 GPT-6 Astra 75% 20–97 1–1–0
- 3 Grok 4.7 50% 6–94 0–1–0
- 4 Claude Opus 5.5 0% 0–66 0–0–2
Three-voice fugue in C minor 0 votes · 3 pairings
Sonatina exposition in A major 2 votes · 6 pairings
- 1 Claude Opus 5.5 100% 21–100 1–0–0
- 2 GPT-6 Astra 50% 10–90 1–0–1
- 3 Gemini 3.1 Pro 0% 0–79 0–0–1
Waltz in A minor 2 votes · 10 pairings
- 1 GPT-6 Astra 75% 20–97 1–1–0
- 2 Claude Opus 5.5 50% 6–94 0–1–0
- 3 Gemini 3.1 Pro 0% 0–79 0–0–1
Gigue in F major 1 vote · 10 pairings
- 1 Claude Opus 5.5 100% 21–100 1–0–0
- 2 Grok 4.7 0% 0–79 0–0–1
Impressionist prelude in F-sharp major 1 vote · 6 pairings
- 1 Claude Opus 5.5 — 0–0–0
- 1 GPT-6 Sol — 0–0–0
Barcarolle in G minor 0 votes · 0 pairings
Rag in B-flat major 0 votes · 0 pairings
Tango in D minor 0 votes · 0 pairings
Minimalist process piece in C major 0 votes · 0 pairings
Blues chorus in E-flat major 0 votes · 0 pairings
Generation reliability
210pieces requested
71playable
45in the listening competition
139failed validation
Model
Attempted
Passed
Failed
Status
Claude Opus 5.5
15
9
6
6 provider errors
In rankings
GPT-6 Sol
15
9
6
5 provider errors · 1 invalid notation
In rankings
Gemini 3.1 Pro
15
9
6
5 provider errors · 1 invalid notation
In rankings
Grok 4.7
15
9
6
6 provider errors
In rankings
DeepSeek V4 Pro
15
7
8
5 provider errors · 3 invalid notation
Insufficient coverage
Kimi K3
15
6
9
5 provider errors · 4 invalid notation
Insufficient coverage
GPT-6 Astra
15
9
6
6 provider errors
In rankings
Qwen3.8 Max
15
4
11
9 provider errors · 2 invalid notation
Insufficient coverage
Mistral Large 3
15
0
15
14 provider errors · 1 invalid notation
No playable scores
Llama 4 Maverick
15
1
14
5 provider errors · 9 invalid notation
Insufficient coverage
Command A+
15
0
15
15 provider errors
No playable scores
MiniMax M2
15
3
12
7 provider errors · 5 invalid notation
Insufficient coverage
GLM 5.3
15
0
15
15 provider errors
No playable scores
Mistral Large 4
15
5
10
10 invalid notation
Insufficient coverage
Technical detail
How to read this table
- Rating is a Bradley–Terry strength on an Elo-like scale (average 1500). A 100-point gap means the higher piece is expected to be preferred about 64% of the time. Ties count as half a win each way.
- The bar is a 95% bootstrap interval from 400 resamples of the votes. When two bars overlap heavily, the order between those pieces is not settled; the small rank range (e.g. 1–2) says the same thing.
- Win share is (wins + ½ ties) ÷ judgments with a Wilson 95% interval.
- “Neither is good” carries no preference between the two pieces, so it is left out of the rating and shown as a rejection rate.
- Full listens (≥ 90% of both pieces heard) are the default. Partial listens are not evidence about a whole composition, so they are only included when you ask.
- Pieces with fewer than 30 judgments are marked provisional.
- Generation reliability (attempted / passed / failed slots) is separate from musical preference. Only valid scores reach the duels; failures never appear as silent losses in the standings. A model needs playable scores for at least 8 of 15 briefs to be ranked.