Note Duel Public pilot

Standings

Rankings

Every model, ranked by blind head-to-head votes, with honest error bars.

Not in the main competition: DeepSeek V4 Pro (7 of 15 playable, insufficient coverage), Kimi K3 (6 of 15 playable, insufficient coverage), Qwen3.8 Max (4 of 15 playable, insufficient coverage), Mistral Large 3 (no playable scores), Llama 4 Maverick (1 of 15 playable, insufficient coverage), Command A+ (no playable scores), MiniMax M2 (3 of 15 playable, insufficient coverage), GLM 5.3 (no playable scores), Mistral Large 4 (5 of 15 playable, insufficient coverage). They are listed under generation reliability, not ranked here.

10 earlier votes are on pairings that are no longer in the main competition. They are kept on record but not counted in these rankings.

Too early to call. 21 votes so far, and every model is still below 30 judgments, so the order below can flip with the next few votes. Each model answered every brief once; listeners hear both pieces blind through the same sampled piano. Add your vote.

1 1–3 Claude Opus 5.5 anthropic/claude-opus-5.5 · 9 briefs provisional 1627 1514–1801 70%40–89% 107W 0T 3L Rejected in 0% of duels · would hear again 10%
2 1–4 GPT-6 Astra openai/gpt-6-astra · 9 briefs provisional 1615 1475–1758 68%39–88% 116W 3T 2L Rejected in 0% of duels · would hear again 0%
3 1–5 Gemini 3.1 Pro google/gemini-3.1-pro-preview · 9 briefs provisional 1452 1275–1622 33%12–65% 93W 0T 6L Rejected in 0% of duels · would hear again 0%
4 3–4 Grok 4.7 x-ai/grok-4.7 · 9 briefs provisional 1435 1407–1455 31%10–64% 81W 3T 4L Rejected in 0% of duels · would hear again 0%
5 4–5 GPT-6 Sol openai/gpt-6-sol · 9 briefs provisional 1371 1348–1374 25%5–70% 41W 0T 3L Rejected in 0% of duels · would hear again 0%

Results by brief

The same model can shine on one brief and stumble on another, so each brief is ranked on its own.

Minuet in G major 5 votes · 10 pairings

  1. 1 Claude Opus 5.5 100% 21–100 1–0–0
  2. 2 GPT-6 Astra 50% 6–94 0–1–0
  3. 3 Grok 4.7 50% 12–88 1–1–1
  4. 4 GPT-6 Sol 50% 10–90 1–0–1
  5. 5 Gemini 3.1 Pro 33% 6–79 1–0–2

Two-voice invention in D minor 2 votes · 10 pairings

  1. 1 Claude Opus 5.5 100% 21–100 1–0–0
  2. 1 GPT-6 Astra 100% 21–100 1–0–0
  3. 3 Gemini 3.1 Pro 0% 0–79 0–0–1
  4. 3 Grok 4.7 0% 0–79 0–0–1

Nocturne in E-flat major 5 votes · 10 pairings

  1. 1 Claude Opus 5.5 67% 21–94 2–0–1
  2. 1 GPT-6 Astra 67% 21–94 2–0–1
  3. 3 Gemini 3.1 Pro 50% 10–90 1–0–1
  4. 4 GPT-6 Sol 0% 0–66 0–0–2

Toy soldier march in C major 3 votes · 10 pairings

  1. 1 Claude Opus 5.5 100% 34–100 2–0–0
  2. 2 GPT-6 Astra 25% 3–80 0–1–1
  3. 2 Grok 4.7 25% 3–80 0–1–1

Modal miniature in D Dorian, 5/4 3 votes · 6 pairings

  1. 1 Gemini 3.1 Pro 100% 21–100 1–0–0
  2. 2 GPT-6 Astra 75% 20–97 1–1–0
  3. 3 Grok 4.7 50% 6–94 0–1–0
  4. 4 Claude Opus 5.5 0% 0–66 0–0–2

Three-voice fugue in C minor 0 votes · 3 pairings

No votes in this view yet.

Sonatina exposition in A major 1 vote · 6 pairings

  1. 1 GPT-6 Astra 100% 21–100 1–0–0
  2. 2 Gemini 3.1 Pro 0% 0–79 0–0–1

Waltz in A minor 1 vote · 10 pairings

  1. 1 GPT-6 Astra 100% 21–100 1–0–0
  2. 2 Gemini 3.1 Pro 0% 0–79 0–0–1

Gigue in F major 1 vote · 10 pairings

  1. 1 Claude Opus 5.5 100% 21–100 1–0–0
  2. 2 Grok 4.7 0% 0–79 0–0–1

Impressionist prelude in F-sharp major 0 votes · 6 pairings

No votes in this view yet.

Barcarolle in G minor 0 votes · 0 pairings

No votes in this view yet.

Rag in B-flat major 0 votes · 0 pairings

No votes in this view yet.

Tango in D minor 0 votes · 0 pairings

No votes in this view yet.

Minimalist process piece in C major 0 votes · 0 pairings

No votes in this view yet.

Blues chorus in E-flat major 0 votes · 0 pairings

No votes in this view yet.

Generation reliability

How many of each model's pieces could be turned into a playable score. These counts are not musical preference scores. The ratings above are conditional on valid submissions: failed pieces are excluded from listening, not hidden or scored as losses. Coverage can be uneven across models; this is a small pilot, not a universal ranking of every system.

A model enters the blind duels and the rankings above once at least 8 of its 15 briefs have a playable score. Below that, its scores, records and earlier votes are kept, but its pairings are left out of the duels and the rankings. A provider error (a rate limit, a rejected request or an empty reply) is counted apart from invalid notation, and a failed piece is not a judgment of musical quality.

210pieces requested
71playable
45in the listening competition
139failed validation
Model Attempted Passed Failed Status
Claude Opus 5.5anthropic/claude-opus-5.5 15 9 6 6 provider errors In rankings
GPT-6 Solopenai/gpt-6-sol 15 9 6 5 provider errors · 1 invalid notation In rankings
Gemini 3.1 Progoogle/gemini-3.1-pro-preview 15 9 6 5 provider errors · 1 invalid notation In rankings
Grok 4.7x-ai/grok-4.7 15 9 6 6 provider errors In rankings
DeepSeek V4 Prodeepseek/deepseek-v4-pro-0813 15 7 8 5 provider errors · 3 invalid notation Insufficient coverage
Kimi K3moonshotai/kimi-k3 15 6 9 5 provider errors · 4 invalid notation Insufficient coverage
GPT-6 Astraopenai/gpt-6-astra 15 9 6 6 provider errors In rankings
Qwen3.8 Maxqwen/qwen3.8-max-0902 15 4 11 9 provider errors · 2 invalid notation Insufficient coverage
Mistral Large 3mistralai/mistral-large-2512 15 0 15 14 provider errors · 1 invalid notation No playable scores
Llama 4 Maverickmeta-llama/llama-4-maverick 15 1 14 5 provider errors · 9 invalid notation Insufficient coverage
Command A+cohere/command-a-plus 15 0 15 15 provider errors No playable scores
MiniMax M2minimax/minimax-m2 15 3 12 7 provider errors · 5 invalid notation Insufficient coverage
GLM 5.3z-ai/glm-5.3 15 0 15 15 provider errors No playable scores
Mistral Large 4mistralai/mistral-large-4-0 15 5 10 10 invalid notation Insufficient coverage
Technical detail

Audited partial publication: 71 of 210 model×brief slots passed fresh Verovio audit; 45 published from 5 model(s) with at least 8 of 15 briefs; 81 same-brief duel(s). Preference rankings only cover published scores; generation reliability is disclosed separately in this sidecar.

How to read this table

Judge a duel