Prototype · Twenty Questions for AI models

How well can AI models play Twenty Questions?

Deep20Bench is a public benchmark for large language models (LLMs). It measures how efficiently they identify a hidden subject through adaptive yes-or-no questions.

How scoring works

The model is told nothing but the broad category (for example, "video game character") and gets up to 50 questions - more than the traditional twenty, giving models more room to finish a round. Each question, wrong guess, or reply that does not follow the required format adds one point. If the model never finds the answer, the round scores 51. Scores are averaged across all rounds, and lower is better.

What this pilot tests

The game combines several abilities.

Each answer should improve the model’s next question.

01

World knowledge

Know which categories and facts are useful.

02

Question strategy

Choose questions that remove many possibilities.

03

State tracking

Use all prior questions and answers to plan the next question.

04

Decision discipline

Make an exact guess before the limit.

How the game is played

The Guesser asks. Three roles determine the answer.

The Guesser is the LLM under test: it asks yes-or-no questions and makes the final guess. For every question, the Oracle must search the live web and cite evidence instead of relying on memory. A blind Reviewer uses that evidence to make an independent second decision on every YES or NO. If the decisions disagree, a blind Judge decides. The Guesser is isolated from this process and receives only the final YES, NO, or UNKNOWN.

Read the full game and answer-checking method →

Current pilot results

How the current runs compare.

Lower is better. A failed trial contributes 51 questions.

Lowest average score in this pilot

Claude Fable 5 (high)

High
View full run →
Question score12.06questions · lower is better95% CI 10.13–13.98

Question score

Lower is better. The blue marker is the average question score. The colored line is its 95% confidence interval (CI). The companion plot shows each exact CI width. Its three bands divide the displayed width scale into equal ranges.

Question score and confidence interval width comparison.
Question scorelower is better95% CI · color = CI widthTighterMiddleWider
CI width stabilityRows follow question-score order · lower is better
  1. 1. Claude Fable 5 (high): 12.06 questions. Lowest score in this pilot. 95% confidence interval of the average: 10.13–13.98 questions. CI width: 3.85 questions; tighter band on the displayed scale. View full run for Claude Fable 5 (high)
  2. 2. Claude Fable 5.1 (high): 12.11 questions. 95% confidence interval of the average: 8.81–15.41 questions. CI width: 6.60 questions; middle band on the displayed scale. View full run for Claude Fable 5.1 (high)
  3. 3. Claude Opus 5 (high): 12.34 questions. 95% confidence interval of the average: 11.41–13.27 questions. CI width: 1.86 questions; tighter band on the displayed scale. Smallest CI width. View full run for Claude Opus 5 (high)
  4. 4. Kimi K3 (high): 12.74 questions. 95% confidence interval of the average: 9.16–16.33 questions. CI width: 7.18 questions; middle band on the displayed scale. View full run for Kimi K3 (high)
  5. 5. gpt-oss-120B (high): 13.63 questions. 95% confidence interval of the average: 9.79–17.46 questions. CI width: 7.67 questions; middle band on the displayed scale. View full run for gpt-oss-120B (high)
  6. 6. Gemini 3.7 Flash (high): 14.03 questions. 95% confidence interval of the average: 10.78–17.28 questions. CI width: 6.49 questions; middle band on the displayed scale. View full run for Gemini 3.7 Flash (high)
  7. 7. Grok 4.6 (high): 14.29 questions. 95% confidence interval of the average: 12.36–16.21 questions. CI width: 3.85 questions; tighter band on the displayed scale. View full run for Grok 4.6 (high)
  8. 8. GPT-5.6 Sol (high): 14.40 questions. 95% confidence interval of the average: 12.84–15.96 questions. CI width: 3.11 questions; tighter band on the displayed scale. View full run for GPT-5.6 Sol (high)
  9. 9. Claude Sonnet 5 (high): 14.74 questions. 95% confidence interval of the average: 12.78–16.71 questions. CI width: 3.93 questions; tighter band on the displayed scale. View full run for Claude Sonnet 5 (high)
  10. 10. Grok 4.5 (high): 15.17 questions. 95% confidence interval of the average: 11.70–18.64 questions. CI width: 6.94 questions; middle band on the displayed scale. View full run for Grok 4.5 (high)
  11. 11. GPT-5 Nano (medium): 15.49 questions. 95% confidence interval of the average: 13.84–17.13 questions. CI width: 3.29 questions; tighter band on the displayed scale. View full run for GPT-5 Nano (medium)
  12. 12. Gemini 3.6 Flash (high): 16.89 questions. 95% confidence interval of the average: 14.08–19.69 questions. CI width: 5.61 questions; middle band on the displayed scale. View full run for Gemini 3.6 Flash (high)
  13. 13. Ox Alpha (high): 17.60 questions. 95% confidence interval of the average: 14.60–20.60 questions. CI width: 5.99 questions; middle band on the displayed scale. View full run for Ox Alpha (high)
  14. 14. GPT-5.6 Luna (high): 17.74 questions. 95% confidence interval of the average: 14.74–20.75 questions. CI width: 6.01 questions; middle band on the displayed scale. View full run for GPT-5.6 Luna (high)
  15. 15. Mistral Medium 3.5 (high): 20.06 questions. 95% confidence interval of the average: 16.11–24.01 questions. CI width: 7.90 questions; middle band on the displayed scale. View full run for Mistral Medium 3.5 (high)
  16. 16. Llama 4 Maverick (non-thinking): 32.23 questions. 95% confidence interval of the average: 26.82–37.63 questions. CI width: 10.81 questions; wider band on the displayed scale. View full run for Llama 4 Maverick (non-thinking)
Pilot result table
RankModelQuestion score95% CIReasoningSuccessContractBenchmarkrun cost
1Claude Fable 5 (high)M-0014 · anthropic
Question score12.06questions · lower is better10.13–13.98
High100%100% $7.87
2Claude Fable 5.1 (high)M-0020 · anthropic
Question score12.11questions · lower is better8.81–15.41
High97%>99% 1 violation$23.61
3Claude Opus 5 (high)M-0006 · anthropic
Question score12.34questions · lower is better11.41–13.27
High100%99% 3 violations$5.47
4Kimi K3 (high)M-0007 · moonshotai
Question score12.74questions · lower is better9.16–16.33
High94%98% 11 violations$10.09
5gpt-oss-120B (high)M-0002 · cerebras
Question score13.63questions · lower is better9.79–17.46
High94%91% 46 violations$4.64
6Gemini 3.7 Flash (high)M-0016 · google-ai-studio
Question score14.03questions · lower is better10.78–17.28
High97%100% $9.58
7Grok 4.6 (high)M-0015 · xai
Question score14.29questions · lower is better12.36–16.21
High100%100% $8.25
8GPT-5.6 Sol (high)M-0010 · openai
Question score14.40questions · lower is better12.84–15.96
High100%100% $9.97
9Claude Sonnet 5 (high)M-0005 · anthropic
Question score14.74questions · lower is better12.78–16.71
High100%97% 16 violations$5.01
10Grok 4.5 (high)M-0008 · xai
Question score15.17questions · lower is better11.70–18.64
High94%98% 10 violations$4.99
11GPT-5 Nano (medium)M-0003 · openai
Question score15.49questions · lower is better13.84–17.13
Medium94%100% $4.59
12Gemini 3.6 Flash (high)M-0004 · google-vertex
Question score16.89questions · lower is better14.08–19.69
High97%>99% 1 violation$9.86
13Ox Alpha (high)M-0017 · stealth
Question score17.60questions · lower is better14.60–20.60
High91%93% 47 violations$9.78
14GPT-5.6 Luna (high)M-0001 · openai
Question score17.74questions · lower is better14.74–20.75
High91%>99% 1 violation$5.88
15Mistral Medium 3.5 (high)M-0012 · mistral
Question score20.06questions · lower is better16.11–24.01
High89%>99% 2 violations$18.55
16Llama 4 Maverick (non-thinking)M-0009 · parasail
Question score32.23questions · lower is better26.82–37.63
None49%100% $3.73
#1Claude Fable 5 (high)anthropic
Question score
12.06
95% CI
10.13–13.98
Success
100%
Benchmark run cost
$7.87
Explore full run · questions, answers & evidence
#2Claude Fable 5.1 (high)anthropic
Question score
12.11
95% CI
8.81–15.41
Success
97%
Benchmark run cost
$23.61
Explore full run · questions, answers & evidence
#3Claude Opus 5 (high)anthropic
Question score
12.34
95% CI
11.41–13.27
Success
100%
Benchmark run cost
$5.47
Explore full run · questions, answers & evidence
#4Kimi K3 (high)moonshotai
Question score
12.74
95% CI
9.16–16.33
Success
94%
Benchmark run cost
$10.09
Explore full run · questions, answers & evidence
#5gpt-oss-120B (high)cerebras
Question score
13.63
95% CI
9.79–17.46
Success
94%
Benchmark run cost
$4.64
Explore full run · questions, answers & evidence
#6Gemini 3.7 Flash (high)google-ai-studio
Question score
14.03
95% CI
10.78–17.28
Success
97%
Benchmark run cost
$9.58
Explore full run · questions, answers & evidence
#7Grok 4.6 (high)xai
Question score
14.29
95% CI
12.36–16.21
Success
100%
Benchmark run cost
$8.25
Explore full run · questions, answers & evidence
#8GPT-5.6 Sol (high)openai
Question score
14.40
95% CI
12.84–15.96
Success
100%
Benchmark run cost
$9.97
Explore full run · questions, answers & evidence
#9Claude Sonnet 5 (high)anthropic
Question score
14.74
95% CI
12.78–16.71
Success
100%
Benchmark run cost
$5.01
Explore full run · questions, answers & evidence
#10Grok 4.5 (high)xai
Question score
15.17
95% CI
11.70–18.64
Success
94%
Benchmark run cost
$4.99
Explore full run · questions, answers & evidence
#11GPT-5 Nano (medium)openai
Question score
15.49
95% CI
13.84–17.13
Success
94%
Benchmark run cost
$4.59
Explore full run · questions, answers & evidence
#12Gemini 3.6 Flash (high)google-vertex
Question score
16.89
95% CI
14.08–19.69
Success
97%
Benchmark run cost
$9.86
Explore full run · questions, answers & evidence
#13Ox Alpha (high)stealth
Question score
17.60
95% CI
14.60–20.60
Success
91%
Benchmark run cost
$9.78
Explore full run · questions, answers & evidence
#14GPT-5.6 Luna (high)openai
Question score
17.74
95% CI
14.74–20.75
Success
91%
Benchmark run cost
$5.88
Explore full run · questions, answers & evidence
#15Mistral Medium 3.5 (high)mistral
Question score
20.06
95% CI
16.11–24.01
Success
89%
Benchmark run cost
$18.55
Explore full run · questions, answers & evidence
#16Llama 4 Maverick (non-thinking)parasail
Question score
32.23
95% CI
26.82–37.63
Success
49%
Benchmark run cost
$3.73
Explore full run · questions, answers & evidence

How to read the pilot

Comparable runs, limited conclusions.

The cohort is small. Shared conditions and public records keep it inspectable.

01

Consistent setup

The same subjects, trial count, question limit, and scoring policy apply.

02

Failures stay visible

Failures receive a declared penalty. Invalid outputs consume turns.

03

Public records

Runs link to subjects, episodes, transcripts, evidence, usage, cost, and timing.

Origin

From a holiday game to a prototype.

Patrick Heusser and Markus Tuor came up with the idea while playing Twenty Questions with the kids. Patrick then designed and built the project.