3 runs per problem · Errors included · Problem & response details available
- Accuracy counts correct answers across all three attempts. Pass@3 counts problems with at least one correct answer. Errors and timeouts count as incorrect.
- Token averages use only runs with reported usage; missing usage is not treated as zero.
- Select a result cell for the recorded problem, reference answer, model response, and all three attempts. Missing responses and usage remain explicitly unreported.
- Tokyo uses a private source dataset. Review access before publishing its problem and response details.
Model Accuracy vs Pass@3
100%
75%
50%
25%
0%
GPT-5.5
Claude Opus 4.8
Gemini 3.5 Flash
K-EXAONE-236B-A23B
DeepSeek V4 Pro
Solar Pro 3
KT Mi:dm 2.0 Base Instruct
Accuracy
Pass@3
Avg Tokens per Reported Run
8.7K
6.5K
4.4K
2.2K
0
K-EXAONE-236B-A23B
Solar Pro 3
Gemini 3.5 Flash
DeepSeek V4 Pro
KT Mi:dm 2.0 Base Instruct
Claude Opus 4.8
GPT-5.5
Avg Tokens / Problem
Entrance Exams brings university entrance examinations together, including verbal reasoning and mathematics. Results remain separate by exam, year, subject, and evaluation language.
Results are reported using Pass@3 metrics to account for generation variance. Detailed execution traces are available for transparency.
Performance Legend
Mastery (100%)
3/3
Strong (66%)
2/3
Weak (33%)
1/3
Fail (0%)
0/3
Leaderboard / Tokyo · Mathematics · English
Change benchmark ↑Scroll horizontally to see all problems. Select a result cell to view the problem and recorded response.


