Deep20Bench · Static publication summary

Deep20Bench results.

Compare official model scores, outcomes, costs, time, and stability.

Official leader

Claude Fable 5 (high)

The current leader has a question score of 12.06. Lower is better.

  1. 1 Claude Fable 5 (high) 12.06 questions
  2. 2 Claude Opus 5 (high) 12.34 questions
  3. 3 Kimi K3 (high) 12.74 questions

Local preview

Open this publication through HTTP.

The app loads small JSON files on demand. Browsers block those requests when this file is opened directly.

npm run --prefix source/publication/site dev -- --host 127.0.0.1 --port 4173

Then open http://127.0.0.1:4173/deep-20-bench/