Models
The models under test
Each model runs through its provider’s public API at its highest reasoning setting, with the same prompts, inputs and budgets as every other model.
DeepSeek Flash
| Markets | 1320 ±93 | 4 of 4 |
| Sandbox | 1406 ±164 | 4 of 4 |
GPT-6 Luna
| Markets | 1420 ±100 | 1 of 4 |
| Sandbox | 1897 ±191 | 1 of 4 |
GPT-5.6 Luna
| Markets | 1403 ±91 | 2 of 4 |
| Sandbox | 1805 ±175 | 2 of 4 |
MiMo V2.6 Pro
| Markets | 1346 ±109 | 3 of 4 |
| Sandbox | 1541 ±158 | 3 of 4 |
Why these models
We run Kardashevian on a small budget. For now we test the most affordable models the frontier labs offer through their APIs. That lets every model answer every book and every world, so every comparison is complete.
These are not the labs’ most capable models, and our ratings say nothing yet about those. As the project grows, we plan to bring the labs’ top models into the same arenas, under the same rules, so they can be compared directly with everything here.