Kardashevian

Mission

Measure what can’t be memorised

Most AI benchmarks are saturated within a year of release: their answers leak into training data, or models simply reach the ceiling. We want evaluations that stay hard as models improve.

The name

The Kardashev scale ranks civilisations by how much energy they can harness, from a planet to a galaxy. We borrow the idea of a scale with no ceiling in sight. Kardashevian measures how much of the real, unknown world a model can handle: futures that haven’t happened and rules it has never been shown.

Three rules

  1. Unsaturable. Every arena has a moving or hidden answer. When a model nears the ceiling, the arena gets harder rather than being retired. When GPT-6 Luna came close to solving our first rover world, we built a much harder one.
  2. Committed before graded. Models answer from frozen inputs, and the answer is decided later by the market or by physics they never see. No one, including us, can tune the result.
  3. Uncertainty shown. Every rating carries its 95% interval, and early results are labelled provisional. We publish our corrections alongside our results.

Which models, and why

We run on a small budget, so for now we test the most affordable models from the frontier labs. Every arena is built to take any model through the same API, prompts and budgets. As we grow, we will add the labs’ top models and rate them on the same scale as everything already here.

What we won’t claim

A high rating here means a model did better than others on these tasks, over the rounds so far. It does not mean a model is superintelligent, safe or suitable for any particular use, and nothing on this site is investment advice.