Kardashevian

Insights · 6 Oct 2026

When a model thinks past its budget

In the 6 October sandbox world, MiMo V2.6 Pro spent all 100,000 output tokens reasoning on one attempt and returned no code.

Every sandbox attempt is one reply from the model: it reads the brief and its results so far, then writes a controller for the rover. Each reply may use up to 100,000 output tokens, and reasoning counts toward that limit. The prompt says so in plain words and asks for the code block before any long explanation.

On its third attempt in the world started on 6 October, MiMo V2.6 Pro used the full 100,000 tokens over 24 minutes and stopped before writing any code. The attempt scored −15, the same as a crash. Its first attempt had already used 88,146 tokens; its second used 54,679 and produced its best practice run of the world, 62.9.

Why this counts

The GPT models and DeepSeek typically finish an attempt in 20,000 to 30,000 tokens. The limit is not tight for the field; one model sometimes reasons past even a generous budget. Fitting useful work into a stated budget is part of what the sandbox measures, so the attempt stands.

Two things make it harder for MiMo than for the others. A model cannot count its own tokens as it thinks, so a number in the prompt is weak guidance. And MiMo's API offers only thinking on or off, while the OpenAI and DeepSeek models run with a reasoning-effort setting their providers use to pace the work.

What changes

We are adding each past attempt's actual token use to the feedback every model receives, for example “attempt 1 used 88,146 of 100,000 output tokens”. It applies to every model equally and only to new worlds, so finished worlds stay comparable.