Method · version 1.0
How we score
Every arena follows the same order: frozen inputs, committed answers, then grading by something the model cannot see or influence. Ratings always carry their uncertainty.
Models and settings
Models run through their providers’ public APIs. OpenAI and DeepSeek models use their maximum reasoning effort; MiMo has thinking switched on, its only setting. Every model in an arena receives the same prompt, the same inputs and the same output budget: 32,768 tokens per Markets prompt and 100,000 per Sandbox reply, reasoning included.
For budget reasons we currently test the lower-cost models from frontier labs, not their flagship models; the Models page explains the choice and our plan to add top models.
Markets
A book starts from a frozen snapshot of 300 instruments: US stocks, ETFs and spot crypto pairs, with adjusted daily closes up to the last completed trading day. Each model answers 15 prompts, one per horizon (1, 3, 9, 12 and 24 months) at each risk level (low, medium, high). Each answer is a long-only portfolio with weights, a maximum expected drawdown and a return forecast for its horizon.
We track every portfolio daily from the snapshot’s closing date against SPY, using adjusted closes. A horizon matures on its calendar due date; its forecast is then graded on the first close on or after that date.
Rating
A cell is one book, one horizon and one risk level. In each cell, every pair of models plays a match decided by which portfolio is further ahead of SPY; SPY itself is a player at 0. Results within 0.5 percentage points are draws. Because the 15 cells of a book share a start date and a market, each book counts as about three matches per pair, and a matured cell counts twice as much as a running one.
Ratings are fitted with a Bradley–Terry model on the Elo scale, with a weak prior toward 1,500 and SPY fixed at 1,500. The 95% interval comes from how sharply the results pin each rating down. Ratings average matured and running portfolios, so they move with the market until each horizon closes; the leaderboard says so whenever no horizon has matured.
How the models invest
The Markets page also describes what goes into the portfolios. Each holding is put in one plain group (bonds, gold and metals, defensive stocks, index and dividend funds, energy and power, tech and chips, crypto, or other stocks), plus cash. “What has helped” is each group’s move against SPY while held, weighted by position size, over portfolios still running. Only symbols and weights are published; the models’ written reasoning stays private.
Sandbox
Each world is generated from a seed: a 40-metre arena with walls of rocks broken by one or two narrow gaps, dead-end pockets, 1 to 4 checkpoints to visit in order, and a goal. The physics is hidden: motor dead zone and imbalance, a 100 to 350 ms command delay, compass bias and drift, GPS noise, imperfect range sensors, wind with gusts, and ice. Every layout is checked to be solvable.
A model writes a controller in JavaScript and runs it on the practice course. It sees its score, the outcome and a trace of the run, then tries again: 8 attempts in all. Its best controller is then scored on 10 unseen courses generated the same way. Reaching the goal earns 100 plus up to 50 for speed; otherwise up to 90 for progress through the checkpoints; a crash costs 15.
Rating
Each finished world is a round: every pair of players meets on their unseen-course scores, with results within 3 points as draws. A careful reference controller that does not learn is fixed at 1,500, and a random controller is included for scale. The same Bradley–Terry fit as in Markets gives ratings and intervals. The learning score is the mean best-so-far practice score across a world’s 8 attempts, so a model that finds a good controller early scores higher.
A world is left out of ratings when no model returned any code, or when a model was stopped early by a provider failure and not resumed.
What we publish
The site shows numbered releases only. Each release is built in our lab with the same code that computes the lab’s own numbers, checked so it carries no prompts, model replies or code, and stored unchanged. You can download every release from the Data page.
Changelog
| Date | Area | Change |
|---|---|---|
| 7 Oct 2026 | Markets | Markets page adds how the models invest: results by risk level, where the money goes, what has helped or hurt, and each model’s most-held holdings. |
| 6 Oct 2026 | Markets | Market snapshot through 5 Oct added; new books freeze it. |
| 2 Oct 2026 | Ratings | Elo fit fixed in both arenas. It used a fixed step size that oscillated; it now converges. DeepSeek Flash moved from first to last in Markets. See the insight “Correcting our ratings”. |
| 30 Sep 2026 | Markets | An internal spending cap paused the 29 Sep book after 3 of 4 models had answered. The missing DeepSeek Flash prompts ran on 30 Sep with the same frozen prices. |
| 28 Sep 2026 | Sandbox | Output budget raised from 32,000 to 64,000 and then 100,000 tokens per reply; worlds run on the lower limits are withdrawn from ratings. |
| 28 Sep 2026 | Sandbox | World generator v2: walls with gaps, ordered checkpoints, five range sensors and harder hidden physics, after v1 came close to being solved. |
| 28 Sep 2026 | Sandbox | Worlds started before long model calls stopped being cut off are withdrawn from ratings. |
| 27 Sep 2026 | Markets | 1-month horizon added: 15 portfolios per model per book. |