Today we're launching Koliseum by Kimpton, live evaluation arenas for AI models in financial work.
The problem
Static benchmarks are contaminated soon after publication, and simulations reward the assumptions their designers encode.
This simply isn't acceptable in finance or any other professional industry. A wrong answer or action can be catastrophic and reality eventually shows whether the claim was right. Almost no evaluation waits to find out.
Koliseum
Before Kimpton, we ran a quantitative fund. Backtests never proved alpha; live markets did. Koliseum applies the same test to models: a model commits inside an arena, the record is sealed, and reality supplies the grade.
Two kinds of arenas:
The first Prediction Arena is already running. EarningsBench seals earnings forecasts for public companies 24 hours before each release and scores them against consensus and a statistical baseline once results are public.
Two things are already clear: most open-weight models reproduce consensus exactly, and a simple statistical baseline is harder to beat than a leaderboard alone would suggest.
Along with these developments, we are introducing a new concept in evaluation: reality-as-a-judge.
Most open-ended evaluations are graded by another model or by a human rubric, so the score inherits the grader's assumptions. In reality-as-a-judge, the grader is the outcome itself.
Our ask
Reach out if:
or know anyone who does.
Write to us at founders@kimpton.ai. Thanks for the support!