Inaugural season · rules version 1.1 · scoring v6
Win by getting closer to reality.
Predictions remain editable until Friday, October 30, 2026 at 23:59:59 UTC. A verified public release before then closes entries for the affected model at its official publication time.
Prize
The #1 eligible forecaster receives a single-use digital card loaded with up to US$100 to buy one month of ChatGPT Pro 5x or Claude Max 5x. The prize is guaranteed even with one eligible entrant. Frontier Forecast Arena funds it; OpenAI and Anthropic do not sponsor it.
Eligibility
Publish both model calls before locking. Entrants must be at least 18, legally able to receive the prize, and use one account per person. The operator, its contractors, and their households are ineligible. Automated, duplicate, or manipulated entries are void.
Model deadline
A named model must become generally available through an official API or first-party paid product by October 31, 2026 at 23:59:59 UTC. An announcement, rumor, benchmark leak, or private preview is not a release. If the exact subject never becomes public, “no public model” resolves as correct.
Winner and ties
The highest settled Elo wins. If Elo is tied, lower inaugural-round error wins. If both remain exactly tied, the US$100 card value is split equally. Future seasons may publish separate prizes; career Elo continues across seasons.
Reproducible scoring contract
Lower error is better. Each model produces an error from 0 to 1; the round error is the average of Astra and Fable 5.1.
Astra
- 20% public existence or generation label
- 15% input and output price pair
- 30% three Max benchmark values
- 15% AA output tokens per task
- 20% lineup, Sol relationship, and launch timing
Fable 5.1
- 20% public existence
- 20% input and output price pair
- 40% three Max benchmark values
- 20% AA output tokens per task
Price and token error = min(|log10((prediction + 0.01) ÷ (actual + 0.01))|, 1). Benchmark error = |prediction − actual| ÷ 100. Categorical error is 0 for a match and 1 for a mismatch; grouped categories use their mean. Stored scores are rounded to six decimals.
If a model does not become public, its conditional price, benchmark, efficiency, and structure components are void and are not reweighted; existence keeps its fixed 20% model weight. If a forecaster chose “no public model” but it does launch, the predecessor snapshot shown when publishing is the frozen fallback for the hidden conditional fields.
Official resolution sources
Price uses the first generally available official USD rate per million uncached text input and output tokens, excluding batch, cache, regional, and temporary promotional discounts.
Performance uses Artificial Analysis Intelligence Index v4.1.1, Humanity's Last Exam from the same AA evaluation, AA weighted output tokens per task, and Terminal-Bench v4 with the named model, agent version, and Max setting. The first qualifying public result published by November 30, 2026 is used. A metric without a qualifying result by then is void for everyone; it is never replaced by a cherry-picked proxy.
Every settled profile will show predicted value, official value, error, observation date, and metric-specific source. Corrections require a public note and the same scoring version.
Elo and fulfillment
Everyone begins at 1500 Elo. Each settled round compares every pair of eligible forecasters: lower error is a win, equal error a draw. Expected score = 1 ÷ (1 + 10^((opponent Elo − your Elo) ÷ 400)); the pairwise updates are averaged with K = 32.
The winner is notified at the private email saved at sign-in, with one reminder, and has 14 calendar days to respond. We then schedule a short redemption window and privately deliver the single-use card details so the winner can buy the chosen plan. We never request an account password. Availability, taxes, currency conversion, and any amount above US$100 are the winner's responsibility. If the winner is ineligible or does not respond, the next eligible rank is offered the prize.
Questions or disputes: prize@frontierforecast.tech.