Skip to content
Vaticinus

Test the model. Test the method.

A forecasting harness has to improve the answer enough to justify its added cost. We publish the comparisons, including the ones that fail.

A fair comparison uses the same model, questions, evidence and information cutoff. One arm makes a direct forecast; the other uses the harness. We retain errors and abstentions alongside issued probabilities, then compare Brier scores, coverage, cost and latency.

What the current record says

TestResultWhat it establishes
Eight 2025 FOMC meetingsLlama: direct 0.235, harness 0.213 Brier. Qwen: direct 0.235, harness 0.250.No reliable harness advantage. The apparent Llama gain came from one answer with contradictory arithmetic.
Three-event prospective pilotDirect: 9/9 issued. Harness: 0/9 issued.Timeouts and HTTP 400 failures; the exact provider error was not retained. The run also violated its transport-error stop rule. This exploratory pilot cannot establish a forecasting improvement.

Recorded results as of 22 September 2026. Lower Brier is better. These small, correlated cohorts do not measure broad forecasting skill; later software repairs do not change their original scores.

Historical replay is useful, with limits.

An older checkpoint can be tested on events that occurred after its release. That lets us investigate resolved outcomes without waiting months. We must still document the checkpoint, freeze the question and supply only evidence available at the simulated issue date.

A stated training cutoff does not prove which weights a hosted endpoint serves. A date printed on a page retrieved today does not prove that its contents were available then. Our FOMC replay documents both limitations. It is a historical experiment, not a prediction originally issued in 2025.

Prospective evidence earns the claim.

Short-horizon questions let us record forecasts before outcomes, then score them as they resolve. We publish the selection rule and preserve every attempted question. A model must beat a declared baseline; the harness must add value over that same model.

The evaluation walkthrough explains the existing studies. The open runner supports new comparisons. Independent benchmarks such as ForecastBench provide an external reference; their leaderboard results are not scores for this software.