Forecasting Geopolitics and Current Events
Transforming forecasting from a ‘game of chance’ to a ‘game of skill’.
Innovating and gamifying current events or geopolitical issues requires first converting them into event contracts. An event contract cannot be open-ended: it must pose a question that resolves unambiguously within a fixed window, with a defined start date and end date. That is what transforms the activity from punditry into informed, time-specific analysis.
“Tensions in the region may well escalate, and the coming period bears close watching.”
— Senior analyst, televised panel
A forecast is a six-tuple. Remove any element and watch it collapse.
Every Geopol forecast is the object ⟨ q, p, h, t₀, A, R ⟩. The paper claims each element does work — that removing any one of them collapses the construction into punditry. Don't take that on faith. Remove them.
p — 0.22
h — by 31 December 2026
t₀ — issued 18 August 2026; evidence base frozen
A — 4 named scenarios · mechanisms · evidence links (see §3)
R — notification published in the official gazette
The number is computed from the structure. Change the structure.
The engine is never asked “what is the probability?” It must produce named scenarios — each with a mechanism, actors, a timeline, and evidence links — and the probability is the mass on the scenarios where the proposition holds. Re-weigh the scenarios below and watch the number follow. It cannot do anything else: one is computed from the other.
FX reserves fall through the import-cover floor; the central bank moves to formal controls to halt the drain.
A multilateral package arrives with disbursement conditions; controls become unnecessary and politically costly.
A contested succession triggers capital flight faster than reserves can absorb; controls follow within weeks.
Import licensing quietly tightens but no gazette notification is issued — the criterion R is never met.
Weights renormalise to sum to 1, exactly as in the paper. A reader who disputes the number doesn't argue with a decimal — they argue with a specific link in a specific chain. That's the point of the artifact.
Whoever chooses the class chooses the prior. Now you choose.
Before examining the specifics of a case, the forecaster should hold the empirical frequency of events like this one as a prior. But “like this one” is a decision, not a fact — and the decision determines the answer. Pick a reference class for the question “Coup attempt in State X within 12 months?”
Each anchor was defensible in isolation — and they differ by a factor of eight. That's why the class, its defining criteria, and its case count travel with the anchor into the artifact. An undisclosed reference class is an appeal to authority with a decimal point.
Pooled estimates hedge. The correction sharpens — in log-odds.
Combining estimates that each carry part of what is collectively known produces a consensus closer to the midpoint than the information warrants. The correction scales the pooled estimate p̄ in log-odds space: it preserves ordering, sharpens magnitude, and never crosses 0.5. Drag a.
A fixed a is a prior, not a finding. Where a recalibration map has been estimated from the resolved record (§5, below), it doesn't audit this value — it supersedes it.
Try to beat honest reporting. You can't — provably.
Forecasts are scored with the Brier score, BS = (p − o)² — a strictly proper rule. Your expected score is minimised by reporting exactly what you believe, and the minimum is unique. Below, you privately believe q. Choose what to report, and watch your expected score.
Shading toward safety: costs exactly (p − q)² = 0.040 in expectation, every time.
Watch a biased system correct itself — using nothing but its own record.
This simulator is an overconfident forecaster: directionally right, numerically overstated. It doesn't know that. Resolve its forecasts against (simulated) reality and its calibration curve appears — bias becomes an observable with a magnitude and a sign. Then apply the recalibration map ĉ, learned from the resolved record alone.
No record yet. A forecast that is never resolved is a write that never commits.
At small n the dots jitter — precision near the tails costs 1/√n, which is why a calibration curve earned purely in real time is years per class, not weeks. The dots below the diagonal on the right and above it on the left are the overconfidence signature. Recalibration is monotone: it fixes the systematic overstatement without reordering which events the system called more likely. Reliability falls toward zero; discrimination survives, except where the record itself couldn't justify it.
The claim names its preconditions. Each one can fail. Read how.
Convergence asserted without assumptions is exactly the unfalsifiable promise the paper argues against. Four conditions carry the claim — expand each to see what protecting it means, and precisely how the loop degrades when it breaks.
One last trap: an honest system can still pick easy questions.
Strict propriety is a per-question property. It guarantees no individual forecast improves by misreporting — and guarantees nothing about which questions were asked. Both systems below are scrupulously honest. Toggle the question mix.
System K — flatters its record
Asks only questions whose answers were never much in doubt: “Will the incumbent government of a stable state survive the month?”
System L — asks what matters
Asks genuinely contested questions: transitions, escalations, controls — the ones a reader actually needs.
On raw scores, System K looks four times better. It isn't. Flip the toggle: difficulty becomes visible, and the gap nearly vanishes. This is why any figure Geopol publishes carries its question mix and is reported as a skill score. An aggregate without its question mix is not a track record. It is a selection.