Join the Movement
Geopol — live todayGlobal Risk Exchange — in IFSCA sandbox testing
Home/Geopol/Methodology
Methodology · interactive companionEarning the Right to Forecast · Aug 2026
THE EPISTEMIC POSITION

Forecasting Geopolitics and Current Events

Transforming forecasting from a ‘game of chance’ to a ‘game of skill’.

Innovating and gamifying current events or geopolitical issues requires first converting them into event contracts. An event contract cannot be open-ended: it must pose a question that resolves unambiguously within a fixed window, with a defined start date and end date. That is what transforms the activity from punditry into informed, time-specific analysis.

Sample Question A — Expert commentary

“Tensions in the region may well escalate, and the coming period bears close watching.”

— Senior analyst, televised panel

Whatever happens next, this sentence survives. But go ahead — try.
Sample Question B — A Geopol forecast
Will Country X impose formal capital controls by 31 December 2026?
Yes / No
A well-framed question: one precise event, a fixed deadline, two outcomes. Check it against (simulated) reality to reveal Geopol's forecast and how it scores.
Probability tape · sample Geopol outputStrait of Hormuz disruptionP 0.18 ▼G7 sanctions escalationP 0.41 ▲EM sovereign downgradeP 0.27 ▼Red Sea shipping haltP 0.33 ▲Central bank surprise cutP 0.22 —Cross-border capital controlP 0.12 ▼Major election upsetP 0.46 ▲Commodity supply shockP 0.29 —Probability tape · sample Geopol outputStrait of Hormuz disruptionP 0.18 ▼G7 sanctions escalationP 0.41 ▲EM sovereign downgradeP 0.27 ▼Red Sea shipping haltP 0.33 ▲Central bank surprise cutP 0.22 —Cross-border capital controlP 0.12 ▼Major election upsetP 0.46 ▲Commodity supply shockP 0.29 —
The forecast lifecycle — Figure 1, made walkable
WHAT A FORECAST IS

A forecast is a six-tuple. Remove any element and watch it collapse.

Every Geopol forecast is the object ⟨ q, p, h, t₀, A, R ⟩. The paper claims each element does work — that removing any one of them collapses the construction into punditry. Don't take that on faith. Remove them.

✓ This is a forecast — it can be wrong, measurably
q — Country X imposes formal capital controls
p — 0.22

h — by 31 December 2026

t₀ — issued 18 August 2026; evidence base frozen

A — 4 named scenarios · mechanisms · evidence links (see §3)

R — notification published in the official gazette
DETERMINING THE DRIVERS

The number is computed from the structure. Change the structure.

The engine is never asked “what is the probability?” It must produce named scenarios — each with a mechanism, actors, a timeline, and evidence links — and the probability is the mass on the scenarios where the proposition holds. Re-weigh the scenarios below and watch the number follow. It cannot do anything else: one is computed from the other.

Will Country X impose formal capital controls by 31 December 2026?R: notification published in the official gazette · anchored base rate for this class: 11% (see §4)
s1 · Reserve depletion forces controlss ⊧ q

FX reserves fall through the import-cover floor; the central bank moves to formal controls to halt the drain.

evidence links: 5 · actors: central bank, finance ministry · timeline: Q3–Q4
0.14
s2 · External program with conditionalitys ⊭ q

A multilateral package arrives with disbursement conditions; controls become unnecessary and politically costly.

evidence links: 7 · actors: IMF, cabinet · timeline: Q3
0.36
s3 · Political transition accelerates outflowss ⊧ q

A contested succession triggers capital flight faster than reserves can absorb; controls follow within weeks.

evidence links: 3 · actors: ruling party factions · timeline: any quarter
0.08
s4 · Muddle-through: informal restrictions onlys ⊭ q

Import licensing quietly tightens but no gazette notification is issued — the criterion R is never met.

evidence links: 6 · actors: commercial banks, regulator · timeline: continuous
0.42
scenarios where q holdsscenarios where it does not
0.22praw — affirmative mass, Σ w(s) over s ⊧ q

Weights renormalise to sum to 1, exactly as in the paper. A reader who disputes the number doesn't argue with a decimal — they argue with a specific link in a specific chain. That's the point of the artifact.

THE BRIDGE: BASE-RATE ANCHORING

Whoever chooses the class chooses the prior. Now you choose.

Before examining the specifics of a case, the forecaster should hold the empirical frequency of events like this one as a prior. But “like this one” is a decision, not a fact — and the decision determines the answer. Pick a reference class for the question “Coup attempt in State X within 12 months?”

4.0%the anchor the decomposition must argue away from
Broad and stable. Nearly 500 cases — the estimate barely moves if you drop a decade. But is a coup in 1960s Latin America really “an event like this one”? The anchor is defensible, and vague.

Each anchor was defensible in isolation — and they differ by a factor of eight. That's why the class, its defining criteria, and its case count travel with the anchor into the artifact. An undisclosed reference class is an appeal to authority with a decimal point.

EXTREMIZING

Pooled estimates hedge. The correction sharpens — in log-odds.

Combining estimates that each carry part of what is collectively known produces a consensus closer to the midpoint than the information warrants. The correction scales the pooled estimate p̄ in log-odds space: it preserves ordering, sharpens magnitude, and never crosses 0.5. Drag a.

0.000.000.250.250.500.500.750.751.001.00pooled estimate p̄published p
a = 1 is the identity — no correction. a > 1 pushes mass away from the hedge.
logit(p̄)−0.847
a · logit(p̄)−1.356
published p0.205

A fixed a is a prior, not a finding. Where a recalibration map has been estimated from the resolved record (§5, below), it doesn't audit this value — it supersedes it.

SCORING: WHY HONESTY IS THE UNIQUE OPTIMUM

Try to beat honest reporting. You can't — provably.

Forecasts are scored with the Brier score, BS = (p − o)² — a strictly proper rule. Your expected score is minimised by reporting exactly what you believe, and the minimum is unique. Below, you privately believe q. Choose what to report, and watch your expected score.

0.000.000.250.250.500.500.750.751.001.00reported probability pexpected Brier score (lower is better)honesty p = q
Shade toward safety, inflate to look decisive — try any strategy.
E[BS] if honest (p = q)0.210
E[BS] as reported0.250
cost of misreporting (p − q)²+0.040

Shading toward safety: costs exactly (p − q)² = 0.040 in expectation, every time.

THE CONVERGENCE ENGINE

Watch a biased system correct itself — using nothing but its own record.

This simulator is an overconfident forecaster: directionally right, numerically overstated. It doesn't know that. Resolve its forecasts against (simulated) reality and its calibration curve appears — bias becomes an observable with a magnitude and a sign. Then apply the recalibration map ĉ, learned from the resolved record alone.

0.000.000.250.250.500.500.750.751.001.00stated probabilityobserved frequencyperfect calibration
resolved pairs n0
Brier score—
reliability (calibration error)—
resolution (discrimination)—

No record yet. A forecast that is never resolved is a write that never commits.

What you should notice

At small n the dots jitter — precision near the tails costs 1/√n, which is why a calibration curve earned purely in real time is years per class, not weeks. The dots below the diagonal on the right and above it on the left are the overconfidence signature. Recalibration is monotone: it fixes the systematic overstatement without reordering which events the system called more likely. Reliability falls toward zero; discrimination survives, except where the record itself couldn't justify it.

THE CONVERGENCE CLAIM, STATED CONDITIONALLY

The claim names its preconditions. Each one can fail. Read how.

Convergence asserted without assumptions is exactly the unfalsifiable promise the paper argues against. Four conditions carry the claim — expand each to see what protecting it means, and precisely how the loop degrades when it breaks.

When it holdsOutcomes are determined by search toward the world — never toward the score the system would prefer. Resolution is kept structurally separate from forecasting.
When it breaksThe loop inverts. The system stops learning to forecast accurately and starts learning to resolve favourably. This condition is protected institutionally, not mathematically — it is the one most worth attacking.
When it holdsNo reporting strategy beats honesty — you verified this yourself two sections up. Every parameter fitted against the score inherits a target that points at truth.
When it breaksThe cheapest condition to satisfy, the easiest to lose by accident: a scoring scheme adjusted for presentational reasons can silently stop being proper — and strategic reporting becomes optimal again.
When it holdsAn event that occurs is observable through some channel at resolution time. Under-determined questions are voided — they leave the record instead of entering it as false negatives.
When it breaksThe quietest failure: an event occurs, no channel records it, absence reads as absence-of-event. The low forecast scores as correct and the map learns to go lower — the loop reinforces the error. In poorly observed domains the honest claim narrows: probabilities converge to the frequency of observed occurrence, and the reader is entitled to weigh a forecast accordingly.
When it holdsThe resolved sample covers the probability range, and the world doesn't change faster than it is sampled. Recent resolutions are weighted more heavily; calibration is fitted within event classes, guarded by a minimum sample.
When it breaksGeopolitics genuinely stresses this one. Regimes change; a relationship that held for a decade breaks in a season. Under drift, convergence becomes tracking — calibration error bounded by the rate of drift relative to the rate of resolution.
THE CONTRACT

One last trap: an honest system can still pick easy questions.

Strict propriety is a per-question property. It guarantees no individual forecast improves by misreporting — and guarantees nothing about which questions were asked. Both systems below are scrupulously honest. Toggle the question mix.

System K — flatters its record

Asks only questions whose answers were never much in doubt: “Will the incumbent government of a stable state survive the month?”

Raw Brier Score0.040

System L — asks what matters

Asks genuinely contested questions: transitions, escalations, controls — the ones a reader actually needs.

Raw Brier Score0.180

On raw scores, System K looks four times better. It isn't. Flip the toggle: difficulty becomes visible, and the gap nearly vanishes. This is why any figure Geopol publishes carries its question mix and is reported as a skill score. An aggregate without its question mix is not a track record. It is a selection.

The closing position

The system is not asking to be believed. It is asking to be checked.

What this system offers is not a better guess. It is a different kind of object than a guess: auditable in its reasoning, falsifiable in its claims, and self-correcting by construction. Whether it forecasts well is a question its record will answer — which is the only way that question has ever been properly answered.