VUCA
VUCA NEWS · METHODOLOGY

How the numbers are made

Every number on this site is computed by code from stored inputs, and those inputs are kept so any number can explain itself. This page describes the current methods plainly. When a method changes, the change is versioned — history is never recomputed.

THE PIPELINE

Sources (RSS, GDELT, official feeds, scraped outlets) are ingested continuously and routed to dynamics by topic. An extraction model proposes atomic claim candidates with quotes. A clustering pass folds near-duplicate reports into one lead and counts the fold as corroboration. An adversarial verification model recommends approve/review/reject with rationale — carrying each source's measured reliability and any state-affiliation label into that judgment. Then a human editor decides. The only exception: claims from tier-1 sources with a clean AI verify and confidence ≥ 0.8 may auto-publish under a policy the pipeline cannot widen — and sources that are state-controlled or below a reliability bar are excluded from even that.

THE VUCA SCORE (ENGINE V0)

Each dynamic scores 0–100 on four components, daily. In brief: Volatility — anomaly of event/claim arrival rates vs a 28-day baseline, plus measured world-state telemetry: our live AIS vessel samples, and (since 2026-07-08) IMF PortWatch daily chokepoint-transit counts — each gate series evaluated on its own latest published day, so PortWatch's ~2-day lag can never read as a false collapse. Uncertainty — verification reject rate, extractor confidence spread, forecaster dispersion, hedged language. Complexity — actor count, relationship variety, domain breadth. Ambiguity — contested-claim share, unresolved candidates, and measured worldview divergence between published factions. The headline is the component mean. Young dynamics blend against an analyst-seeded baseline; the real-signal weight rises automatically with history and is disclosed on each dossier's score breakdown. Confidence reflects source coverage, freshness, provenance, and agreement. Every daily snapshot stores its full input vector — the “why this score” module on each dossier renders those stored inputs directly.

FORECASTS

Questions are binary, dated, and resolvable by a stranger from public sources — criteria are written before opening. Submissions are scored with Brier and log scores at resolution; the crowd aggregate excludes our AI baseline (@vuca_ai), which is scored as a benchmark and published win-or-lose. The full board is on /accuracy.

SOURCE RELIABILITY

Sources are scored from their measured record in this pipeline only: P = approvals + 0.5·corroborations given + 0.5·corroborations received; N = rejections + 2·later contests + verifier flags; score = 100·(P + K/2)/(P + N + K) with K = 6 shrinking new sources toward a neutral 50. Duplicate folds are neutral. Identity and state-affiliation labels are machine-drafted with cited evidence and publish only after editor review. Full ledger: /sources.

CONTESTED REALITY INDEX

A weighted blend — ambiguity 40%, worldview divergence 35%, contested-claim share 15%, forecaster disagreement 10% — renormalized over whichever inputs exist for a dynamic, so the index is honest from day one and sharpens as coverage deepens: /divergence.

STRUCTURAL DASHBOARDS

Dossier dashboards (dollar security, PLA activity, the Ukraine Initiative Index) are built from named public series — FRED, IMF COFER, Taiwan MND bulletins, Oryx confirmed losses, DeepState map layers — each cited on the widget, with method notes where our computation adds a step (e.g. geodesic area used for momentum only). Structural series never feed the live VUCA score; they are context, and the engine excludes their namespaces explicitly.

REPLAY CALIBRATION · THE ENGINE AUDITS ITSELF

Every published snapshot stores its full scoring input vector, so the whole record can be re-scored from scratch. We periodically replay it with the current formulas and publish the result — if a formula changed since a snapshot was scored, the replay names the day.

WINDOW 2026-07-042026-07-11 · 210 SNAPSHOTS · 34 SITUATIONS
COMPONENT REPLAY: 100% EXACT (840/840)
HEADLINE REPLAY (SEEDED): 63/64 EXACT · 1 WITHIN ±1 (COMPONENT ROUNDING)
MOVES ≥3PTS: 38 — WORLD-DRIVEN 26 · REPORTING 8 · COMPOSITION 4
LEAD TIME: BUILDING (0/10 SAMPLES BEFORE WE SHOW A NUMBER)
REPLAYED 2026-07-11 · ARTIFACT: /calibration.json

Current formulas reproduce the entire published record. Moves classified by their stored sub-inputs: world = measured events/telemetry/transits shifted; reporting = coverage volume shifted; composition = actor/contestation structure shifted. We publish this split because an index that mostly moves on its own coverage is an attention meter, not a world meter.

KNOWN LIMITS

Coverage is uneven across theaters; visually-confirmed loss data undercounts; assessed-control maps carry their makers' judgment; young dynamics lean on seeded baselines; and models make mistakes — which is why machine output is human-gated by default (the sole exception, the labeled tier-1 auto-publish lane, is policy-bounded and auditable), and why the record of our misses stays public.