Documentation

Methodology

How OpenWatch turns raw, noisy global data into structured intelligence. This page describes what the scores mean, where the data comes from, and where the platform's limitations are. If you're using OpenWatch for research or investment decisions, read the limitations section.

How to Read Our Scores

OpenWatch uses two opposite color conventions, so the same color can mean very different things depending on the metric. Conviction scores run high-is-good (green); instability and threat scores run high-is-bad (red). Use this legend to read any badge on the platform.

78STRONGConviction / Signal Breadth

Higher = more independent signal families agree. Green is stronger conviction; amber is weaker. Breadth of agreement, not a price target.

82CRITICALInstability / Threat

Higher = more risk or disruption. Red/amber is elevated; green is calm. This is the opposite color direction from conviction.

63%Probability

A market-implied percentage from a prediction market — what traders collectively price an event at, not an OpenWatch forecast.

#4Rank

Position in a sorted list (e.g. 4th by signal volume). Relative ordering only — a #1 is not necessarily high on any absolute scale.

All scores are analytical signals, not recommendations. See the limitations section below before acting on any of them.

1 — Signal Collection

OpenWatch continuously pulls from roughly 1,000 open-source feeds — wire services, government press releases, financial filings, think-tank publications, regional newspapers, and specialist newsletters. Collectors run on a rotating schedule; most feeds refresh every 15–60 minutes. The total corpus is over 157,000 processed signals.

Each incoming item is deduplicated by URL and content hash before processing. Items shorter than 120 characters are discarded as noise (headline fragments, social-media stubs). Remaining items enter the enrichment pipeline: language detection, entity extraction, geographic tagging, and sentiment scoring.

SourceTypeUpdateCoverage
GDELT 2.0News events15 minGlobal, 65+ languages
ACLEDConflict eventsWeeklyGlobal
USGSEarthquakes5 minGlobal, M2.5+
WHO / ProMEDDisease outbreaksHourlyGlobal
UN OCHA ReliefWebHumanitarianHourlyGlobal crises
SEC EDGARFilings (US)HourlyUS-listed companies
HKEX, EDINET, Companies HouseFilings (intl)DailyHK, JP, UK
Curated RSS (1,000+ feeds)News, gov, NGOsHourly45+ languages
Central banksRate decisionsEvent-basedG20 + select EM
Sanctions lists (OFAC, EU, UK)DesignationsDailyGlobal
STOCK Act filingsEquity disclosuresDailyUS Congress
Commodity exchangesPrices, supplyDailyEnergy, ag, metals
IMF World Economic OutlookMacro indicatorsWeekly130+ countries
World Bank Open DataReserves, debt serviceWeekly180+ countries
Polymarket, Kalshi, Manifold, PredictItPrediction marketsHourlyUS politics, geopolitics, macro

2 — Country Risk Score (0–100)

The country score is a single number reflecting recent signal intensity and quality for a given country. It measures how much attention the country is drawing from credible sources right now, weighted by the quality of those signals — not an absolute geopolitical risk index.

score = avg_quality × √(volume_deviation)
  • avg_quality — the mean quality score of signals collected in the last 30 days. Quality is derived from source authority (established outlets score higher), article length, and whether the item was picked up by multiple independent publishers. Range: 0–1.
  • volume_deviation— how much the current 7-day signal volume departs from the country's own 30-day rolling baseline. A value of 1.0 means no change; above 1.0 means elevated volume. The square root dampens extreme spikes — a 4× volume surge contributes a 2× multiplier, not 4×.

Scores are recomputed hourly. A country with no recent signals returns to its baseline rather than holding a stale high score. Score changes over 7 and 30 days are surfaced as trend percentages.

Each country is also assigned a posture (opportunity, risk, or neutral) and a recommended action (ENTRY, EARLY, WATCH, REDUCE, EXIT, AVOID) based on score level and trend direction.

3 — Signal Frequency Bands

The five frequency bands describe where a country's current signal volume sits relative to its own baseline — not relative to other countries. A quiet country moving from QUIET to MODERATE is a bigger change than a perpetually active country sitting at MODERATE.

BandWhat it means
QUIETBelow-baseline volume. Fewer signals than usual for this country.
LOWNear-baseline. Normal coverage levels.
MODERATE1.5–2× baseline. Elevated but not unusual for a news cycle.
ELEVATED2–4× baseline. Sustained above-normal attention; warrants review.
HIGH4× baseline or above. Major event driving coverage spike.

Thresholds are calibrated per-country over a 30-day rolling window. A single-day spike does not upgrade the band unless sustained for at least 48 hours.

4 — Sentiment Analysis

Every signal is scored on a sentiment dimension from –1.0 (most negative) to +1.0 (most positive). The primary model is XLM-RoBERTa, fine-tuned on geopolitical and financial news. When XLM-RoBERTa is unavailable, a BART-based fallback handles classification.

Country-level sentiment is decomposed into three components:

  • Level— the current 7-day average sentiment, normalized against the country's own long-term mean. A country that always scores –0.3 is not “negative” in the alert sense — level captures deviation from that country's norm.
  • Velocity — the rate of change over the last 7 days. High negative velocity means sentiment is deteriorating fast; positive velocity means conditions are improving. Alerts fire on velocity thresholds, not level alone.
  • Dispersion — how spread out individual signal scores are around the mean. High dispersion means mixed signals — some sources positive, others very negative. High dispersion is often an early warning sign before a consensus forms.

Sentiment scores are directional indicators, not precise measurements. Model performance degrades on technical financial language, satire, and highly idiomatic text.

5 — Congressional Trades

The STOCK Act (2012) requires members of Congress and their spouses to disclose stock transactions within 45 days of the trade date. OpenWatch ingests these disclosures from the House and Senate portals and normalizes them into a single searchable database currently covering over 5,500 disclosures spanning purchases, sales, and exchanges across 2,975 Representatives, 174 Senators, and 524 Executive Branch filers.

Each record shows the member, transaction type, asset, amount range (STOCK Act permits range-based reporting, not exact amounts), and both the transaction date and the disclosure date. Committee assignments and party affiliation are cross-referenced for filtering.

Return % calculation: where possible, OpenWatch estimates a return by comparing the asset price on the trade date to the current price. This number has hard limitations:

  • Amount ranges (e.g., “$1,001–$15,000”) make precision impossible. Returns are calculated at the midpoint.
  • The 45-day disclosure lag means the trade happened weeks before you see it.
  • Members may report in bulk on the last day of the window — actual timing within the range is unknown.
  • Partial sales, options, and funds with multiple underlying positions complicate simple price-based return math.

The return percentage is a rough heuristic for identifying patterns, not a precise accounting of any member's gains or losses.

6 — Prediction Markets

OpenWatch aggregates markets from four platforms: Polymarket (crypto-settled, high liquidity), Kalshi (CFTC-regulated, real money), Manifold (play money, broad topic coverage), and PredictIt (real money, US political focus). Markets are fetched hourly and matched to geopolitical topics and countries using keyword and NLP matching.

Two quality dimensions are surfaced for each market:

  • Match Quality — how confidently we matched this market to the associated country or topic. High = the market title explicitly names the country/event. Low = inferred from keywords.
  • Market Quality — a composite of liquidity, volume, and number of active traders. Markets with very few traders or tiny volume are flagged as thin. Thin-market prices can be moved by a single participant and should not be read as reflecting genuine crowd wisdom.

Markets with fewer than 10 active traders or under $500 in total volume are marked as thin. The platform surfaces both deep and thin markets — read the quality scores before drawing conclusions from a probability.

Platform Edge

When a scenario branch has an OpenWatch model probability assigned, each matched market chip shows a Platform Edge— the gap between the OpenWatch scenario model's estimate and the market's calibrated price, expressed in percentage points.

A positive edge means OpenWatch's model assigns a higher probability than the market — signals may be pricing in something the market has not yet reflected. A negative edge means the market is more bullish than the model.

  • High Conviction — gap ≥ 15pp. The divergence is large enough to warrant closer investigation.
  • Medium Conviction — gap 8–14pp. Meaningful but within normal uncertainty range.
  • Low Conviction — gap 3–7pp. Minor divergence; interpret cautiously.

Platform Edge is not investment advice. Model estimates carry uncertainty and the market may be right. Always read both the scenario context and the market quality scores before acting on any edge signal.

7 — Structural Risk Profile

Signal frequency alone cannot distinguish a country in structural crisis from one experiencing a brief news spike. The Structural Risk Profile on each country page adds six slow-moving macroeconomic indicators sourced from the IMF World Economic Outlook and World Bank Open Data, updated weekly:

  • Debt-to-GDP — General government gross debt as a share of GDP. Above ~90% is typically flagged as elevated.
  • Fiscal Balance — Government net lending/borrowing as a share of GDP. Persistent deficits above −5% indicate structural fiscal stress.
  • Current Account — External current account balance as a share of GDP. Large sustained deficits (<−4%) signal dependency on foreign capital.
  • FX Reserves — Expressed in months of import coverage. Below 3 months is a standard vulnerability threshold.
  • Debt Service Ratio — Total debt service as a share of GNI. Above 20% constrains fiscal flexibility.
  • CDS Spread — 5-year credit default swap spread in basis points; market-implied probability of sovereign default.

These indicators provide context for interpreting signal-frequency scores but do not yet feed into the country score formula above. Integration into the scoring model is planned for a future release.

8 — Data Retention

Signal text is subject to tiered retention windows based on the significance score assigned during enrichment. Lower-significance items are pruned on a rolling schedule to manage database size; higher-significance items are kept longer for historical research.

TierSignificance ScoreFull Text Retained
Low< 4.53 months
Medium4.5 – 6.512 months
High≥ 6.536 months

Signal metadata — source URL, publication date, country codes, and significance score — is retained indefinitely for historical audit and trend analysis. Only full signal text is subject to these retention windows. Pruning runs nightly.

9 — Limitations

Not financial advice

Nothing on OpenWatch constitutes investment advice. Country scores, congressional trade patterns, and prediction market probabilities are informational signals, not recommendations. Always do your own due diligence before making any financial decision.

  • English-language source bias. The vast majority of sources publish in English. Events covered primarily by non-English media — particularly in Central Asia, sub-Saharan Africa, and Southeast Asia — will be systematically underrepresented. Scores for these regions should be read with extra skepticism.
  • Disclosure lag. Congressional trades are reported up to 45 days after the fact. By the time you see a disclosure, the opportunity may already be priced in.
  • Model errors. XLM-RoBERTa and BART can misclassify sentiment on technical text, satire, and culturally specific idioms. Named entity recognition struggles with transliterated names and alternative spellings. Country assignment errors are most common for multi-country stories.
  • Volume without quality is not a crisis.A media event — a major speech, an election anniversary — can spike a country's signal volume without any change in underlying conditions. The quality weighting dampens pure-volume spikes, but the mechanism isn't perfect.
  • Thin prediction markets. Some topics have deep, liquid markets. Others have one thin market with a handful of traders. The platform surfaces both — read the quality scores before drawing conclusions.
  • Analytical outputs, not recommendations. Signal scores, posture classifications, and investment implications represent structured hypothesis generation, not advice.

Last updated: May 2026

3485b22