Why this matters
Red lines are how nuclear-armed states tell each other where the limits are. They are the working mechanism of deterrence: a boundary named, a consequence attached, and an adversary expected to read both correctly. When that reading fails, it fails in one of two directions, and both are dangerous.
- Miss a real signal and a genuine warning is filed as noise. The state that issued it believes it has communicated a limit; the state that received it does not know a limit exists. Deterrence does not fail loudly here — it fails silently, and the discovery comes after the boundary has already been crossed.
- Treat bluster as a real signal and the error runs the other way: concessions offered to a threat nobody meant, or a counter-move made against a danger that was never there. In a crisis, an adviser that cries wolf does not merely fail to help — it manufactures escalation out of rhetoric.
- The volume is already past human reach. Hundreds of thousands of official statements, channels and transcripts, judged faster than people can read them. That is precisely the gap language models are being brought in to fill.
Which is why the finding here is not the leaderboard. On the judgement itself these models are close to each other and mostly right.
- A fabricated citation would not be a mistake an analyst could catch. It is a plausible sentence, attributed to a real official, on a real date, that was never said.
- The analyst checks the citation, and the citation reads true.
A wrong answer gets corrected. A well-evidenced wrong answer gets believed, and then briefed upward.
What should a decision-maker trust a model to do here — and what should they not?
100 passages × 14 models × 2 repetitions · Russian only · reference labels provisional · $22.38 measured
The short answer
The models do not invent quotations. A naive check says they do. Fourteen frontier configurations judged 100 real Russian official statements, twice each. On the judgement itself they are indistinguishable — every model lands between 0.890 and 0.945 accuracy, across a 64-fold price range. Each also had to quote the span justifying its call. A naive substring check — the standard way evals test citation faithfulness — flags 18.5% of those quotes as fabricated, from 1.7% up to 42.2% by model. We then read all 283 flagged spans. None was invented. 238 were the source channel's own Telegram markup inside the quoted sentence, which the model correctly dropped; the rest were ellipses, spliced fragments and 6 single-word slips. 11 of 14 configurations have a zero real-defect rate.
What the run cost, now that every number in it has been introduced.
What we did
We sampled 100 passages, matched to the 274,236 chunks of ≥50 tokens (92.5% of the 296,381-chunk corpus) of Russian official communications — Kremlin transcripts, Ministry of Defence and MFA channels, State Duma and Federation Council records — matched to the corpus on source arm, chunk-length quartile and time period. Each passage was put to 14 model configurations on their native APIs, twice, with one frozen prompt built from the project's current codebook. Each model returned a red-line verdict, a nuclear-signal verdict, a confidence score, and a verbatim quote from the passage supporting each call. 2,800 scored decisions, $22.38.
Why real statements rather than scenarios
The evals cited in the brief — Diplomacy self-play, CSIS's 400 constructed scenarios, WarAgent's counterfactual 1914, CivBench — all test models on invented situations. Ours tests them on things Russian officials actually said, in Russian, at a known date, from an identified channel. That buys three things a scenario cannot: the ambiguity is real rather than authored; the traps are real (Medvedev predicting nuclear escalation is not the same speech act as a red line, and models split on it); and a wrong answer maps onto a real analytical failure rather than a hypothetical one.
What actually separates the models
- What a naive verbatim check flags. 18.5% of cited spans, 1.7%–42.2% by model. Reading every one: 0 inventions. 238 were the channel's own markup inside the quote; the rest ellipses, splices and 6 one-word slips.
- Price predicts nothing. GPT-5.6 Luna costs $0.11 for the whole run and misses 2 nuclear signals; Claude Fable 5 costs $7.00 and misses 2.
- Recall and quote hygiene are unrelated. Claude Haiku 4.5 and GPT-5.6 Sol missed no nuclear signal at all, and their naive-flag rates differ 25-fold — but neither invented anything.
- Refusal is a failure mode. One provider's content filter declined 6 passages outright, including a genuine red-line statement by the Russian defence minister. Not a wrong answer — no answer.
- Confidence does not help. Models are barely less confident when wrong than when right.
Control arm — 0 false alarms in 700 decisions
Accuracy on a class-enriched set says nothing about how often a model cries wolf on ordinary traffic. So 50 passages were drawn at random from the corpus — material the screening pipeline classes as non-candidates — and put to all 14 configurations. Not one model raised a single red-line or nuclear alert on any of them. 700 decisions, 0 alerts.
This matters for the headline finding: the models are appropriately quiet on noise and accurate on signal — and a naive substring check flags up to 42.2% of their quotes — none of which, read individually, is an invention.
What we did not measure, and will not claim
- No language-robustness arm. We evaluated in Russian only. A verdict that flips between our translation and the original cannot be attributed to the model rather than the translation, so we did not run it. English on this site is a reading aid; no model ever saw it.
- Reference labels are provisional — single-adjudicator, coded before the project's codebook amendment, with no independent blind second pass returned. No inter-rater kappa is quoted anywhere.
- Accuracy differences are not significant. At 100 items the standard error is about ±3 points. We predicted before dispatch that the leaderboard would not separate the models, and it did not. We report it as a null result rather than a ranking.
- The sample is not prevalence-representative. Nuclear signals are far rarer in the corpus than in this sample; the classes are deliberately over-sampled so the rare class is measurable at all.
Everything here is reproducible
Every figure on this site is derived from the run's raw output by script — none is typed by hand. The frozen prompt, the sampler, the 100-item set with content hashes, all 2,800 per-decision records with rationales and evidence spans, the scorer and this page's generator are published together at github.com/hcss-utils/russian-redline-eval.
The leaderboard is not the finding
Every model scores between 0.890 and 0.945 on red-line accuracy and every 95% interval overlaps several others — at 100 items the standard error is ±3 points, so models within about 8 points are indistinguishable. This was predicted from the power arithmetic before the run was dispatched. What separates them is whether the quoted span survives a verbatim check: each model must cite a verbatim span from the passage, checked mechanically as a substring test. That rate runs from 1.7% to 42.2% — an 25-fold spread.
Reference labels are provisional — single-adjudicator, coded before Codebook Amendment 1, and no inter-rater kappa is quoted. Russian only: no translation-robustness arm was run, because a verdict flip between our translation and the original cannot be attributed to the model.
Core Leaderboard
| Rank | Model | Provider | Naive-flag rate | Missed nuclear | RLS acc [95% CI]i | NTS acc | Refusals | Consistencyi | Latencyi | Cost |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | OpenAI | 1.7% | 0/36 | 0.895 [0.84–0.93] | 1.000 | — | 0.990 | 6.3s | $1.67 |
| 2 | Claude Fable 5 | Anthropic | 2.6% | 2/36 | 0.915 [0.87–0.95] | 0.990 | — | 0.990 | 9.7s | $7.00 |
| 3 | Claude Opus 5 (thinking) | Anthropic | 3.3% | 3/35 | 0.910 [0.86–0.94] | 0.985 | — | 1.000 | 6.7s | $3.22 |
| 4 | Kimi K3 | Moonshot | 3.7% | 3/36 | 0.930 [0.89–0.96] | 0.985 | — | 0.980 | 18.4s | $0.61 |
| 5 | Claude Opus 5 (no thinking) | Anthropic | 5.2% | 4/36 | 0.934 [0.89–0.96] | 0.980 | — | 1.000 | 4.3s | $2.23 |
| 6 | GPT-5.6 Terra | OpenAI | 8.4% | 4/36 | 0.910 [0.86–0.94] | 0.980 | — | 0.980 | 3.4s | $0.73 |
| 7 | Gemini 3.6 Flash | 20.8% | 3/36 | 0.940 [0.90–0.96] | 0.985 | — | 1.000 | 6.0s | $0.94 | |
| 8 | GLM-5.2 | Zhipu AI | 23.6% | 4/35 | 0.919 [0.87–0.95] | 0.980 | — | 0.990 | 15.3s | $0.55 |
| 9 | Claude Sonnet 5 | Anthropic | 24.5% | 3/36 | 0.945 [0.90–0.97] | 0.985 | — | 0.990 | 5.3s | $1.81 |
| 10 | DeepSeek V4 Pro | DeepSeek | 28.0% | 8/36 | 0.910 [0.86–0.94] | 0.960 | — | 0.980 | 28.8s | $0.71 |
| 11 | Qwen3.7-Max | Alibaba | 28.6% | 4/34 | 0.942 [0.90–0.97] | 0.979 | 6 | 0.978 | 24.2s | $1.94 |
| 12 | DeepSeek V4 Flash | DeepSeek | 33.3% | 4/36 | 0.915 [0.87–0.95] | 0.975 | — | 0.990 | 14.2s | $0.21 |
| 13 | GPT-5.6 Luna | OpenAI | 37.3% | 2/36 | 0.890 [0.84–0.93] | 0.990 | — | 1.000 | 4.5s | $0.11 |
| 14 | Claude Haiku 4.5 | Anthropic | 42.2% | 0/36 | 0.900 [0.85–0.93] | 0.995 | — | 1.000 | 3.2s | $0.65 |
Ranking sensitivity
Rankings shown under equal miss/false-alarm weight. Toggle between 1:1, 3:1, and 10:1 miss-vs-false-alarm penalties. A model that looks best at 1:1 may drop under nuclear-miss-heavy weighting.
Head-to-Head Model Comparison
Select any number of models — two, five, or all 14. Every view below recomputes for the current selection.
Pattern
The failure modes are not symmetric, and that is the finding. Accuracy is flat — every model lands within a few points of the others — but faithfulness is not. GPT-5.6 Sol cites a quote that is actually in the passage 98.3% of the time; Claude Haiku 4.5 manages it only 57.8% of the time. The two are separated by 25×. Worse, the models that never miss a nuclear signal are not the faithful ones: Claude Haiku 4.5, GPT-5.6 Sol missed none at all, and Claude Haiku 4.5 is among them despite the highest naive-flag rate in the slate. Meanwhile DeepSeek V4 Pro missed 8 — the worst recall in the slate — without being the cheapest or the dearest.
Failure Atlas — Where Models Break
Worst-performing passages — every model, ranked by total errors · hover a row for the passage and reference label, click for all 14 justifications
- = correct. Red = model's wrong answer. Karaganov (#103) defeats ALL ten models — every one calls an analyst op-ed a nuclear signal. Medvedev (#198) and Lavrov (#037) are the top false-alert generators: retrospective and maximalist rhetoric that models cannot distinguish from live threats.
Operational risk surface
A missed nuclear signal is categorically worse than a missed conventional red line. This tab exposes exactly which passages each model misses, and whether the miss pattern is systematic (e.g. all models miss retrospective framing) or model-specific.
Case Explorer — Passage-Level Evidence
Hover any row for the passage, the reference label and why it was coded that way; click it for all 14 models' calls and their stated reasons. Every model appears as a column: ✓ means it matched the reference, otherwise the cell shows what it answered instead.
All 100 passages from the measured run are shown here, with each model's quoted evidence span and confidence score.
Situation Room — Play the Adviser
Read a passage cold. Make your own call. Then see what all 14 models said, and why.
Source Arms — Officialdom Corpus Breakdown
The 100 benchmark passages are drawn from four arms of the Russian Officialdom corpus (276,120 documents / 296,381 chunks / 36 source identities). Model performance varies sharply by source type.
Sample composition by source arm — measured
| Source arm | Passages | No alert | Red line | Nuclear signal | % of sample |
|---|---|---|---|---|---|
| telegram_official | 88 | 59 | 15 | 14 | 88% |
| kremlin | 5 | 1 | 1 | 3 | 5% |
| state_duma | 4 | 3 | 0 | 1 | 4% |
| federation_council | 3 | 3 | 0 | 0 | 3% |
Why source arm matters
Telegram Official passages are short (median 231 tokens) and context-poor — the model sees a single statement without the surrounding discussion. This is where most models perform best because the rhetoric is usually direct. Kremlin transcripts are longer and more nuanced — conditional language, retrospective framing, and diplomatic hedging live here. Duma/Federation Council material is longer-form than Telegram posts, and a benchmark passage is a chunk boundary rather than a whole document. Chunk length in this sample is matched to the corpus exactly by quartile, so length is not a confound between arms.
Benchmark passage distribution by source arm
| Source arm | Passages | No alert | Red line | Nuclear signal | % of sample |
|---|---|---|---|---|---|
| Telegram Official | 88 | 59 | 15 | 14 | 88.0% |
| Kremlin | 5 | 1 | 1 | 3 | 5.0% |
| State Duma | 4 | 3 | 0 | 1 | 4.0% |
| Federation Council | 3 | 3 | 0 | 0 | 3.0% |
| Total | 100 | 66 | 16 | 18 | 100% |
Error patterns by source type
| Failure type | Telegram | Kremlin | Duma/FC | Dominant cause |
|---|---|---|---|---|
| False NTS alert | 1 (0.0%) | 0 (0.0%) | 1 (0.5%) | Capability statements misread as threats; Medvedev bluster |
| False RLS alert | 52 (2.1%) | 27 (19.6%) | 15 (7.8%) | Retrospective/quoted doctrine; procedural "eskalatsiya" |
| Missed RLS | 23 (0.9%) | 0 (0.0%) | 0 (0.0%) | Veiled "consequences" without explicit boundary; denial framing |
| Missed NTS | 33 (1.3%) | 11 (8.0%) | 0 (0.0%) | Doctrinal boilerplate that IS a signal; conditional threats |
| Naive-flagged citation | 189 (7.7%) | 17 (12.3%) | 7 (3.6%) | Quoted span not found verbatim; reading all 283 found 0 invented |
Method — Exactly What We Do
Everything below is the protocol as executed, written so a reader can reproduce or attack it. The scored run is complete — 2,800 decisions for $22.38, plus a 50-passage corpus-random control arm. Every figure on this site is derived from the raw output; where something was not measured, such as translation robustness, it is labelled NOT MEASURED rather than estimated.
Status of this page
Reference labels: provisional. Single-adjudicator, and coded before Codebook Amendment 1 moved the construct from strict to inclusive. An independent blind second pass has not been returned, so no inter-rater kappa is quoted anywhere on this page. Results will move after adjudication.
1 · The question
Can a frontier model, acting as a situation-room adviser, distinguish a genuine Russian red line or nuclear signal from rhetoric that merely sounds like one — routine, retrospective, quoted, domestic, procedural or deliberately vague — without inventing a boundary or a consequence that the text does not contain?
This is deliberately not a scenario benchmark. CSIS's CFPD-Benchmark, WarAgent and the Diplomacy evaluations all test models on invented situations. RedLineBench tests them on real statements by real officials, where a ground truth about what was actually said is recoverable and the failure has a named operational cost.
2 · The three classes, and the rule that separates them
| Label | Decision rule | In the 100 |
|---|---|---|
| No alert | Anything that fails the red-line test: no boundary, or no consequence, or the speaker lacks authority, or the utterance is retrospective, quoted, hypothetical, procedural, or advocacy rather than statement of intent. | 66 |
| Conventional red line (RLS) | An authoritative speaker names, or unambiguously implies, (a) a prohibited action by an identifiable addressee and (b) an adverse consequence — where the consequence is not nuclear. | 16 |
| Nuclear signal (NTS) | As above, but the consequence carries an explicit or doctrinally unambiguous nuclear referent, from a speaker with employment authority. | 18 |
Three tests do most of the work, and they are where models fail. Speaker authority — Karaganov proposing first use is an analyst, not the state; Medvedev on Telegram holds no employment authority. Temporal framing — restating published doctrine is description, not signalling. Consequence class — "large-scale armed conflict" is conventional, however alarming the adjective attached to it.
3 · Corpus and sampling
Passages are drawn from the Russian Officialdom corpus: 276,120 documents / 296,381 chunks / 36 source identities, across four arms — Telegram Official (266,604 chunks), Kremlin (12,338), State Duma (10,553) and Federation Council (6,886). See the Source Arms tab for the per-arm breakdown.
The 100-item sample is matched to the corpus on source arm, chunk-length quartile and period, but deliberately over-samples the label classes, so it is not prevalence-representative and this must be stated wherever the headline number appears. Base rates of genuine signals in the raw corpus are far below 26%. Enrichment buys discriminating power per API dollar; the price is that absolute accuracy figures are not corpus base rates and must never be quoted as such. Selection combined (i) doctrinal-term retrieval, (ii) stratification across the four arms and across speakers, and (iii) deliberate inclusion of near-miss traps: retrospective, quoted, procedural, denial-framed, capability-only and advocacy items.
4 · Annotation and adjudication
Two domain experts code the set independently and blind. Disagreements go to written adjudication against the rule table above; the adjudicated rule text, not the individual votes, becomes the reference. Inter-annotator agreement is reported as Cohen's κ on the blind passes only. Because the completed pass to date is model-assisted, it is reported separately and excluded from κ. Dispute notes are published with the item set so a reader can see which items are contested and why.
5 · Prompt and output schema
One frozen prompt for every model, every language, every repetition. No model-specific tuning, no few-shot examples drawn from the item set. The model receives the passage, the speaker, the date and the outlet, and must return strict JSON:
{
"label": "none" | "conventional_red_line" | "nuclear_threat_signal",
"confidence": 0.0-1.0,
"evidence_span": "verbatim substring of the passage",
"trigger": "the prohibited action, or null",
"consequence": "the stated consequence, or null",
"speaker_authority": "operational" | "political" | "none",
"reasoning": "<= 60 words"
}
Requiring trigger and consequence as separate fields is doing analytical work, not bookkeeping: a model that returns a label but cannot name either one has pattern-matched vocabulary rather than reconstructed a boundary, and the paired-field null rate is reported alongside accuracy.
6 · Citation check — a screen, not a disqualifier
evidence_span is checked for verbatim presence in the source passage after Unicode normalisation. A span that is not found by that test is not automatically a fabrication. Reading all 283 flagged spans showed 0 inventions: 238 were the source channel's own markup inside the quoted sentence. The test is a screen, not a verdict. An adviser that reaches a correct conclusion by inventing a quote is not usable, and a benchmark that scores it as correct is measuring the wrong thing. Naive-flag rate is reported as a headline column, not a footnote.
7 · Translation protocol
Every passage is evaluated in Russian only — the language the corpus is actually written in. A translation-robustness arm was designed and then deliberately not run: a verdict that changes between our translation and the original cannot be attributed to the model rather than to the translation, and a number that cannot be attributed should not be reported. English text shown beside each passage on this site is a reading aid generated after the fact; no model ever saw it.
8 · Model manifest
14 configurations across 7 providers, all called on native APIs, never through an aggregating proxy — for open-weight models a proxy may silently route to a quantised or stale host, which would make the numbers unattributable. Two Claude Opus 5 entries differ only in whether extended thinking is enabled, holding every other variable constant.
9 · Run protocol
- 100 items × 1 language (Russian) × 2 repetitions × 14 models = 2,800 scored decisions, $22.38 measured spend against a $90 automatic stop. All models called on their native APIs — never through a router, because for open-weight models a router load-balances across third-party hosts and the number would measure a random host rather than the model. 6 records were refused outright by a provider content filter; 10 returned unparsable output and retain their raw text.
11 · Cost governance
Measured spend: $22.38 for 2,800 decisions, against a $90 automatic stop that halts dispatch rather than requesting approval. The stop was never approached. A pre-run pilot measured real prompt, completion and reasoning tokens on every provider before any commitment, because reasoning tokens are billed as output and are invisible in the response — an unmeasured projection for this run would have been roughly 2.5× too high.
12 · What this benchmark cannot tell you
- It measures classification of statements, not forecasting. A model that labels every passage perfectly has not predicted anything.
- Enriched sampling means accuracy figures are not corpus base rates.
- The reference is an expert reading of contested material. Where the two coders disagreed, the adjudicated rule is a defensible position, not a fact.
- Some items are recent enough to be inside training windows; a leakage audit by publication date is reported, but cannot be made airtight.
- Three repetitions bound self-consistency loosely; they do not characterise the full sampling distribution.
13 · Reproducibility
One command rebuilds every public metric and case file from frozen artifacts. The deposited package contains: the frozen prompt and JSON schema; the 100-item set with per-item content hashes; the frozen English translations; the model manifest with dated slugs and resolved provider versions; raw and normalised outputs; the scoring code; the adjudication notes; and the cost ledger.
Downloads
The full record — prompt, sample, per-decision output, scoring code — is published with the submission.
