← back to dashboard  ·  method  ·  Türkçe

What we measure, and how

This page is not marketing — it is the product. For a closed score to be trusted in this field you need either a BlackRock-sized brand or an auditable methodology. The second is within our reach.

The core claim

An official gazette is not what a state says, but what it does. It is published before it reaches the news, its date is exact, and it is binding. A speech shows intent; a regulation shows a committed resource. This system counts the second kind.

Role Gap  Δᵢ = Kᵢ − Rᵢ Capacity minus assigned role. Δ>0 revisionist pressure · Δ<0 proxy fatigue.

The binary choice was dropped: separate K and R per document

The first schema forced K/R as a binary: a document was either a capacity signal or a role signal. That was the model's most contestable assumption, and it was wrong. A “basing agreement” is both — it grows the host state's actual power and imposes an obligation on it. Under a binary schema one of those two opposed readings disappears.

each document   k ∈ [−5,+5]  ·  r ∈ [−5,+5]  ·  Δ contribution = k − r A basing agreement can now be k=+3, r=+4 → Δ contribution −1: it confers power but commits more.

Δ flow is a flow, not a stock. The Δ = K − R in the thesis layer is a stock: an actor's accumulated position, assigned by hand. The Δ measured on the dashboard is the net pressure produced within a window. They are not the same unit and sit in separate boxes.

Only their signs are compared. If the flow's sign matches the stock assumption it reads “supports”; if it is opposite, “contradicts”. This is exactly where the thesis is tested — a contradiction says either the assumption is wrong or the trend is turning.

Δ is hidden when label coverage is low. Δ is computed only from labelled documents; an unlabelled document contributes zero. At low coverage this does not merely make Δ incomplete — it makes it incomparable across time. If August is 100% labelled and March 20%, the series shows a fake August spike, precisely where fracture detection looks.

So below 85% coverage the figure is not shown as “small” — it is treated as invalid. The volume layer is unaffected: it needs no labels and is always complete, which is why fracture detection works regardless of labelling progress.

But coverage alone is not enough. Measured: in the United Kingdom's 90-day window, 16 of 16 documents were labelled — 100% coverage — yet only 9 of them were relevant. 100% coverage does not turn nine documents into a hundred. Δ now requires a second, independent condition: at least 12 relevant documents in the window.

And the two together are still not enough — a third was added on 24 August 2026. Both conditions above are built on “the documents we hold” and neither can see a document that never arrived. An unlabelled document at least appears in the denominator; a document never fetched appears nowhere — so coverage reads 100%, because a document that does not exist is not counted as missing.

Measured: the Turkish source failed DNS resolution three nights running, and because the daily loop only looks two days back, one day was never asked for again. Turkey's 90-day Δ was computed from 83 days rather than 90, while label coverage read 97.9%: the figure looked sound. Δ is now also invalid when ingestion coverage within the window falls below 95%.

The reason for invalidity is diagnosed from the outside in — ingestion gap first, then label coverage, then document count. Saying “coverage is low” about an incomplete archive points the reader at the wrong thing: what needs fixing is not the labelling but going back to the source for that day. A wrong diagnosis is worse than none.

The referee — Δ's error margin

The rest of the system measures states; this layer measures the instrument. It exists because of one finding: of 99 relevant labels, 49 carry r=0 and only 20 carry k=0. The labeller assigned zero on the role axis to half the documents, and on the capacity axis to far fewer. Since Δ = k − r, this is a systematic bias that makes every state look like a capacity builder — and indeed all four countries came out with positive Δ.

This is not a model error but a behaviour: “capacity” is something visible (an authority, a budget, a permit), whereas “role” is a theory-laden judgement. Writing zero on the axis you are unsure about is a reasonable reflex — but done 99 times in a row it produces an artefact, not a measurement. A human cannot do this; they get tired and ask “am I always writing zero in this field?” A machine does not ask.

On 20 August 2026 the same 40 documents were labelled blind twice more: once by the same model (test–retest), once by a different model (inter-rater). The sample is seeded, so all three rounds cover exactly the same documents.

κ (quadratic weighted)test–retestinter-rater
relevance+0.754+0.441
field+0.847+0.908
k (capacity)+0.773+0.779
r (role)+0.902+0.798
k drift (per document)−0.45−0.44

The expectation was that the role axis would come out weak, being a theory-laden judgement. Measurement said the opposite — r is more reliable than k in both rounds.

But this does not solve the r=0 problem — it changes its meaning. Across all four measurements, 55–60% of documents received r=0, and two different models write zero on the same documents. So r=0 is not noise but a stable default.

A caveat: two instruments from the same model family reading the same instruction text is not independent evidence. Whether the zeros come from the documents or from the definition of r can only be settled by a human round. Reliability is not validity: an instrument that makes the same error consistently scores high.

On the capacity axis both rounds scored 0.45 points lower per document than the base round. Two independent instruments drifting the same way suggests the problem lies not in the noise of a round but in the base round itself. That is not scatter but directional drift, and across a 45-document window its total exceeds Δ itself. So the two failure modes are tested separately:

band (random) = Δ recomputed 400 times with drift-removed disagreement → a symmetric 95% interval around the point estimate
drift (directional) = where Δ lands when the mean per-document difference is applied Δ's sign must survive both. If it does not, neither the magnitude nor the direction is reported, and no thesis test is run in that window.

Combining the two into one band was the first version's flaw, and measurement exposed it: Turkey's 90-day band came out as [−18.33 … −5.45] — the band excluded its own point estimate (+2.41). The figure was not wrong, its meaning was: the band was asserting “Δ's true value is −12”, and there is no basis for that claim; we do not know which round is right. Separating them is more conservative: a narrow band plus a large drift means the instrument is consistently unstable, and a reading that looks only at the band would call it sound. That is the most misleading case of all.

The band is not produced by a second implementation of Δ, but by calling the production aggregation function again with perturbed labels. A separate formula would silently drift from the very figure it claims to measure — the same decision as in the time machine.

The weakest link is not the scores but the relevance decision. With a different labeller, 8 of the 33 documents that entered Δ in the base round dropped out entirely (24%). That is a far larger lever than ±1 point shifts: a shifted score moves Δ by one unit, a dropped document erases its entire contribution. In the first version this sat on the “not modelled” list; a measured but unmodelled figure is more dangerous than an unmeasured one, because it is assumed to have been accounted for. It is now inside the band.

The band is fed by the inter-rater round when one exists; the test–retest round is a known lower bound and serves only as a fallback, labelled as such in the dashboard.

The third gate — criterion information. Band and drift were not enough; on 21 August 2026 the EU corpus showed why: when both rounds wrote r=+1 on almost every document, κ on the role axis collapsed to −0.00, the band NARROWED because the rounds "agreed", and the figure looked solid. A narrow band there is not evidence — it is the criterion erasing variance. So Δ's sign must pass three gates: the band must not cross zero, the drift must not flip the sign, and the round feeding the band must have κ ≥ 0.40 on the Δ axis. The three are AND-ed; if one fails, the figure is not reported.

The band is measured on the corpus it claims to measure. Applying disagreement measured on one source to another is a silent assumption — the EU's largest figure was once withdrawn for exactly that. Each country's band is now fed by a blind round run on that country's own corpus where one exists (EU, US, Türkiye and Canada); otherwise it is fed by the general round and labelled "general" in the dashboard. The published JSON carries, for every band, which round it came from and how old that round is.

What is not modelled (the irrelevant→relevant direction, disagreement over field assignment, documents the filter never surfaced, and systematic bias shared by both labellers) widens the band further, so the published band is a lower bound on reality.

The human anchor — what is still missing

Every κ value and every band above comes from model rounds. Ten rounds have been run (four of them on the EU, US, Türkiye and Canada corpora themselves) and all ten are from the same model family. That makes the last item on the “not modelled” list something other than a footnote: no model round can see a bias the family shares.

What was measured: all four rounds scored the capacity axis lower than the base round (−0.29 … −0.46 per document), and in all four the share of r=0 was 46–59%. That means one of two things — either the base round scores capacity too high, or all four rounds make the same mistake. No model round can tell those apart, because the instrument doing the measuring belongs to the same family.

The same gap runs through calibration: 165 of the gold set's 165 labels are model-generated. So the published precision and recall measure how well the filter matches the model, not the truth. This page has said so since the system was built.

On 24 August 2026 the infrastructure to close this gap was built: an offline worksheet showing the same 40 documents blind, with full text and the scoring ladder alongside. Once the round is run it becomes the band's primary source — a model round measures reliability, a human round measures validity for the first time. The round has not been run yet, and until it is, this remains a gap rather than a plan.

28 August 2026 — the owner's decision: the system is run entirely by AI and the human round is not being run. The slot was left open deliberately, not closed: the worksheet exists and the day it is filled in, the band switches to it automatically. Until then every band on this page comes from a model round and is labelled as such — model output is never written up as a human measurement.

One component of the shared bias has been measured

“Bias shared by both labellers” is not one thing but two: both come from the same model family and both read the same scoring ladder. The second is measurable — and on 25 August 2026 it was measured: the same 40 documents, the same different model, but with the ladder withheld, using only the definitions of k and r.

with ladderladder-free
share of r=0 (repeat round)54%91%
relevance κ+0.353+0.771
relevance drop-out31%3%
Δ = k − r · κ+0.852+0.685

Most of the published role signal is the rulebook itself. Without the ladder the model leaves the role axis empty in 37 of 40 documents. The “r=0 pile-up” this page has flagged from the start appears here in its pure form: the pile-up is the model's natural behaviour and the ladder partly corrects it. So the measured agreement on r is not two independent judgements converging — it is largely both of them applying the same rule.

But the ladder makes relevance agreement worse. Its explicit “irrelevant” lists give the second labeller grounds to say “this document is not our subject”, dropping documents the base round kept. This was previously attributed to the body text becoming visible; it has now been isolated — the cause is not the text but the ladder.

No published figure changed as a result, and the band continues to be fed by the with-ladder round. That is correct: the band must model the variability of a labeller following the published method, and the method includes the ladder. What changed is how much of a lower bound the band is — “a lower bound on reality” now carries a number: for two labellers who do not share the method, Δ agreement is not 0.852 but 0.685. The ladder-free round does not replace the human anchor: it measures the rulebook's share, not the model's.

Calibration now reports the human and model references separately. The reason is forward-looking: the moment the first human label enters the set, a single combined figure becomes an average of two references and the reader can no longer tell which one they are looking at. Worse, while human labels remain a minority the combined figure is almost entirely the model figure — creating the impression that a human reference now exists. In a system that measures its own error margin, hiding whose judgement the reference is is the most expensive thing to lose.

Scoring ladders — published copy (28 August 2026)

The single source of the criterion the labeller reads is the agent definition; it is summarised here for the reader. Severity runs 1 (administrative, routine) … 5 (regime-changing). k moves actual power, r moves the assigned or assumed role; both can be given on the same document and negative values are real.

corpusdocument typekr
US
Federal Register
AD/CVD final determination, dumping/subsidy found+2+1
initiation of investigation/review · preliminary result+1+1
no dumping · postponement · scope list · schedule notice (NOT irrelevant)0+1
OFAC/State SDN designation · UNSC 1267 implementation0+3
232/301 tariff proclamation · relaxation of export controls+3…+4 · +20…+2 · −2
EU
EUR-Lex
CFSP restrictive measure — listing · extension · delisting0+3 · +2 · −2
sanctions notice: body says listing / extension / identity correction / data-protection notice0+2 / +1 / 0 / irrelevant
state aid: "no objections" / 108(2) investigation / negative decision — k only for defence (+2), nuclear-grid-semiconductor (+1)0…+2+1 / +2 / +3
defence-industry programme · energy security of supply · critical raw materials+3…+4 · +30…+1 · 0
signature/ratification of international agreement · preparatory acts (one rung down) · CJEU judgment+1 · ≤1 · 0+3 · ≤1 · +1
Canada
Gazette Part II
Special Economic Measures / UN Act listing · delisting · regime extension0 · 0 · 0…+1+3 · −2 · +3
Criminal Code List of Entities addition · annual review0+3 · +1
surtax / safeguard · surtax remission+2 · −1…0+1 · +1
Export/Import Control List · Nuclear Safety and Control Act import-export+1+2
Türkiye
Resmî Gazete
UNSC 1267 asset freeze · lifting of a freeze0+3 · −2
parliamentary Article 92 motion (cross-border deployment) · UN-mandated participation extension+4 · +2+3 · +3
ratification of international agreement · import regime/safeguard (when the body is readable)0 · +2+2 · +1

Recognising the pattern and writing the score from memory is forbidden. If the distinguishing field cannot be found in the body, the default is not applied: a low severity is given, the rationale says so, and the document goes to the calibration queue. Adopted as principles: establishing or repealing a whole regime sits one rung above a single listing (Iran snapback r=+4, lifting of the Syria regime r=−3). Every edge case not on the ladder is kept in a dated list in the agent definition and counts as open until it becomes a rule.

The pipeline

  1. Ingest — daily cron, raw documents from the source. Deduplicated by hash, written to a permanent archive.
  2. Normalise — HTML/PDF → plain text. Scanned PDFs go through OCR; confidence is stored.
  3. Rule filter — weighted key terms. A free layer; it removes the obviously irrelevant.
  4. Labelling — per document: field, orientation, k and r scores, target actor, rationale. Every label is stored with its author; click a document in the panel to see its rationale.
  5. Aggregation — no model here. Code produces the score, by a deterministic formula.
  6. Publish — daily static JSON. No server-side computation.
  7. Commentary — one model call a day; its only input is the set of figures that passed the referee. Every number and country name in the text is checked by code against that input, forecasting language is forbidden, and commentary that fails the gate is not published. With no reportable window the model is not called at all — silence is a legitimate output.
  8. Referee — rounds that relabel the same documents blind; Δ's error margin is measured here and every figure is published together with its band (below).

The archive rule

Fetch a source once, never fetch it again. The raw document, its download timestamp and its hash are kept permanently. The reason: when the taxonomy changes the archive is reprocessed, the source is not refetched. If a source site deletes its history or moves behind a paywall — as happened with Japan's Kanpō archive — this is all that remains. The archive is this project's real asset, not the model output.

Role archetypes

Country colours on the globe are role archetypes: hegemon, revisionist, rising, bloc member, proxy, double-bound, isolated, buffer, periphery. If the thesis claims that states are positioned by their assigned role rather than by their power, then that is what the map should carry.

These assignments are assumptions, not measurements. The measured layer is separate and looks different on the globe: amber pillars (decision volume) and targeting links.

A single assignment is a simplification. India is both rising and double-bound, Iran both revisionist and isolated, Ukraine both buffer and proxy. Only the dominant one is shown — the same limit the binary K/R choice carried, and it must be stated as plainly.

Two images are produced: one to be looked at (archetype colours) and one to be read (each country filled with a unique code colour, no antialiasing). Clicking resolves the country by reading a pixel at the hit point's UV — running a point-in-polygon test in the browser across 177 countries would be both slow and wrong at the antimeridian.

Sources

countrysourceclasslicence
United StatesFederal Register (JSON API)APublic domain — 17 U.S.C. §105
TürkiyeResmî Gazete (HTML/PDF)CUnclear — legal opinion needed before republication
United Kingdomlegislation.gov.uk (Atom)AOpen Government Licence v3.0
PolandDziennik Ustaw (ELI API)AOfficial legal text — not copyrightable
European UnionEUR-Lex / CELLAR (SPARQL) — Union-level acts onlyACommission reuse decision 2011/833/EU
CanadaCanada Gazette, Part II (SOR/SI, fortnightly HTML)BReproduction of Federal Law Order (SOR/97-5)

Class A: official API or bulk download · B: regular structure, no API · C: scraping + PDF/OCR · D: restricted access. Turkish content is archived and enters the measurement, but its full text is not republished here; only the title, source link and derived score are shown. For the EU, member-state acts, consolidated texts and duplicate case-law records are out of scope; for Canada only Part II (registered regulations) is taken, not Part I (notices and proposals) — the system counts what is done.

Why not The Gazette

The Gazette was tried first for the UK, and measured: of the 516 notices published on 19 August 2026, 484 were corporate insolvencies, personal insolvencies and probate notices. The wrong source for measuring state behaviour. The real counterpart of the Federal Register is the statutory instrument stream — legislation.gov.uk.

A second trap found there: the feed returns 20 records per page and hides the rest behind a rel="next" link. Without pagination, every day with more than 20 items silently lost the remainder — more insidious than returning nothing, because partial data looks entirely normal. It was caught by noticing that the daily maximum was exactly 20 on nine separate days: a number repeating at the top of a distribution is a ceiling, not data.

The rule filter

This layer's job is not to decide correctly. Its job is to keep the obviously irrelevant half of the hundreds of daily documents away from the expensive layer. Its threshold is therefore deliberately loose: a wrong elimination is expensive, a wrong pass costs a few cents.

score = Σ(positive term weight) − Σ(negative term weight) A term in the title carries full weight, in the body 45%. Threshold 1.0; if no positive term appears in the title the threshold rises to 2.2. Every matched term is recorded — which decision was made and why stays auditable.

Negative terms do not eliminate a document, they lower its score. A strong positive term can beat a negative one.

The raised body threshold was measured, not guessed: in the first version only 39% of passing documents were genuinely relevant, and the dominant failure was documents whose body mentioned a topic term while the title was about something else entirely — a car-rental regulation passed because the word “sanction” appeared in its text. Gazette titles are deliberately descriptive.

Calibration — the one place the system can audit itself

There is no ground truth; the hand-labelled gold set is the only truth we have. The filter is measured against it at every version. The two error types are not equal: a miss is expensive (that document never reaches the labelling layer), a false pass is cheap.

versionprecisionrecallwhat changed
r139%100%first version — passes everything
r290%56%negative list + body threshold; cut too much
r389%95%gaps in the positive list closed

r2's collapse was not caused by the threshold but by omissions in the positive list: antidumping existed only in its hyphenated form (anti-dumping) and the Federal Register does not hyphenate it. A single hyphen missed ten documents. Without the gold set this would have been invisible — which is the whole point of having one.

Live calibration figures are in the dashboard's system tab, alongside the values actually in production.

OCR — and a silent error we measured

In Türkiye, presidential decisions, international treaties and board rulings are published as scanned PDFs. Their content is an image; they appear to have no text layer. Our first threshold was simple: “a PDF yielding fewer than 180 characters is scanned.” It was wrong.

We measured the distribution across 83 PDFs and it was bimodal: Poland's genuine text PDFs yield 3,400–4,200 characters per page, Türkiye's scans 4–206. Nothing in between.

206 is not a coincidence. A scanned page does carry a text layer — but it is not content, it is the masthead: “13 August 2026 THURSDAY · Resmî Gazete · No: 33339”.

A total-character threshold mistook that boilerplate for content. A 22-page scanned court ruling counted as “has text” on the strength of 459 characters and never entered the OCR queue. Silently. Of 83 PDFs, 49 were effectively scanned while only 37 were flagged.

is it scanned = (extracted characters ÷ page count) < 400 Measured per page, not in total. A wrong OCR is cheap (a few seconds of processing); a missed OCR silently becomes a wrong label.

Whose confidence is the confidence?

Human review was initially triggered by the document average. Looking at the output showed that this was wrong: the Türkiye–Saudi Arabia visa-exemption agreement runs to 17 pages with an average confidence of 58.8. But the decision text is on the first page and is clean — the low score comes from the maps and coordinate tables on the following 16 pages.

human review = first-page confidence < 70 Not the document average. In a gazette the decision text is always on the first page; failing to read the annexes is not failing to read the decision.

OCR output does not alter the raw archive; it is derived data, stored separately with its confidence. If the engine or language changes it is regenerated without returning to the source. The Turkish language pack is mandatory: without ğ, ş, ı, İ, ö, ü, ç the output silently breaks keyword matching.

Aggregation

intensity(country, window) = Σ [ score × field weight × 0.5(age/90) ] 90-day half-life: a new decision outweighs an old one, but the old one is not zeroed.
fracture = (7-day average ÷ 30-day average) ≥ 1.8  and  ≥ 3 documents in 7 days Not a prediction — a deviation measurement: “more decisions than usual are being produced here.” It does not say why.

Windows are divided by the number of days available, not by their nominal length, and fracture is not computed until the archive is at least 30 days deep. The earlier 21-day rule produced an artefact visible in the time series: Turkish fractures began exactly on the archive's 21st day, because a 30-day baseline computed from 21 days of data and still divided by 30 is roughly 30% too low, inflating the ratio by about 43%.

Against silent breakage

Gazette sites change structure without notice. When the ingest layer breaks it must not silently return empty: “zero documents today” is an alarm, not a normal day.

But some sources genuinely do not publish. The Federal Register does not publish at weekends or on federal holidays; legislation.gov.uk produces 0–9 documents a day. If we cannot tell the two apart we either miss real failures or get false alarms twice a week and stop reading alarms — the second being more dangerous. The calendar is therefore source-specific.

For Türkiye no fixed holiday calendar is hard-coded. On religious holidays the index page returns HTTP 200 with a notice instead of documents: “pursuant to Presidential Decree No. 10, the Official Gazette is not published today.” That notice is read directly — if the calendar changes the code need not, and “not published” never gets confused with “the parser broke”.

Known limits

  1. The gold set was also produced by a model. The first 92 labels were assigned document by document, by hand — but the hand was a language model's. Better than no reference; not the same as a human-verified one. The real job of weekly human calibration is to correct this set, not to enlarge it.
  2. A labeller cannot grade its own measure. The daily agent may not add its output to the gold set; if it did, the system would confirm its own error as reference and drift silently.
  3. The K/R distinction is still interpretive — the binary was removed, but whether a document deserves k=+3 or k=+4 is judgement. Consistency of the severity scale matters more than individual accuracy: Δ is a time series, and a drifting criterion corrupts it.
  4. Ease of access is inversely correlated with thesis relevance. The actors that build the containment triangles — China, Russia, India, Pakistan — are all in the hard-access class. Taking the easy path would build a Western-centric index while the thesis's actual claim stays unmeasured. Source class is therefore visible on every country card.
  5. Unsupervised loops amplify their own errors. Hence the target is not “fully autonomous”: daily automatic production plus weekly human-approved calibration.

Positioning

The claim is not “I measure the world correctly.” The claim is: I convert state behaviour into a traceable unit, and I show exactly how I convert it.

The taxonomy, weights and decay coefficient on this page are the values actually used in production — not a separate marketing text. Back to dashboard · thesis model · Türkçe