This page is not marketing — it is the product. For a closed score to be trusted in this field you need either a BlackRock-sized brand or an auditable methodology. The second is within our reach.
An official gazette is not what a state says, but what it does. It is published before it reaches the news, its date is exact, and it is binding. A speech shows intent; a regulation shows a committed resource. This system counts the second kind.
The first schema forced K/R as a binary: a document was either a capacity signal or a role signal. That was the model's most contestable assumption, and it was wrong. A “basing agreement” is both — it grows the host state's actual power and imposes an obligation on it. Under a binary schema one of those two opposed readings disappears.
Δ flow is a flow, not a stock. The Δ = K − R in the thesis layer is a stock: an actor's accumulated position, assigned by hand. The Δ measured on the dashboard is the net pressure produced within a window. They are not the same unit and sit in separate boxes.
Only their signs are compared. If the flow's sign matches the stock assumption it reads “supports”; if it is opposite, “contradicts”. This is exactly where the thesis is tested — a contradiction says either the assumption is wrong or the trend is turning.
Δ is hidden when label coverage is low. Δ is computed only from labelled documents; an unlabelled document contributes zero. At low coverage this does not merely make Δ incomplete — it makes it incomparable across time. If August is 100% labelled and March 20%, the series shows a fake August spike, precisely where fracture detection looks.
So below 85% coverage the figure is not shown as “small” — it is treated as invalid. The volume layer is unaffected: it needs no labels and is always complete, which is why fracture detection works regardless of labelling progress.
But coverage alone is not enough. Measured: in the United Kingdom's 90-day window, 16 of 16 documents were labelled — 100% coverage — yet only 9 of them were relevant. 100% coverage does not turn nine documents into a hundred. Δ now requires a second, independent condition: at least 12 relevant documents in the window.
And the two together are still not enough — a third was added on 24 August 2026. Both conditions above are built on “the documents we hold” and neither can see a document that never arrived. An unlabelled document at least appears in the denominator; a document never fetched appears nowhere — so coverage reads 100%, because a document that does not exist is not counted as missing.
Measured: the Turkish source failed DNS resolution three nights running, and because the daily loop only looks two days back, one day was never asked for again. Turkey's 90-day Δ was computed from 83 days rather than 90, while label coverage read 97.9%: the figure looked sound. Δ is now also invalid when ingestion coverage within the window falls below 95%.
The reason for invalidity is diagnosed from the outside in — ingestion gap first, then label coverage, then document count. Saying “coverage is low” about an incomplete archive points the reader at the wrong thing: what needs fixing is not the labelling but going back to the source for that day. A wrong diagnosis is worse than none.
The rest of the system measures states; this layer measures the instrument. It exists because of one finding: of 99 relevant labels, 49 carry r=0 and only 20 carry k=0. The labeller assigned zero on the role axis to half the documents, and on the capacity axis to far fewer. Since Δ = k − r, this is a systematic bias that makes every state look like a capacity builder — and indeed all four countries came out with positive Δ.
This is not a model error but a behaviour: “capacity” is something visible (an authority, a budget, a permit), whereas “role” is a theory-laden judgement. Writing zero on the axis you are unsure about is a reasonable reflex — but done 99 times in a row it produces an artefact, not a measurement. A human cannot do this; they get tired and ask “am I always writing zero in this field?” A machine does not ask.
On 20 August 2026 the same 40 documents were labelled blind twice more: once by the same model (test–retest), once by a different model (inter-rater). The sample is seeded, so all three rounds cover exactly the same documents.
| κ (quadratic weighted) | test–retest | inter-rater |
|---|---|---|
| relevance | +0.754 | +0.441 |
| field | +0.847 | +0.908 |
| k (capacity) | +0.773 | +0.779 |
| r (role) | +0.902 | +0.798 |
| k drift (per document) | −0.45 | −0.44 |
The expectation was that the role axis would come out weak, being a theory-laden judgement. Measurement said the opposite — r is more reliable than k in both rounds.
But this does not solve the r=0 problem — it changes its meaning. Across all four measurements, 55–60% of documents received r=0, and two different models write zero on the same documents. So r=0 is not noise but a stable default.
A caveat: two instruments from the same model family reading the same instruction text is not independent evidence. Whether the zeros come from the documents or from the definition of r can only be settled by a human round. Reliability is not validity: an instrument that makes the same error consistently scores high.
On the capacity axis both rounds scored 0.45 points lower per document than the base round. Two independent instruments drifting the same way suggests the problem lies not in the noise of a round but in the base round itself. That is not scatter but directional drift, and across a 45-document window its total exceeds Δ itself. So the two failure modes are tested separately:
Combining the two into one band was the first version's flaw, and measurement exposed it: Turkey's 90-day band came out as [−18.33 … −5.45] — the band excluded its own point estimate (+2.41). The figure was not wrong, its meaning was: the band was asserting “Δ's true value is −12”, and there is no basis for that claim; we do not know which round is right. Separating them is more conservative: a narrow band plus a large drift means the instrument is consistently unstable, and a reading that looks only at the band would call it sound. That is the most misleading case of all.
The band is not produced by a second implementation of Δ, but by calling the production aggregation function again with perturbed labels. A separate formula would silently drift from the very figure it claims to measure — the same decision as in the time machine.
The weakest link is not the scores but the relevance decision. With a different labeller, 8 of the 33 documents that entered Δ in the base round dropped out entirely (24%). That is a far larger lever than ±1 point shifts: a shifted score moves Δ by one unit, a dropped document erases its entire contribution. In the first version this sat on the “not modelled” list; a measured but unmodelled figure is more dangerous than an unmeasured one, because it is assumed to have been accounted for. It is now inside the band.
The band is fed by the inter-rater round when one exists; the test–retest round is a known lower bound and serves only as a fallback, labelled as such in the dashboard.
The third gate — criterion information. Band and drift were not enough; on 21 August 2026 the EU corpus showed why: when both rounds wrote r=+1 on almost every document, κ on the role axis collapsed to −0.00, the band NARROWED because the rounds "agreed", and the figure looked solid. A narrow band there is not evidence — it is the criterion erasing variance. So Δ's sign must pass three gates: the band must not cross zero, the drift must not flip the sign, and the round feeding the band must have κ ≥ 0.40 on the Δ axis. The three are AND-ed; if one fails, the figure is not reported.
The band is measured on the corpus it claims to measure. Applying disagreement measured on one source to another is a silent assumption — the EU's largest figure was once withdrawn for exactly that. Each country's band is now fed by a blind round run on that country's own corpus where one exists (EU, US, Türkiye and Canada); otherwise it is fed by the general round and labelled "general" in the dashboard. The published JSON carries, for every band, which round it came from and how old that round is.
What is not modelled (the irrelevant→relevant direction, disagreement over field assignment, documents the filter never surfaced, and systematic bias shared by both labellers) widens the band further, so the published band is a lower bound on reality.
Every κ value and every band above comes from model rounds. Ten rounds have been run (four of them on the EU, US, Türkiye and Canada corpora themselves) and all ten are from the same model family. That makes the last item on the “not modelled” list something other than a footnote: no model round can see a bias the family shares.
What was measured: all four rounds scored the capacity axis lower than the base round (−0.29 … −0.46 per document), and in all four the share of r=0 was 46–59%. That means one of two things — either the base round scores capacity too high, or all four rounds make the same mistake. No model round can tell those apart, because the instrument doing the measuring belongs to the same family.
The same gap runs through calibration: 165 of the gold set's 165 labels are model-generated. So the published precision and recall measure how well the filter matches the model, not the truth. This page has said so since the system was built.
28 August 2026 — the owner's decision: the system is run entirely by AI and the human round is not being run. The slot was left open deliberately, not closed: the worksheet exists and the day it is filled in, the band switches to it automatically. Until then every band on this page comes from a model round and is labelled as such — model output is never written up as a human measurement.
“Bias shared by both labellers” is not one thing but two: both come from the same model family and both read the same scoring ladder. The second is measurable — and on 25 August 2026 it was measured: the same 40 documents, the same different model, but with the ladder withheld, using only the definitions of k and r.
| with ladder | ladder-free | |
|---|---|---|
| share of r=0 (repeat round) | 54% | 91% |
| relevance κ | +0.353 | +0.771 |
| relevance drop-out | 31% | 3% |
| Δ = k − r · κ | +0.852 | +0.685 |
Most of the published role signal is the rulebook itself. Without the ladder the model leaves the role axis empty in 37 of 40 documents. The “r=0 pile-up” this page has flagged from the start appears here in its pure form: the pile-up is the model's natural behaviour and the ladder partly corrects it. So the measured agreement on r is not two independent judgements converging — it is largely both of them applying the same rule.
But the ladder makes relevance agreement worse. Its explicit “irrelevant” lists give the second labeller grounds to say “this document is not our subject”, dropping documents the base round kept. This was previously attributed to the body text becoming visible; it has now been isolated — the cause is not the text but the ladder.
No published figure changed as a result, and the band continues to be fed by the with-ladder round. That is correct: the band must model the variability of a labeller following the published method, and the method includes the ladder. What changed is how much of a lower bound the band is — “a lower bound on reality” now carries a number: for two labellers who do not share the method, Δ agreement is not 0.852 but 0.685. The ladder-free round does not replace the human anchor: it measures the rulebook's share, not the model's.
Calibration now reports the human and model references separately. The reason is forward-looking: the moment the first human label enters the set, a single combined figure becomes an average of two references and the reader can no longer tell which one they are looking at. Worse, while human labels remain a minority the combined figure is almost entirely the model figure — creating the impression that a human reference now exists. In a system that measures its own error margin, hiding whose judgement the reference is is the most expensive thing to lose.
The single source of the criterion the labeller reads is the agent definition; it is summarised here for the reader. Severity runs 1 (administrative, routine) … 5 (regime-changing). k moves actual power, r moves the assigned or assumed role; both can be given on the same document and negative values are real.
| corpus | document type | k | r |
|---|---|---|---|
| US Federal Register | AD/CVD final determination, dumping/subsidy found | +2 | +1 |
| initiation of investigation/review · preliminary result | +1 | +1 | |
| no dumping · postponement · scope list · schedule notice (NOT irrelevant) | 0 | +1 | |
| OFAC/State SDN designation · UNSC 1267 implementation | 0 | +3 | |
| 232/301 tariff proclamation · relaxation of export controls | +3…+4 · +2 | 0…+2 · −2 | |
| EU EUR-Lex | CFSP restrictive measure — listing · extension · delisting | 0 | +3 · +2 · −2 |
| sanctions notice: body says listing / extension / identity correction / data-protection notice | 0 | +2 / +1 / 0 / irrelevant | |
| state aid: "no objections" / 108(2) investigation / negative decision — k only for defence (+2), nuclear-grid-semiconductor (+1) | 0…+2 | +1 / +2 / +3 | |
| defence-industry programme · energy security of supply · critical raw materials | +3…+4 · +3 | 0…+1 · 0 | |
| signature/ratification of international agreement · preparatory acts (one rung down) · CJEU judgment | +1 · ≤1 · 0 | +3 · ≤1 · +1 | |
| Canada Gazette Part II | Special Economic Measures / UN Act listing · delisting · regime extension | 0 · 0 · 0…+1 | +3 · −2 · +3 |
| Criminal Code List of Entities addition · annual review | 0 | +3 · +1 | |
| surtax / safeguard · surtax remission | +2 · −1…0 | +1 · +1 | |
| Export/Import Control List · Nuclear Safety and Control Act import-export | +1 | +2 | |
| Türkiye Resmî Gazete | UNSC 1267 asset freeze · lifting of a freeze | 0 | +3 · −2 |
| parliamentary Article 92 motion (cross-border deployment) · UN-mandated participation extension | +4 · +2 | +3 · +3 | |
| ratification of international agreement · import regime/safeguard (when the body is readable) | 0 · +2 | +2 · +1 |
Recognising the pattern and writing the score from memory is forbidden. If the distinguishing field cannot be found in the body, the default is not applied: a low severity is given, the rationale says so, and the document goes to the calibration queue. Adopted as principles: establishing or repealing a whole regime sits one rung above a single listing (Iran snapback r=+4, lifting of the Syria regime r=−3). Every edge case not on the ladder is kept in a dated list in the agent definition and counts as open until it becomes a rule.
Fetch a source once, never fetch it again. The raw document, its download timestamp and its hash are kept permanently. The reason: when the taxonomy changes the archive is reprocessed, the source is not refetched. If a source site deletes its history or moves behind a paywall — as happened with Japan's Kanpō archive — this is all that remains. The archive is this project's real asset, not the model output.
Country colours on the globe are role archetypes: hegemon, revisionist, rising, bloc member, proxy, double-bound, isolated, buffer, periphery. If the thesis claims that states are positioned by their assigned role rather than by their power, then that is what the map should carry.
These assignments are assumptions, not measurements. The measured layer is separate and looks different on the globe: amber pillars (decision volume) and targeting links.
A single assignment is a simplification. India is both rising and double-bound, Iran both revisionist and isolated, Ukraine both buffer and proxy. Only the dominant one is shown — the same limit the binary K/R choice carried, and it must be stated as plainly.
Two images are produced: one to be looked at (archetype colours) and one to be read (each country filled with a unique code colour, no antialiasing). Clicking resolves the country by reading a pixel at the hit point's UV — running a point-in-polygon test in the browser across 177 countries would be both slow and wrong at the antimeridian.
| country | source | class | licence |
|---|---|---|---|
| United States | Federal Register (JSON API) | A | Public domain — 17 U.S.C. §105 |
| Türkiye | Resmî Gazete (HTML/PDF) | C | Unclear — legal opinion needed before republication |
| United Kingdom | legislation.gov.uk (Atom) | A | Open Government Licence v3.0 |
| Poland | Dziennik Ustaw (ELI API) | A | Official legal text — not copyrightable |
| European Union | EUR-Lex / CELLAR (SPARQL) — Union-level acts only | A | Commission reuse decision 2011/833/EU |
| Canada | Canada Gazette, Part II (SOR/SI, fortnightly HTML) | B | Reproduction of Federal Law Order (SOR/97-5) |
Class A: official API or bulk download · B: regular structure, no API · C: scraping + PDF/OCR · D: restricted access. Turkish content is archived and enters the measurement, but its full text is not republished here; only the title, source link and derived score are shown. For the EU, member-state acts, consolidated texts and duplicate case-law records are out of scope; for Canada only Part II (registered regulations) is taken, not Part I (notices and proposals) — the system counts what is done.
The Gazette was tried first for the UK, and measured: of the 516 notices published on 19 August 2026, 484 were corporate insolvencies, personal insolvencies and probate notices. The wrong source for measuring state behaviour. The real counterpart of the Federal Register is the statutory instrument stream — legislation.gov.uk.
A second trap found there: the feed returns 20 records per page and hides the rest
behind a rel="next" link. Without pagination, every day with more than 20 items
silently lost the remainder — more insidious than returning nothing, because partial data
looks entirely normal. It was caught by noticing that the daily maximum was exactly 20 on nine
separate days: a number repeating at the top of a distribution is a ceiling, not data.
This layer's job is not to decide correctly. Its job is to keep the obviously irrelevant half of the hundreds of daily documents away from the expensive layer. Its threshold is therefore deliberately loose: a wrong elimination is expensive, a wrong pass costs a few cents.
1.0;
if no positive term appears in the title the threshold rises to 2.2.
Every matched term is recorded — which decision was made and why stays auditable.Negative terms do not eliminate a document, they lower its score. A strong positive term can beat a negative one.
The raised body threshold was measured, not guessed: in the first version only 39% of passing documents were genuinely relevant, and the dominant failure was documents whose body mentioned a topic term while the title was about something else entirely — a car-rental regulation passed because the word “sanction” appeared in its text. Gazette titles are deliberately descriptive.
There is no ground truth; the hand-labelled gold set is the only truth we have. The filter is measured against it at every version. The two error types are not equal: a miss is expensive (that document never reaches the labelling layer), a false pass is cheap.
| version | precision | recall | what changed |
|---|---|---|---|
| r1 | 39% | 100% | first version — passes everything |
| r2 | 90% | 56% | negative list + body threshold; cut too much |
| r3 | 89% | 95% | gaps in the positive list closed |
r2's collapse was not caused by the threshold but by omissions in the positive list:
antidumping existed only in its hyphenated form (anti-dumping) and the
Federal Register does not hyphenate it. A single hyphen missed ten documents. Without the gold
set this would have been invisible — which is the whole point of having one.
Live calibration figures are in the dashboard's system tab, alongside the values actually in production.
In Türkiye, presidential decisions, international treaties and board rulings are published as scanned PDFs. Their content is an image; they appear to have no text layer. Our first threshold was simple: “a PDF yielding fewer than 180 characters is scanned.” It was wrong.
We measured the distribution across 83 PDFs and it was bimodal: Poland's genuine text PDFs yield 3,400–4,200 characters per page, Türkiye's scans 4–206. Nothing in between.
206 is not a coincidence. A scanned page does carry a text layer — but it is not content, it is the masthead: “13 August 2026 THURSDAY · Resmî Gazete · No: 33339”.
A total-character threshold mistook that boilerplate for content. A 22-page scanned court ruling counted as “has text” on the strength of 459 characters and never entered the OCR queue. Silently. Of 83 PDFs, 49 were effectively scanned while only 37 were flagged.
400
Measured per page, not in total. A wrong OCR is cheap (a few seconds of processing);
a missed OCR silently becomes a wrong label.Human review was initially triggered by the document average. Looking at the output showed that this was wrong: the Türkiye–Saudi Arabia visa-exemption agreement runs to 17 pages with an average confidence of 58.8. But the decision text is on the first page and is clean — the low score comes from the maps and coordinate tables on the following 16 pages.
70
Not the document average. In a gazette the decision text is always on the first page;
failing to read the annexes is not failing to read the decision.OCR output does not alter the raw archive; it is derived data, stored separately with its confidence. If the engine or language changes it is regenerated without returning to the source. The Turkish language pack is mandatory: without ğ, ş, ı, İ, ö, ü, ç the output silently breaks keyword matching.
Windows are divided by the number of days available, not by their nominal length, and fracture is not computed until the archive is at least 30 days deep. The earlier 21-day rule produced an artefact visible in the time series: Turkish fractures began exactly on the archive's 21st day, because a 30-day baseline computed from 21 days of data and still divided by 30 is roughly 30% too low, inflating the ratio by about 43%.
Gazette sites change structure without notice. When the ingest layer breaks it must not silently return empty: “zero documents today” is an alarm, not a normal day.
But some sources genuinely do not publish. The Federal Register does not publish at weekends or on federal holidays; legislation.gov.uk produces 0–9 documents a day. If we cannot tell the two apart we either miss real failures or get false alarms twice a week and stop reading alarms — the second being more dangerous. The calendar is therefore source-specific.
For Türkiye no fixed holiday calendar is hard-coded. On religious holidays the index page returns HTTP 200 with a notice instead of documents: “pursuant to Presidential Decree No. 10, the Official Gazette is not published today.” That notice is read directly — if the calendar changes the code need not, and “not published” never gets confused with “the parser broke”.
The claim is not “I measure the world correctly.” The claim is: I convert state behaviour into a traceable unit, and I show exactly how I convert it.
The taxonomy, weights and decay coefficient on this page are the values actually used in production — not a separate marketing text. Back to dashboard · thesis model · Türkçe