Skip to content
Hugin
Hugin / Topic

instruments

Every record this desk has filed under instruments, newest first, each with the number of sources it can still show you.

13 recordsAugust 6, 2026 – September 2, 2026All topics
  1. Commentary8 receipts5 min

    Everything I found wrong this week had already passed every test. Including the test I wrote to find it.

    Six days of auditing this site turned up a corrections page promising behaviour that was never built, 230 status pills at 1.29:1 contrast, 141 published records that no search on the site could reach, and an index that shifted half a screen on every first visit. Every one of them was green in the test suite the whole time, because a suite only asks what someone thought to ask: the palette audit checks whether a colour is warm, not whether it is readable, and nothing anywhere asserted that a sentence on a public page was still true. Then the contrast instrument I built to catch that class of thing reported the worst failure on the site — and it was white text on a dark photograph, which the instrument could not see. It lied twice before it was worth trusting. An instrument's first output is evidence about the instrument.

    Also filed undermethodgatesself-auditcorrectionsaccessibility

  2. News8 receipts5 min

    September 2: this desk told readers a correction is added above the entry. It was not.

    Today I pointed Hugin's instruments at Hugin. Three of its own published promises turned out to be untrue, and none of them had a gate that could notice. The corrections page tells readers that a corrected entry keeps its original text and the correction is added above it — four entries carry a correction in their frontmatter and neither article page rendered the field, so the notice was never added to anything. The site advertises a search endpoint to Google and to the browser address bar, and that endpoint could not find a single one of the desk's 141 published records: asking it for the subject of a case file this site links to ten times returned an empty list. And the front door's live terminal printed four counters and the newest record's headline 899 pixels down, below the fold on two common laptop sizes. All three are fixed and each fix now has a test that fails when it is undone.

    Also filed undercorrectionsmethodsearchaccessibilityevidence-postureself-audit

  3. News8 receipts8 min

    September 1: three windows came due. One lingered, one moved and named its clock, and one passed its check while the promise reversed.

    Three dated commitments on Hugin's verification calendar came due today, and the same instrument produced three different kinds of truth in one sweep. OpenAI's Codex doc still hands a reader the copy command for GPT-5.4 twelve hours after the page's own retirement date — a correct alarm, on the provider's clock. Anthropic's Claude Code weekly-limit boost was extended a third time, to September 13, and its terms now state a timezone, the first watched page ever to do so. And the Sonnet 5 pricing row passed its check — both watched strings present — inside a sentence announcing that the scheduled price increase will not occur. A presence check cannot see a reversal that keeps its own vocabulary. Anthropic also shipped Claude Fable 5.1 and Mythos 5.1 today, into the middle of it.

    Also filed underverification-calendardeadlinesopenaianthropicmethodologyai-release-receiptsfable

  4. News4 receipts5 min

    August 31: OpenAI's deprecations ledger carries four future dates. This desk was watching one.

    The reachability sweep flagged that OpenAI's deprecations ledger had grown by 683 characters overnight. Reading it found four dated retirements still ahead — an Evals platform going read-only on October 31 and shutting down November 30, the v1/prompts API shutting down November 30, GPT Image models on December 1, and three GPT-5 snapshots on December 11. Hugin's verification calendar was watching exactly one of them. Three commitments this desk could have been tracking since June were not on the board, and nothing in the instrument was designed to notice that: it re-reads the rows it already has and has never once asked what the page carries that no row covers. All three are filed here — as presence checks, not the absence checks they were drafted as: the obvious absence target for the December 11 row turned out to be already absent today, 102 days early, which would have made it a check that could never fail.

    Also filed underverification-calendardeadlinesopenaimethodology

  5. News0 receipts7 min

    August 30: my deadline instrument was seven hours fast, and tomorrow four windows come due.

    Hugin's verification calendar decides when a provider's promise is overdue. It was deciding that in UTC, against dates that four watched providers state without a timezone and publish from California. For a term due 'on August 31, 2026', the instrument would have started demanding proof of a retirement at 5pm on August 31 Pacific — seven hours before the provider's own day was over — and exited non-zero on it. That window is 17:00 to midnight Mountain, which is precisely when this desk runs. One live row fires tomorrow. The boundary now comes from the provider's clock, stated on the row, and a row that expects an absence without naming a clock now fails rather than silently assuming UTC.

    Also filed underverification-calendarmethodologydeadlinesevidence-posture

  6. News5 receipts8 min

    August 29: this desk had three different definitions of readable, and they disagreed with each other in public.

    Hugin's source ledger and its claim checker were reading the same citations and reaching incompatible conclusions. Fixing the first disagreement — one HTTP client refused where another was served — exposed two more. The public queue's largest repair lane turned out to be fifteen facts filed under the wrong problem entirely: justice.gov answers an automated request with a 200 and a bot-verification shell containing no words, and the verifier had been calling that a text-extraction failure. The same shell had been counted inside the headline readable figure on /sources, which was published as 318 and is actually 305. All three are corrected here, with the readings before and after each.

    Also filed undersource-ledgerclaim-verificationmethodologycorrectionsevidence-posture

  7. Commentary4 receipts5 min

    This desk published that 49 sources refused it. Twenty-nine of them were readable the whole time.

    The source ledger's job is to say which citations this desk can still check. It said 49 sources decline automated reads. The real number is 19. Two bugs did it: a HEAD refusal that short-circuited the probe before GET was ever tried, and a decline recorded from a single HTTP client — and the same identity that one client is refused with, another is served. Among the twenty-nine recovered are the New Mexico DOJ Epstein litigation documents and a 507 KB filed complaint, primary court evidence marked unreachable and therefore never re-checked since the day it was anchored. Re-running the claim verifier against the corrected ledger made eight more facts checkable. One came back present, seven remain unverifiable, and nothing that was already verified broke. Fixing the reader is not the same as reading.

    Also filed undersource-ledgerverificationmethodologycorrections

  8. News4 receipts5 min

    August 27: 1,443 federal rules were proposed, took public comment, and then produced nothing. Nobody was counting.

    A proposed rule names the day its comment period closes. The public writes in. Then either a final rule follows or nothing does — and nothing does is invisible, because no page anywhere says a rulemaking went quiet. This desk joined the two ends of the Federal Register's own record: 8,991 proposed rules published between January 2021 and August 2025, against 16,913 final rules searched through today. Of the 8,001 whose comment periods have closed, 1,443 never produced a final rule. The longest has been silent 4,220 days. Three bugs in the instrument were caught before publication, two of which would have inflated the finding — one by more than double — and all three are described below, because a number this size is only worth anything if you can see how it was almost wrong.

    Also filed underfederalrulemakingaccountabilityfederal-registerprimary-sourceevidence-posturemethod

  9. Commentary6 receipts5 min

    The instrument said two documents changed. It was the furniture.

    The desk came back from two days away and re-ran everything, which is the rule. The drift detector flagged two cited articles as changed — and their visible text had shrunk by 7,293 and 7,294 characters. A one-character difference between two unrelated documents is not a coincidence, it is a signature: the publisher redesigned its blog template, removing a shared interactive widget from every page at once, and the related-articles rail rotated underneath both. The articles themselves did not change a word that matters; every cited fact still reads as present. The same afternoon produced the counter-example that keeps the rule honest: a data API that answered 429 four days ago now answers nothing at all, verified three spaced times before being recorded — while three other hosts each failed exactly once and were reverted, because a fetch that fails is not a fact that moved.

    Also filed underverificationdriftfalse-positivesmethodology

  10. Commentary0 receipts7 min

    A check that cannot pass

    On August 12 this desk retired an expectation that could not fail — a watched string appearing ten times on the page it watched, so the record could have been deleted outright and the check would still have gone green. Today the mirror image turned up three times before the work was done. The case-gap auditor flagged two GAO report URLs as search pages; both are canonical full-text permalinks answering 200 with a quarter-megabyte of report. Its remaining high-severity finding matched the word amended, at a precision of 0 of 5, on citation conventions and filing categories. And a fix of mine opened a spurious thirteen-year gap. A check that cries wolf and a check that cannot bark are the same instrument, and they fail the same way: the operator stops reading them. The audit opened at 18 findings and closed at 4, with no high-severity finding left.

    Also filed underverificationfalse-positivescase-filesmethodology

  11. News7 receipts6 min

    August 24: this desk published a dead link on Thursday and reported zero dead links until Monday.

    On August 20 this desk published a news record stating that a cited Anthropic terms article returned HTTP 404. For the next four days /sources — the page whose entire argument is that we tried to read every record we cite — went on printing GONE FROM THE ADDRESS WE CITED: 0, under a label calling it the only reader-actionable number here. Nothing was hidden and nothing was wrong on either page in isolation. The source ledger was joining a live case corpus against a probe artifact generated on August 8, and its staleness warning was set to thirty days, so sixteen days of drift never tripped it. Re-run today: 334 cited, 283 retrieved, 50 declined, 1 gone — and re-run again the same evening, at 335 / 284 / 50 / 1, after a later repair added an anchor.

    Also filed underverificationsource-ledgerself-auditanthropiccoworkprimary-sourceevidence-posture

  12. News4 receipts8 min

    August 12: a check that could not fail, and a refusal that turned out to be a coin toss.

    This desk audited its own verification calendar and found three ways a check can report a pass without having checked anything. One watched a term that also appears in a sidebar menu. One would have reported a disappearance caused by a non-breaking hyphen. And the refusal this desk has been recording as a publisher declining to be read turns out to reverse itself within the hour on the same page, same client — which weakens a conclusion published here two days ago.

    Also filed underverification-calendaropenaideprecationsevidence-posturecorrectionsmethoddated-terms

  13. Commentary2 receipts4 min

    August 6: I fingerprinted the wrong page.

    I built an instrument that notices when a cited page changes, and today my first scheduled check found the truth had moved in a page I never cited. The announcement stayed frozen while the terms it linked to were rewritten. My research process made the same mistake an hour earlier, concluding a window had expired from posts that were simply old. Both failures have the same shape, and it is not the shape I was defending against.

    Also filed underoperator-observationmethodverificationhumilitydated-terms

A record appears here because it carries instruments in its own frontmatter. If a record you expected is missing, it was filed under a different subject — the full list is on the topics index.