Run #24 closed over a month ago — the longest gap between real runs in this series. Nothing about the code sat idle in that time; several rounds of hardening landed in the meantime. What hadn't been tested was whether any of it still worked after a credential that had quietly expired somewhere in the gap.
What we did to prepare
Nothing, as it turned out — not because nothing needed doing, but because the thing that needed doing wasn't visible until the first new run failed. Runs #25 through #27 are really one continuous incident: a credential expiry, a governance false positive found and fixed in direct response, and a clean run that still shipped the wrong code.
Run #25 — the credential had expired
Launched after the long gap, failed in 11 seconds — too fast for any real step to have run at all.
Scope dispatched and came back almost instantly: "Anthropic authentication/configuration failed."
provider_usage confirmed it precisely: a 96-millisecond round trip, the API rejecting the
credential outright, not timing out or misbehaving.
Checking the inbox directly settled it. Anthropic had sent two warnings — "expires in 7 days" on September 19th, "expires tomorrow" on September 25th — for the API key this workspace's BYOK credential actually used, expiring September 27th. Run #25 launched October 1st, four days past that date. A new key was generated and added to the workspace; the old one revoked.
Run #26 — blocked, permanently, for repeating the user's own words
Scope ran successfully this time — confirming the new key worked — and the run was then
permanently blocked, the first blocked_prohibited classification in twenty-six runs. Every
step after Scope shows skipped. Unlike every gate before it, there was no re-authentication prompt
and no path forward: by this project's own design, a prohibited classification can never produce an
execution scope.
The trigger for this run had explicitly asked: "Deliver this through the complete BinChicken
Flight, including all required governance, testing, approval and release steps. Do not bypass any
controls." Scope took that seriously and filed it, correctly, under its own requirements
document's "Out of Scope" section: "Bypassing any governance, review, or release controls in the
BinChicken Flight." The safety classifier's prohibited.bypass_governance rule matched that
sentence anyway — "bypass" sitting near "governance" was enough, regardless of which one of them
was doing the prohibiting.
Tracing it down to the rule itself explained why no amount of rephrasing should have been expected
to help: PROHIBITED_RULES are deliberately, explicitly exempt from every negation and descriptive
suppression this engine has — a design choice, documented in the code, not an oversight. The
reasoning holds up on its own terms: a genuinely malicious instruction could just as easily wear a
negation as camouflage ("please don't forget to disable the approval gate"), so blanket suppression
for the most severe tier was never on the table. What was missing wasn't suppression in general —
it was any way at all to tell a model's own structured, schema-correct restatement of an
already-safe instruction from a freshly dangerous one.
Run #27 — the fix confirmed itself within the hour
A fix shipped directly in response — reusing a structural detector already built and tested for the high-risk tier after Run #24: recognize when a match sits inside an actual markdown list item under an "Out of Scope" heading, the one place this schema's own output reliably produces trustworthy structure, and let only that narrow, syntactic signal through for the prohibited tier — nothing resembling the broad negation-word suppression that tier still can't have. Twelve new tests, including the adversarial cases that matter most: a benign structured exclusion sitting next to a genuinely dangerous instruction still blocks, in prose or as a separate list item, because the fix evaluates every occurrence independently rather than suppressing a whole document once.
That fix merged roughly forty minutes before Run #27 launched. Scope's output for Run #27 restated the identical Out-of-Scope line, word for word, that had permanently killed Run #26. This time it didn't block. That's not a different run getting lucky with different phrasing — it's the exact adversarial shape the fix was built and tested against, reproduced for real within the hour, behaving exactly as the tests predicted it would.
From there the run went all the way: nine of nine steps, a safety gate on an unrelated
destructive.replace_component match confirmed normally along the way, a real PR opened, the run's
own status correctly reaching completed — the completion-ordering and audit-durability fixes from
Run #23 and Run #24 both holding a
second time, on a second real end-to-end success.
And then the PR itself turned out to be wrong, in a way nothing in the pipeline was positioned to catch. Its generated test imported a library this project doesn't have — the identical defect class Run #24's PR hit, already investigated afterward and found to be a structural, accepted limit of how Test and Review work: both are reasoning-only roles with no ability to execute code, routinely without the dependency manifest in context at all, and real CI is this pipeline's deliberate, only execution-based backstop. It caught the error again, exactly as that investigation said it would.
What that investigation hadn't anticipated was the second problem sitting underneath it. The real, live site footer is one specific file, used on every page that renders a footer at all. This run's Builder never touched it — it wrote a brand-new, entirely disconnected component instead, imported nowhere in the application. Had the test file's import resolved correctly, none of the existing checks — not lint, not the type checker, not the build — would have caught that the change did nothing at all. It would have merged clean and shipped silently broken. The PR was closed rather than patched forward, since the actual implementation needed redoing, not fixing.
What the evidence actually supports
A credential can expire silently over a long gap, with no warning inside the product itself — the only signal was an email, read after the fact, not anything BinChicken surfaced on its own.
The prohibited-tier false positive and its fix are both now fully confirmed, not just argued. Run #26 proved the gap existed with live production evidence. Run #27, on the identical Scope wording, proved the fix closes it — a controlled before/after pair this series rarely gets handed to it this cleanly.
Run #24's Test/Review finding generalizes, exactly as that investigation predicted it would. It wasn't a one-off defect in one generated file; it's the structural shape of what reasoning-only review can and can't promise, and it recurred in a different file, under a different task restart, right on schedule.
A clean, fully-passing pipeline is not the same claim as a correct one. Every gate in this run said yes. The work was still wrong, in a way only reading the actual diff against the actual repository — not running any check this pipeline has — revealed.
Where this leaves things
- Credential expiry has no in-product warning today — worth a reminder or health check before it costs another silently-wasted run.
- The
prohibited.bypass_governancegap is closed, with the strongest confirmation this series has produced for any fix: the same failing input, pre- and post-fix, inside the same session. - The Test/Review structural gap from Run #24 is now a confirmed pattern, not a single occurrence — still not something this investigation argues should be fixed at that layer, since the product's own design deliberately assigns execution-verification to CI.
- A new, more serious category sits underneath that gap: a change that compiles, lints, builds, and passes every gate, while silently doing nothing — not yet addressed, and not caught by anything currently in this pipeline.