Back to blog
Aug 31
dogfooding
byok
anthropic
provider-reliability
governance
engineering-log

Run #24: A Clean Success, and a CI Check That Actually Worked

Run #23 closed with its mystery fully explained — a missing database constraint value had been quietly breaking both the success and failure paths of the real GitHub handoff — and a fix for it, alongside a Builder truncation fix, shipped before this run launched.

What we did to prepare

Two real fixes landed. The audit_events_type_check constraint was widened to include all seven github_* audit event types Run #23's tracing had found missing — the gap explaining why a genuinely successful run never reported itself as complete. Separately, Builder's truncation pattern from Run #22 got its own dedicated investigation: comparing every real attempt against this task (six, not three) found the created test file's size — not the modified file, which stayed almost constant — was the volatile driver of output length, and a targeted prompt clarification shipped for that specific cause. Neither the token ceiling nor the provider timeout was touched, since the investigation also found the two were already close to co-binding — raising one without the other would likely have just traded one failure mode for another.

What happened

This run's task was different from the recent string of attempts — a dynamic copyright year in the footer, not the run-detail timestamp. Scope ran, and its output tripped a high_risk safety gate on two rules at once: "Touches CI/CD configuration" and "Touches infrastructure or deployment configuration." Scope then ran a second time, producing an independently-worded version of the same requirements — and tripped an equivalent gate again.

Both gates were false positives, and clearly so: both requirements documents state, in their own Out-of-Scope sections, that CI/CD and infrastructure would explicitly not be touched. Traced directly against the actual rule (modules/safety/policy.ts): high_risk.cicd_change matches the bare string "CI/CD" anywhere in a document, with no requirement that it appear near an actual change. Its sibling rule, high_risk.infrastructure_change, was hardened against exactly this failure mode after Run #6 — required to sit near a real change-verb before firing. cicd_change never received the same fix. Both gates were re-authenticated and confirmed, and the run continued.

From there, everything ran clean: Architecture, Security, UX, Build, Test, Review, Approval, and Release all completed without incident, two ordinary approval gates decided along the way. Nine of nine steps — for the third time now — and this time everything downstream of that worked too. A real branch, a real commit, a real draft pull request: #79. The run_completed audit event fired, and — for the first time on a fully successful real-GitHub run — the run's own database status actually updated to completed, with all five github_* audit events present and no gaps. The fix held.

The interface caught up this time, too. Reopening the run's page mid-flight (after "nothing popped up" on the first check) restarted the run's own update loop and rode it through live to a proper "Hatching complete" card, linking straight to the real PR — no manual guesswork about whether a reload would surface anything, the way Run #23 had required.

Then a genuine problem surfaced in the PR itself. Its generated test file imported @testing-library/react and used a jest-dom matcher — neither of which exists in this project's dependencies, which are bare vitest only. tsc failed outright in CI; the Vercel preview deployment failed as a direct consequence. BinChicken's own Test and Review steps had both passed this change before it ever reached GitHub — they evaluate generated code by reasoning about it, not by running it against this project's real toolchain, so a plain, verifiable compile error made it all the way to a real pull request. Real CI is what actually caught it, correctly blocking the PR from being mergeable.

What the evidence actually supports

Three separate, previously-broken things all held at once. The audit-constraint fix, the Builder truncation fix, and the client-liveness restart-on-reload behavior each did exactly what they were built to do, on the first real run to exercise all three together. Nothing here is an argument that they'll always hold — it's the first genuine evidence that they do.

The CI/CD safety rule is a structural false-positive generator, not a coincidence. It fired twice, independently, on two different phrasings of the same negative statement, because the rule as written cannot distinguish "this touches CI/CD" from "this document contains the word CI/CD." A sibling rule already carries the fix this one needs.

The pipeline's own review steps cannot catch what only running the code would catch. This is a different category of gap from anything found before — not a provider timeout, not a database constraint, not a UI staleness problem, but a limit on what AI review can verify about code it never executes. Real CI covered for it this time. Whether that's sufficient as a permanent answer is an open question, not one this run settles.

Where this leaves things

  • Every fix shipped after Runs #20/21/22/23 is now confirmed working, together, on one real run.
  • The high_risk.cicd_change rule is a known, precisely-located gap, structurally identical to one already fixed in its sibling rule — flagged, not yet touched.
  • A real compile failure reached a real pull request past two AI review gates, caught only by the CI pipeline that actually runs the code — the clearest evidence yet that AI review and execution-based verification are checking different things, not substitutes for each other.
  • PR #79 remains open and correctly failing CI — not merged, pending a decision on whether to fix the test file's approach or close the PR outright.