Runs #25–27 ended with a pull request that compiled, passed every gate, and changed nothing a user could see: Builder had written a brand-new footer component that nothing imported, while the real footer sat untouched. Three fixes were written in response. Run #28 was the first run to exercise any of them.
What we did to prepare
Three changes merged between the two runs, the last one seven minutes before Run #28 launched.
The first was about visibility. Tracing Run #27 against real production data showed why the real footer
never reached Builder: the context scorer counted flat keyword matches, and components/marketing-footer.tsx
scored 1, matching only the word "footer", which ranked it about 280th of roughly 298 candidates for
every role except UX. The fix weights rarer words more heavily. Re-run against the same repository and the
same title, the footer now ranks in the top five for Architecture, Build, UX, Review and Test.
The second was about the human gates. Run #27's own Review and release notes had both said, in plain words, that the implementation was a disconnected parallel component. A person approved the gate anyway, and the gate hadn't made that warning meaningful: it showed only the gated step's own artifact, and the review's individual findings were discarded the moment they were reduced to one confidence number. Gates now carry the artifacts of the steps they depend on, plus a plain-language brief.
The third was about the premise itself. The request behind Runs #24 through #28 asks to make a "hard-coded copyright year" dynamic. The footer contains no copyright text of any kind, a task-authoring fact the investigation recorded plainly and left alone. Nothing in the pipeline had ever checked claims like that, and Scope can't: by design it reads no repository files. So the Architect, the first role with real file content, now must record an explicit result, confirmed, contradicted, unresolved, or not applicable, and a response that omits it fails the step. The first version of this fix inferred "confirmed" from the absence of a finding. A review before merge caught that silence and verification looked identical, and replaced it with a required field. The result is meant to appear at the security gate as its own line. It is advisory by design: the human can still approve.
What happened
Scope, Architecture, Security, UX, Build, Review, Test and the Approval step each took 38 to 50 seconds, with the usual parallel pairs. Scope's input was 796 tokens, the request and little else, so it restated the false premise as fact, hedged only with an "e.g." and a placeholder company name.
Architecture recorded the contradiction exactly as designed. Its stored result carries
existingStateGrounding: contradicted, with confidence 0.4. In prose it went further than the schema
required. It wrote that its whole approach was "conditional on that grounding result being surfaced to the
human approver before Build is allowed to run," headed its component list "if the human approves proceeding
despite contradiction," and offered an alternative, naming it "the more honest framing": reject the request
as stated and re-scope it as a pure creation request, which "avoids silently inventing a requirement."
The roles after it kept asking the same question. UX asked whether the human approver had "explicitly acknowledged that no copyright year currently exists" and "consciously approved reframing this as a net-new copyright addition." Security asked the same, and called it "the explicit precondition the architecture artifact places on Build being allowed to run." It also asked how anyone would verify the change was actually rendered, rather than a disconnected component that passes tests.
The only human decision before Build was the security gate. It was requested at 07:04:19 and approved at 07:04:55, 36 seconds later, with no comment. Its evidence included the architecture artifact and the security review.
Build then did the work, on the real file this time. It modified components/marketing-footer.tsx
rather than inventing a new component, so the Run #27 failure didn't recur. Its patch added a fourth span
and described itself as "net-new." The patch and two later artifacts then recorded the human's consent as a
fact: Build's notes said the contradiction had been surfaced to the approver and the architecture "explicitly
describes proceeding as a net-new feature addition"; Review said a human approver "approved proceeding as a
net-new addition"; the release notes said the approver "was informed and approved proceeding as a deliberate
creation task."
The release gate was rejected, three minutes and 44 seconds after it opened. The comment names the
principle precisely. The request was to make an existing year dynamic while preserving the footer. The
Flight established that no such year exists, then treated that as authorisation to add a new notice and a
fourth footer element. "A contradicted existing-state premise must not silently convert a MODIFY request into
a CREATE request." Release was skipped, the run ended failed at 07:12:08, and no branch or pull request was
ever created.
What we found tracing it
Detection worked at the provider boundary. The Architect's explicit result is stored on its execution record, the required-field enforcement did its job, and the prose artifacts downstream all carry the contradiction. That half of the third fix is confirmed in production.
The structured result never reached the gate. The artifact table has no column for it, nothing in the persistence code or the migrations mentions it, and the function that rebuilds an artifact's metadata from a database row restores seven fields, none of them this one. The approval panel reads the result from that metadata, and by design shows nothing when it is absent. The run API loads artifacts through exactly that function. So the dedicated "existing-state check" line most likely never appeared at the security gate. We can't see the screen, so that is inferred from the code path, not observed. It is the same shape of gap as Run #23's audit constraint: code that is correct when held in memory, running against storage that was never extended to match.
Some signals did survive. Confidence is persisted, and the security review and architecture artifacts both sat at 0.4. By the second fix's design, that gate's brief should have read red, "Blocking issue found," with the line "Approving will continue the Flight despite this unresolved concern." The architecture's own prose, naming the contradiction, was attached as evidence. The release gate, by the same function, should have read neutral, because the only reviewing-role artifact there, the code review, sat at the ambiguous 0.7. If both inferences hold, the gate with the strongest structured signal was approved after 36 seconds, and the one with none was rejected after the person read what it carried. The release notes there stated the contradiction outright.
The deeper break was in what an approval means. The Architect made its recommendation conditional on one thing: that a human explicitly approve the reframing. The workflow has two gates and each offers only approve or reject; rejecting fails the run, and a "changes requested" outcome exists in the data model but has no route behind it. So that precondition had nowhere to be satisfied except the security gate's Approve button, and three artifacts downstream recorded that button as the satisfied precondition. The design record written seven minutes earlier says the opposite, that approving never means inventing a new requirement. The behaviour and the stated design disagreed, and the artifacts inherited the behaviour.
What the person at the security gate actually read is unknown. The database holds what was attached to the gate, not what was rendered or scrolled.
The test would likely have failed CI again. Build's test imports @testing-library/react, which
package.json doesn't contain, the third footer run to reach Build to do so. Review raised it as an open
question: it couldn't confirm the dependency from the files it had been given. It passed anyway at 0.7
confidence, which fits what the earlier investigation of Test and Review found about what reasoning-only
roles can verify. The run stopped before CI could weigh in, so that failure is inferred from three
precedents, not observed.
What the evidence actually supports
Of Run #27's defects, one is closed on evidence, one is half-closed, and one is open. The wrong-file failure did not recur. The false premise was detected and written down, but its dedicated signal to the human most likely did not arrive. What an approval is allowed to mean is untouched.
A recorded consent needs a recorded decision behind it. Three artifacts asserted that a human had approved a specific reframing. The decision log holds one approval, of a security gate, with no comment. The claim wasn't invented from nothing, since the Architect had said approval was required, but nothing tied it to anything the person had actually been asked.
The later gate worked because of what it carried. The release gate showed the code review, the test plan and the release notes together, which is what the second fix was for, and the person reading them stopped the run.
Where this leaves things
- Detection is confirmed working; delivery to the gate is not. A gap between what the Architect records and what the artifact table can store sits between the two. Nothing here changes it; this is what the trace established.
- The approval-as-consent gap is open and well characterised. There is no action that means "I accept that this request is being reframed," so the nearest approval gets cited as that. What an approval should be allowed to mean is a decision for outside this post.
- The task itself is still unchanged. Five runs have now been given a request that contradicts the repository. Run #28 holds the only gate rejection recorded in the audit trail.
- The test-dependency import remains open for a third time, now with Review noticing it and being unable to settle it.