Skip to main content

When the Agent Maintains the Agent's Code

· 10 min read
Calvin Cheng
Shape what gets built and the value it creates.

The Governance Speed Problem addressed how fast governance can move. This post addresses a more fundamental question: what does readiness mean when the system being governed was not written by a human and is not maintained by a human? The Other Dimension series assumed a human authored the AI-assisted output and a human maintained it. At Sau Sheong's Levels 3 and above — autonomous agents and collaborative agent networks — that assumption breaks. The code being reviewed was written by an agent. The changes being proposed are written by an agent. The human reviewer is evaluating agent work on agent work. Every readiness component shifts meaning.

Ownership clarity does not mean the owner wrote the code. It means the owner can explain — under pressure — what the code does and why it changed.

The brownfield turn​

Most conversations about AI in software development focus on greenfield: agents building new things. The harder problem — and the one that arrives sooner in any organisation with production systems — is brownfield: agents maintaining existing systems that were themselves built with agent assistance. This is not a future concern. It is the present reality for any team that adopted agent-assisted development six months ago and now needs to modify, extend, or debug what was produced.

At NUS, systems built with AI assistance in one quarter become the maintenance target for the next quarter. The Data Engineering team's AI-generated pipelines need modification when upstream schemas change. The Integration team's AI-generated API mappings need updates when partner specifications evolve. The agent that maintains the code did not write the original version. It is reading code that another agent wrote, inferring intent from structure rather than memory, and proposing changes based on pattern rather than understanding. The human reviewer is now evaluating a proposed change to code they did not write, made by an agent that did not write the original, in a system whose design rationale may exist only in the structure of the code itself.

Ownership clarity without authorship​

The Invisible Line defined ownership clarity as: a named human who understands what the system does and has accepted responsibility for its outputs. When the human authored the system — or at least directed its authorship — understanding followed naturally from the act of creation. You understand what you built because you built it.

When the agent builds it, understanding must be constructed independently of authorship. The named reviewer did not write the code. They must still be able to answer the four questions: what does this system do in production, what data does it touch, what happens when it is wrong, and who is woken up? But now there is a fifth question: can you explain why the code is structured this way — not because you designed it, but because you have studied it enough to defend the design decisions embedded in it?

This is not an unreasonable standard. Engineers have always inherited codebases they did not write. The difference is scale and frequency. When agent-generated code is the majority of the codebase and agent-proposed changes arrive daily, the reviewer cannot rely on institutional memory or authorship knowledge. They need a different kind of understanding — one built from reading, testing, and interrogating rather than from creating.

At Hedera, I explained aBFT architecture to enterprise teams who had not designed it. The explanation required understanding deep enough to answer adversarial questions — not because I wrote the consensus mechanism, but because I had studied it until I could defend its design choices and name its limits. That is the standard for ownership clarity in agent-maintained systems: understanding achieved through study rather than authorship, deep enough to withstand scrutiny.

Failure-mode awareness when nobody remembers the original​

Try to Break It First argued that failure-mode awareness comes from deliberately trying to break the system before deployment. That practice assumed someone on the team understood the system well enough to construct meaningful adversarial scenarios. When the system was agent-built and the person constructing the scenarios did not author it, the failure scenarios tend toward the generic — standard boundary testing, common input fuzzing, expected error paths. The subtle failure modes — the ones specific to this system's architecture and this system's assumptions — require understanding that is harder to achieve without authorship.

The Esco Micro freezer prototype taught me that the most dangerous failure modes are the ones specific to the interaction between disciplines — what happens when the mechanical behavior and the firmware response disagree. In agent-generated systems, the equivalent is the interaction between components that were generated independently — each internally coherent, but making assumptions about the other that were never explicitly coordinated. An agent maintaining one component does not know what assumptions the agent that built the adjacent component encoded. A human reviewer cannot know either, unless the architecture makes those assumptions explicit.

This shifts the constructive payload. For agent-maintained systems, failure-mode awareness requires architecture that makes assumptions visible — not just for human readers, but for agents that will maintain the code later. Interface contracts, assumption documentation, and boundary specifications are not nice-to-have comments. They are the mechanism that enables meaningful failure-mode work on systems whose authorship history is opaque. The pre-deployment failure scenario exercise from the series still applies, but it must be preceded by an assumption-surfacing exercise: what does this system assume about its inputs, its adjacent systems, and its operating conditions? If those assumptions are not documented, the failure scenarios will miss the interesting failures.

Confidence calibration in recursive generation​

How Confident Should You Be? established that confidence calibration means understanding how reliable the system's outputs are under different conditions. For agent-maintained systems, there is a recursive dimension: how confident should you be in the agent's proposed changes to a system whose original confidence profile was established for the original code?

At GoNetZero, confidence in a carbon calculation depended on the assumptions embedded in the model, the data quality of the inputs, and the conditions under which the model was validated. A change to the model code — even a small one — could shift the confidence profile without changing the headline output. The same applies when an agent modifies agent-generated code: the change may look correct (tests pass, output format unchanged) while shifting the conditions under which the system is reliable. The confidence profile of the modified system is not the same as the confidence profile of the original system, even if the modification was small.

This means confidence calibration for agent-maintained systems must be change-aware. Each agent-proposed modification should be evaluated not just for correctness (does it do what was requested?) but for confidence impact (does it change the conditions under which this system's outputs are reliable?). The confidence conversation from the series — what conditions make outputs reliable, and what degrades reliability — must be repeated not just at initial deployment but at each significant modification. This is expensive. The portfolio triage mechanism from the previous posts determines where the expense is justified.

Recovery design when the fix is also agent-generated​

The Feedback Loop Is the Thing argued that recovery design is about how fast you learn from failure and how effectively corrections reach the right person. When an agent maintains the system, the recovery loop has an additional node: the correction itself may be agent-generated. A failure is detected, routed to the responsible human, and the human directs an agent to fix it. The agent proposes a fix to code it did not write, for a failure whose root cause it infers from symptoms rather than understanding from memory.

The MDM hard-delete incident at NUS is instructive here. The failure occurred because the SOP for a specific scenario lived in a human's head rather than in a system. Imagine the same failure in an agent-maintained system: the SOP lives in no human's head because no human authored the system. The failure mode was not documented because the generating agent did not know it was important. The recovery agent cannot fix what it does not understand was a design choice rather than a bug.

Recovery design for agent-maintained systems requires explicit design rationale — not just what the system does, but why it does it that way. When a failure occurs and an agent is directed to fix it, the agent needs to distinguish between code that is wrong (should be changed) and code that is deliberately constrained (should not be changed without understanding the constraint). Without documented rationale, every quirk looks like a bug, and every fix risks undoing a deliberate choice whose purpose was never written down.

The compounding opacity problem​

Each generation of agent-on-agent modification increases opacity. The first generation: a human directs an agent to build a system. The human has intent; the agent has capability; the code reflects both. The second generation: an agent modifies the system to meet a new requirement. The modifying agent infers the original intent from structure. Some intent is preserved; some is lost in translation. The third generation: another agent modifies the modified system. The original intent is now two layers of inference away from the current code.

This is not different in kind from human-maintained legacy systems. It is different in speed. Human-maintained code accumulates opacity over years. Agent-maintained code can accumulate the same opacity in months because the modification cadence is so much faster. The governance mechanisms from the series — ownership clarity, failure-mode awareness, confidence calibration, recovery design — all degrade faster when modification happens at agent speed. The readiness assessment cadence must match the modification cadence, not the human memory of when the system was last reviewed.

What this demands of architecture​

The series argued that readiness is organisational, not technical. This post argues for an exception — or rather, a refinement. In agent-maintained systems, certain technical architecture choices either enable or prevent the organisational readiness the series advocates. Specifically:

Explicit interface contracts enable ownership clarity. The named reviewer can understand the system at the boundary level without understanding every implementation detail, because the contracts tell them what each component promises and what it requires.

Assumption documentation enables failure-mode awareness. The pre-deployment failure scenario exercise can target the interesting failures — assumption violations — rather than generic boundary tests.

Confidence metadata enables confidence calibration across changes. Each component carries a record of the conditions under which it was validated, so a modification can be evaluated against the original confidence profile.

Design rationale documentation enables recovery design. The recovery agent — human or AI — can distinguish deliberate constraints from incidental structure, and fix failures without undoing design choices.

These are not new ideas in software architecture. They are well-understood practices that have been optional in human-authored systems because institutional memory and author availability compensated for their absence. In agent-maintained systems, institutional memory does not exist at the authorship level, and author availability means nothing because the author has no memory of the act. The practices become load-bearing.

Two questions for your engineering leadership​

  1. For your longest-running AI-generated system, how many modification generations has it been through — and can a named human still explain the design rationale for its current structure? If the answer is uncertain, the next agent-proposed change is being reviewed without the context needed to evaluate it.

  2. Does your architecture mandate explicit interface contracts and assumption documentation for agent-generated components — or are those optional practices that depend on individual discipline? In an agent-maintained world, optional documentation is no documentation.


Beyond The Other Dimension · Follow-up 3 of 3 · Previous: The Governance Speed Problem