Do I Need These 7,100 Lines?

A working feature isn't the finish line when most of the code is unneeded risk.

7 minute read

A developer examining two contrasting code diffs on a monitor, one sprawling and complex, the other clean and compact.

A developer’s pull request shows up for a working feature and thorough end-to-end tests backing it. The implementation works as promised, passing every validation check without a hitch. Yet when you ask how the architecture hooks together under the hood, they hesitate. The code runs and the test suite passes, but nobody can explain what the system actually does behind the interface.

That gap between functional execution and human comprehension is quietly turning into a serious operational hazard. When teams celebrate delivery speed without interrogating the footprint left behind, they trade immediate progress for long-term fragility.

The Weight of Unchecked Scaffolding

I ran into this pattern firsthand on a current project. Our team is encouraged to use AI tools for prototyping and feature development. An engineer brought me a backend capability designed to resolve a longstanding architectural bottleneck. On paper, it was stellar work. The problem was real, the solution addressed customer pain directly, and the feature functioned end to end.

Then came the pull request review. The implementation spanned roughly 7,500 lines of code with an additional 6,000 lines of automated tests. When I asked the engineer to walk me through how data moved through the flow and the safety guards, they admitted they couldn’t explain the internal mechanics. The AI agent had generated the code, run the unit suite, and delivered a green build.

This disconnect matters because software rarely lives in isolation. Without knowing what the code does, you can’t forecast how it will behave under concurrent load, evaluate its interaction with dependent services, or catch performance degradation before an incident occurs. In previous articles, I have explored both cognitive debt and the delegation gap. When developers cannot explain their own implementations, you lose the mental model required to operate the software.

 
I’m well aware there’s an growing sector of the developer community right now that is leaning heavy into “no one care’s how it works if it works.” I will continue to hold on to my crumudgen hat and care about how it works. That’s bigger, opinionated topic for another time.

I sat down that afternoon and spent an hour rewriting the component. Again, the research was excellent and did a great deal of the heavy lifting for how packets needed to be structured and a slew of memory opcode that’d have been tedious to calculate. However, the generated solution had wrapped everything in layers of speculative orchestration. The agent had built redundant wrappers and custom configuration parsers that duplicated existing capabilities in the platform.

Once I stripped away the extra machinery, I pointed the engineer’s original 6,000 lines of tests at my new implementation. Every single test passed.

The final count was roughly 400 lines of code.

94%
reduction in total code footprint
400 lines vs. 7,500 lines, passing 100% of the original 6,000 test lines
me "seeing if I could"

Why Footprint Is Surface Area Risk

When the engineer and I sat down to compare the pull requests, the engineer asked a fair question: if both versions pass all 6,000 lines of tests, why does the complexity of the implementation matter?

The answer comes down to operational risk. Those extra 7,100 lines of code are not harmless scaffolding. Every unneeded function is another place where an unhandled null pointer can panic a worker thread or an unexpected permission check can fail. When code is generated in volume, our exposure to memory leaks and edge cases expands with every extraneous block.

 
Working software is only half the contract. Every line of unneeded code expands the blast radius of future incidents while requiring perpetual maintenance from engineers who did not write it.

Unconstrained agents naturally produce bloated architectures–they’re, for lack of better words, covering their bases. Prompted without explicit guardrails, they default to constructing independent universes. An agent creates custom logging helpers instead of importing the existing application logger. It invents translation layers to handle scenarios that never occur in practice. It writes code to satisfy the prompt’s ambiguity rather than asking whether the problem could be solved by existing primitives.

Unnecessary code acts as an operational liability. The goal of using AI in development cannot simply be producing running software faster. True due diligence requires coaching agents to right-size their output, guiding them toward the minimal footprint that satisfies the requirement.

The Mirror in the Commits

A few days later, that same engineer ran an experiment. Taking both commits, a few of their own PRs they had in flight, and several of my recent pull requests, they prompted Claude to build a profile of my coding style. The goal was practical. By attempting to distill what made my code different from theirs into a clear prompt framework, the engineer hoped to give agents the constraints needed to produce compact, production-ready code on the first attempt.

I thought the experiment was clever right up until I read the profile the model produced. Seeing your own professional habits reflected back by an algorithm feels remarkably like being judged the a machine.

The model identified nine specific rules defining how I approach code:

  1. Build inside the existing application and use established frameworks. Avoid creating secondary loggers, custom configuration layers, or redundant scanners for features already supported by the platform.
  2. Shrink the problem before solving it. Avoid building scaffolding or abstractions for requirements nobody requested.
  3. Deliver the whole feature, proven in tests and validated in actual use.
  4. Keep comments brief and place deep research in audit trail documentation. When code requires paragraphs of commentary to justify its presence, treat that as a signal to simplify the logic.
  5. Keep process artifacts out of the repository. Allow automated build tools to handle generated assets rather than checking them into version control.
  6. Keep the public surface area small.
  7. Native patches fail safe and are checked hard.
  8. Structure tests around pure units, malformed inputs, and verifying that state remains untouched on refusal.
  9. Leave no temporary scripts or scratch files behind.

I smiled while reading the list. The summary was accurate. What struck me was how quickly an AI model could extract decades of hard-won engineering discipline from a small handful of git commits.

The Documentation Half-Life

Turning personal engineering habits into a prompt profile can work because agents are only as disciplined as the knowledge available to them. That reality exposes a two-part investment teams often overlook when adopting AI coding assistants.

First, the up-front commitment to capture architectural intent. In his writing on agentic engineering, Addy Osmani notes that skipping design thinking is precisely why models invent weird abstractions and speculative bloat. Empirical research into coding agents shows they rarely consult traditional API documentation, spending less than two percent of their interactions browsing external manuals. Instead, they lean almost entirely on repository-level context files and architectural boundaries. In earlier writing on cognitive debt, I touched on Margaret-Anne Storey’s concept of intent debt, which defines the loss of captured rationale explaining why a system exists. When an agent enters a codebase with no written boundaries, it assumes no native primitives exist and defensively manufactures its own scaffolding.

Second, the ongoing tax of keeping those documents current. Documentation carries a strict half-life. When human developers read an outdated guide, they notice discrepancies against the live code and adjust their mental models. An agent, by contrast, treats repository files as absolute ground truth, faithfully serializing deprecated libraries and obsolete schema assumptions across thousands of lines of fresh pull requests.

Without an ongoing commitment to update architectural documentation alongside code changes, agents quietly convert stale assumptions into fresh technical debt.

Guiding Agents Toward Simplicity

Keeping documentation fresh protects the boundaries, but the real test happens during code review. Without explicit constraints, an agent will readily generate thousands of lines of orchestration to solve a four-hundred-line problem. Due diligence means coaching our tools toward restraint, establishing strict boundaries, and refusing to accept code simply because the test suite lights up green.

This is your reminder to go check those AGENTS.md, CLAUDE.md, and other agent skill files to ensure you’ve equipped your agents with the right skills and tools to code to your standards.

 
Before merging an AI-assisted pull request, ask what native primitives already solve the problem and whether the feature can stand on a smaller footprint.

True due diligence means guiding agents to solve the problem with the smallest possible footprint, so your team avoids inheriting thousands of lines of unneeded risk.

When your team reviews the next AI-generated feature, are you verifying that the code works, or are you checking how much unneeded surface area you are agreeing to maintain?