Blog

The Amplifier and the Artifact: On the Careful Use of AI in Software Development

Date Published

AI Slop vs Digital Rubber Duck

There are, broadly, two ways to use a generative AI in software development. They produce superficially similar artifacts and radically different engineering cultures.

The first is transactional. A developer needs a function, a regex, a migration script, or a test. They prompt, receive plausible output, skim it, and paste it into the codebase. The interaction is one-shot. The AI is treated as a vending machine: insert a request, extract an artifact.

The second is dialectical. The developer treats the model as a drafting adversary — a tireless generator of candidate reasoning that must be specified, interrogated, corrected, and verified. The interaction is iterative. The developer supplies constraints, inspects omissions, rejects the first plausible answer, and demands that the output justify itself.

The first mode feels productive. It produces code quickly. It is also, in most non-trivial contexts, a mechanism for importing defects at scale.

The Illusion of the First Draft

The transactional mode mistakes generation for engineering. Generation is cheap; engineering is the discipline of establishing that the generated thing is correct, maintainable, and fit for purpose. When the two are conflated, the bottleneck does not disappear — it simply moves downstream, from writing code to reviewing it.

That relocation is dangerous precisely because it is invisible. A developer who writes a function by hand has, by the act of writing, constructed a mental model of its behaviour. A developer who pastes a generated function has not. The model may be correct; the developer's understanding of it is thinner. Multiply this across a codebase and you accumulate a class of debt that is harder to detect than ordinary technical debt: not merely code you would write differently, but code whose behaviour no one has actually reasoned about.

The Specific Failure Modes

The failures are not random. They are patterned, and the patterns are predictable.

Plausible but nonexistent APIs. Models generate usage that looks idiomatic and refers to methods, flags, or library internals that do not exist or have been deprecated. This is most dangerous in large, fast-moving ecosystems where the model's training data lags the current release. The code compiles in the reader's mind and fails in the terminal.

Correct logic, wrong domain. A sort, a pagination scheme, or an authentication check may be syntactically and algorithmically sound while being semantically wrong for the system in question. The model does not know your invariants, your tenancy model, your regulatory constraints, or your historical incidents.

Tests that assert the wrong thing. This is among the most insidious. Asked to write a test, a model will often produce one that passes. But passing and testing are different properties. An assertion may check for equality where membership is required, or may assert against a mock that encodes the same misunderstanding as the implementation. A test suite can grow in size while shrinking in assurance.

Silent security omissions. Generated code frequently omits rate limiting, input sanitisation, authorisation checks, constant-time comparison, or transaction boundaries — not because the model is malicious, but because these are the parts of the specification that were never stated. Absence of instruction becomes absence of defence.

Architectural drift. Each generated component may be locally reasonable and collectively incoherent. Without a human holding the design, the system accretes patterns rather than expressing one.

Register and context mismatch. Generated prose, documentation, and commit messages often miss the intended audience and tone — a minor issue individually, a corrosive one across a project's public surface.

The Verification Asymmetry

The core structural problem is that generation and verification scale at different rates. A model can produce a thousand lines in seconds. A competent engineer cannot meaningfully review a thousand lines in seconds. When generation outpaces review, the unreviewed surplus does not vanish; it becomes latent risk.

This is why the transactional mode is not merely less rigorous — it is unstable. It works only while the volume of generated code remains small enough for informal, in-passing review to be adequate. Beyond that threshold, the practice degrades quietly, and the degradation is attributed to something else: the framework, the requirements, the team, the deadline.

What Disciplined Use Looks Like

The dialectical mode is not slower in any meaningful sense. It is differently sequenced: more effort before and after generation, less effort spent debugging the consequences of unexamined generation.

Supply constraints, not just requests. State the language version, the framework, the architectural layer, the error-handling convention, the security requirements, the test framework, and the audience. Ambiguity in the prompt becomes ambiguity in the artifact.

Demand rationale and trade-offs. Ask why a particular approach was chosen, what alternatives exist, and under what conditions the choice would be wrong. If the model cannot articulate the trade-off, the decision has not been made — it has been defaulted.

Ask what is missing. Explicitly request the failure modes, the edge cases, and the assumptions. Models are better at responding to "what does this not cover?" than at volunteering the same unprompted.

Interrogate the tests. Verify that assertions test what they claim. Distinguish membership from equality, presence from truthiness, and behaviour from implementation. A test that cannot fail is not a test.

Refuse the first plausible answer. The first output is a proposal, not a conclusion. Treat it as one.

Verify against primary sources. For anything load-bearing — APIs, standards, cryptographic primitives, domain rules — confirm against the actual specification, not the model's recollection of it. This is non-negotiable in regulated or safety-critical contexts.

Keep the domain model human. The model may assist in expressing the design; it should not author the design's semantics. Whoever holds the invariants must be a person who can be held accountable for them.

Red-team your own outputs. Ask the model to attack the code it just wrote. Adversarial review is cheap to request and disproportionately productive.

The Cognitive Trade

There is a subtler cost than defects: the atrophy of the capacities that made the defects detectable.

If AI is used to avoid thinking — to skip the formulation of the problem, the enumeration of cases, the design of the interface — then the judgment required to evaluate its output is precisely the judgment that was never exercised. The tool is then not augmenting expertise; it is substituting for its development. Over time, the organisation loses the ability to tell good output from bad, because the criteria were never internalised.

If AI is used to accelerate thinking — to generate candidates against a problem the developer has already framed, to surface counterarguments, to draft and then be corrected — then judgment is exercised more, not less. The developer remains the authority on correctness, and the model compresses the mechanical portion of the work.

The distinction is not about how much AI is used. It is about who holds the specification, who holds the verification, and whether a human being can still explain, on demand, why the system behaves as it does.

A Practical Posture

A reasonable working discipline is this:

The model proposes; the engineer disposes.

Nothing merges that no one has reasoned about.

Every generated test is inspected for whether it can fail.

Every generated dependency, API call, and security control is verified against the source.

Architectural decisions are made by people and recorded in prose, not inferred from a diff.

The rate of generation is bounded by the rate of genuine review.

This is unglamorous. It also happens to be the only version of AI-assisted development that survives contact with production.

Conclusion

AI is an amplifier. It multiplies whatever discipline is applied to it. Applied to careful specification and adversarial review, it compresses the mechanical labour of software and leaves more room for design and judgment. Applied to the avoidance of thinking, it accelerates the production of code that no one understands, at a rate no one can audit.

The difference between the two modes is not sophistication of tooling. It is whether the human remains the author of the specification, the arbiter of correctness, and the person who can explain the system. The tool does not decide that. The engineer does, prompt by prompt.