ztzoff.tech

Sep 21, 2026

The source code may no longer be the source

If an agent can reliably produce the implementation, the durable artifact worth owning may be the specification it was generated from. An essay on intent-first development.

There is a practice spreading through repositories right now that looks like housekeeping. Teams are putting Markdown next to the code it governs: an AGENTS.md at the root, a module document beside each package, the conventions and invariants an agent loads before it writes a line. We argued for this ourselves — an in-repo wiki written for the machine rather than for the new hire.

That idea has already won. This essay pushes on the step after it, which has not won and may turn out to be wrong:

Markdown may not merely live alongside source code. It may become the artifact from which source code is derived.

The interesting question in AI-assisted engineering is no longer "will the model write the code?" It writes the code. The question underneath is harder and much less discussed:

If a machine can reliably produce an implementation, what is the durable artifact humans should own?

The claim here is narrow, and it is not that code stops mattering. Code remains executable truth at runtime. The claim is stranger than that: code may stop being the artifact humans primarily author without ceasing to be the artifact machines execute.

Today:       requirements -> human interpretation -> code -> runtime behavior
Possible:    specification -> agent synthesis -> code -> verification -> runtime behavior

Code has always been an abstraction

Every generation of this profession has handed a layer to a machine.

machine code -> assembly -> high-level languages -> frameworks
-> declarative infrastructure -> domain-specific languages -> ?

Nobody hand-writes register allocation any more. Nobody sshes into a box to edit a config file and calls it deployment. Each rung was abandoned for the same two reasons, in this order: the translation downward became reliable, and the output became verifiable without reading it.

That ordering matters, and it is where the analogy to AI has to be handled carefully. We still read compiler output when it matters — we spent a whole post reading JIT assembly to find out why a bounds check survived. The lower layer never becomes unreadable. It becomes something you inspect on purpose rather than author by default.

A historical progression is an analogy, not a proof. It does not establish that another rung is coming. It only tells you what a rung requires.

AI changes the economics of translation

Specifications have failed repeatedly — waterfall requirements, UML, model-driven architecture — and they failed for one structural reason. A human still had to convert the specification into an implementation by hand. That made the spec pure cost with no leverage: it was written once, diverged in week three, and became decoration by month six. The expensive step was always this one:

intent -> implementation

Coding agents have collapsed the cost of that step by something like an order of magnitude. That is the whole event. When translation is expensive, you store knowledge at the bottom of the stack, in the artifact that runs. When translation gets cheap, the optimal place to store knowledge moves up.

But cheap is not the same as trustworthy. The economics moved; the trust did not move with it. Everything difficult about intent-first development lives in that gap.

What already exists, and what is still a hypothesis

This distinction deserves to be explicit, because the space is loud right now.

Shipping today. GitHub's Spec Kit implements a Specify → Plan → Tasks → Implement → Converge flow in which each phase produces a Markdown artifact that feeds the next. AWS Kiro makes the spec the unit of work: a requirements document, a design document, and a task list with traceability back to the requirements. Agent instruction files are ubiquitous. Plan-then-execute modes are standard in every serious coding agent.

Being bet on commercially. Tessl is the most aggressive version of the thesis — spec-as-source, with code as a regenerable output you are not supposed to hand-edit. At OpenAI, Sean Grove argued in The New Code that specifications rather than prompts or code are the durable unit of programming, and that the code an engineer writes is a minority of the value they contribute.

Still a hypothesis, mine included. That a specification plus a verification suite can carry enough information to regenerate a non-trivial production service — repeatedly, across model versions, and possibly across languages. Nobody has demonstrated that at scale. What follows is an exploration of what would have to be true, not a report of what is.

A specification is not documentation

Most teams already have documents describing their system, and those documents are not what this essay is about. The difference is mechanical, not stylistic.

Documentation describes: here is how the system works. A specification constrains: here is how the system must work. And the test for which one you have is brutally simple — a specification can fail a build.

Specification:    every payment command must be idempotent.
Generated code:   no idempotency key is honored.
Result:           verification fails. The change does not merge.

A first-class source artifact is version-controlled, diffable, reviewed in pull requests, authoritative, readable by both humans and machines, used to produce or validate the implementation, and — the load-bearing property — able to break the build when the implementation disagrees with it.

Documentation drifts because nothing happens when it is wrong. That is the whole difference, and it is not a tooling detail. It is the difference between an artifact and a comment.

Why Markdown, and where it breaks

Markdown is a strong candidate for the container, for unglamorous reasons: humans read it, models parse it, Git diffs it at sentence granularity so review becomes semantic, it is tool-independent, it is already where agents look, and it happily embeds other formats.

It is also genuinely bad at being a specification language. It has no formal semantics, no type system, no validation, and prose is vaguest exactly where engineers are most casual — "a reasonable timeout", "handle errors gracefully", "should be fast". Those phrases survive review because a human reader silently fills them in. An agent fills them in too, just differently, and without telling you.

So the realistic intent layer is Markdown as a container around things that are not Markdown:

prose               rationale, domain semantics, why the constraint exists
front matter        machine-readable attributes and thresholds
schemas             OpenAPI, JSON Schema, SQL DDL, Protobuf
examples            Gherkin scenarios, golden cases, edge cases
properties          invariants stated so a test can enforce them
policy              architecture rules, permissions, data handling

Markdown is the container. It is not the semantics. The thesis is intent-as-source, not Markdown maximalism.

What a specification-first repository looks like

/intent
  /payments
    overview.md          purpose, domain language, boundaries
    requirements.md      behavior the service must exhibit
    invariants.md        what must always be true
    api.md               interface and compatibility guarantees
    security.md          authz model, data handling, threat notes
    failure-modes.md     what happens when each dependency fails
  /identity
    authentication.md
    authorization.md
    threat-model.md

/contracts               the machine-checkable surface
  payment-api.yaml
  events.json
  schema.sql

/policy                  constraints on how anything may be built
  architecture.md
  dependencies.md
  data-handling.md

/tests                   the executable half of the specification
  acceptance/  contracts/  properties/  security/

/decisions
  ADR-001-event-driven-payments.md
  ADR-002-idempotency-strategy.md

/generated
  /services  /clients  /migrations

One directory in that tree carries the entire argument: /generated. In the strong form of this model, you should be able to delete it and rebuild it. Whether you can is the test of whether your intent layer is real. Today, for most systems, you cannot — and the gap between what the spec says and what the code actually does is precisely the knowledge that currently exists only in people's heads.

Here is what a specification dense enough to generate from starts to look like. Note how little of it a conventional README would contain.

# Transfer funds

## Purpose
Move money between two internal accounts while preserving ledger
consistency and preventing duplicate transactions.

## Preconditions
- Both accounts exist and are active.
- Amount is greater than zero and matches account currency.
- Source has sufficient available balance.

## Invariants
1. Money is never created or destroyed.
2. Debit and credit entries always balance.
3. One idempotency key never produces two transfers.
4. A failed transfer leaves no partial ledger state.

## Interface
POST /transfers
  body: sourceAccountId, destinationAccountId, amount, currency, idempotencyKey
  201 -> transfer created
  409 -> insufficient funds
  404 -> unknown or inactive account
  200 -> idempotency key already seen; returns the original result

## Failure behavior
- Any internal failure before commit mutates nothing.
- Ledger writes and event emission share a transactional boundary.

## Performance
p95 under 250 ms at 400 requests per second.

## Security
- Caller requires the transfers:create permission.
- Account ownership is verified server-side, never trusted from the request.
- Emitted events carry no account numbers or personal data.

## Observability
Events: transfer.started, transfer.completed, transfer.failed
Metrics: latency, failure rate, insufficient-funds rate

That document is simultaneously the input to implementation, to test generation, to code review, to the runbook, and to any future regeneration. A README is none of those things.

Three layers, and the one that decides everything

Three stacked layers. An intent layer of requirements, constraints, invariants, architecture policy and rationale, authored by humans and durable. Below it an implementation layer of generated application code, migrations and clients, treated as reproducible build output. A verification layer spanning both, holding contract tests, property tests, static analysis, policy checks, security checks and performance thresholds, which gates the merge. A weak verification layer turns generation into unreviewed change.

The intent layer is human-owned and durable. The implementation layer is generated and, in the strong form, disposable. The verification layer is not a step at the end — it spans both, and it is the only mechanism that converts probabilistic generation into something you are willing to deploy.

Traditional development compared with an intent-first pipeline. In the traditional path, requirements are interpreted by a developer who hand-authors source code that a compiler turns into running software, and the requirements are discarded along the way. In the intent-first path, a versioned specification is synthesised by a coding agent into an implementation which must then pass a verification layer of contract tests, property tests, policy and security checks; failing verification returns to generation and never merges.

An LLM is not a compiler, and the analogy breaks precisely here. Run the same specification twice and you get two different implementations. Reproducibility therefore cannot come from identical output. It has to come from equivalent verified behavior — which is a much weaker guarantee than a compiler gives you, and it has to be earned with machinery:

  • model and agent versions pinned, recorded, and changed deliberately
  • an approved dependency list, enforced in CI rather than in review comments
  • architecture policy checks that fail the build
  • contract tests, property tests, security scanning, performance thresholds
  • provenance metadata attached to every generated artifact
generated_by:
  model: <pinned model version>
  agent_version: 3.4.1
  specification_commit: 7b234af
  policy_commit: c91e02d
  generated_at: 2026-09-21T15:00:00Z

That last block is a supply-chain artifact. Once code is machine-produced at volume, "which model generated this, from which spec, under which policy" becomes a question auditors will eventually ask, and a question you should want to answer before they do.

Verification becomes the scarce resource

If implementation is cheap to produce, then confidence in the implementation is the thing you are actually short of. Tests stop being something that validates code and start being something that defines behavior.

For every accepted internal transfer:
  sum(balances before) == sum(balances after)

That property may well outlive every line of C# that currently satisfies it. It is the more valuable artifact, and it is four lines long.

This is also where the sharpest failure mode lives. A verification layer written by the same agent that wrote the implementation, from the same misreading of the same ambiguous sentence, proves nothing — a model cannot grade its own homework. The gate has to be deterministic and independently derived, which is the same discipline as a verifier-gated agent loop: generation is allowed to be probabilistic only because something non-probabilistic decides whether the result ships.

What happens to code review

Consider what a pull request looks like when the human-authored diff is the intent.

 ## Transfer limits
-Maximum transfer: $10,000
+Maximum transfer: $25,000

 ## Authorization
+Transfers above $10,000 require the transfers:approve-high-value permission.

Five lines changed. The generated diff underneath might be fourteen hundred. But the reviewable unit — the thing a human should be spending judgment on — is the behavior change and its justification, not the plumbing. The questions shift:

  • What behavior changed, and why?
  • Which constraints moved, and who is affected by the move?
  • Are the invariants still complete after this change?
  • Did verification actually prove the implementation satisfies them?

That is a better review than most teams get today, where a reviewer approves 1,438 lines of diff by skimming the interesting 40. It is also only as good as the verification layer underneath it. Without one, this is not semantic review — it is ceremony with a smaller diff.

Architecture stops being a document

Agents are extremely good at producing locally correct, globally incoherent systems. Every file is reasonable. The system is a swamp. This makes architecture more important in an intent-first model, not less — but only if architectural intent is expressed in a form that can reject a change.

Services may not read another service's database.
Inter-service communication is asynchronous and event-carried.
Every command accepts and honors an idempotency key.
Public API changes must remain backward compatible for two versions.
No dependency is added without an approved entry.
Personal data must never appear in application logs.

Every one of those is mechanically checkable. The architect's output shifts from a diagram that describes the intended system to a policy an agent cannot violate without failing CI — the same move as drawing an explicit permission map for an agent instead of hoping it behaves.

The counterarguments, taken seriously

"Natural language is ambiguous." Correct, and this is the right instinct. The answer is not unrestricted prose; it is the layered container above — prose for rationale, schemas and types and executable examples for anything that must be exact. Worth noting too: the ambiguity is not new. It currently lives in a developer's head instead of in a reviewable file, where nobody can diff it.

"Code is the only precise specification." This is the strongest objection, and it is partly right. But code is precise about what a system does, not about what it must do. It cannot distinguish an invariant from an accident. Every codebase contains behavior nobody intended and behavior nobody may ever change, and in the source they look identical. That distinction is exactly what the intent layer stores.

"Generated code still needs debugging." It does, and engineers will be reading implementation for many years. The question is not whether humans touch code. It is whether code remains the primary human-authored artifact.

"Models hallucinate." The core constraint, and the reason the model here is specification → generation → deterministic verification and never prompt → code → deploy. If you remove the third step, none of this works — you have just industrialised the production of plausible-looking systems.

"This just moves the debt." The objection I find most persuasive. If the specification is the source, then a stale specification is a production bug rather than a documentation bug, and specification debt compounds the same way code debt does. This implies a discipline — call it specification engineering — that almost nobody, including us, is currently good at.

"Regeneration is expensive." True. Regenerating a large system costs real money and real wall-clock time. The strong form of this model is economically absurd for most systems today, which is why the adoption path below stops well short of it.

"Our system can't be described that cleanly." Often true, and a legitimate stopping point. Legacy systems whose behavior nobody fully knows cannot be specified into existence. Distributed systems exhibit emergent behavior that no component-level specification predicts.

A spectrum, not a switch

LevelWhat it meansWhere teams are
0Code-first, no agentsrare now
1AI-assisted coding, prompt by promptcommon
2Agents implement from tickets or ad-hoc promptscommon
3Structured specifications drive implementation, with verification as the gateearly adopters
4Specifications plus tests regenerate whole componentsexperimental
5Implementation treated as reproducible build outputhypothetical

Most organizations we work with are somewhere around Level 1 to 2. The interesting engineering question is what Level 3 actually requires — because Level 3 pays for itself even if 4 and 5 never arrive. A specification precise enough to generate from is also a specification precise enough to onboard with, review against, and debug from at 3am.

An experiment worth a week

Pick a small internal service you already own and understand. Write its intent layer from scratch: overview, requirements, invariants, interface, security, failure modes, observability. Write the contracts. Write the acceptance and property tests. Then hand an agent only those artifacts — not the existing code — and ask it to implement the service.

Then measure the interesting things:

  1. What information was missing that you had to supply by hand?
  2. What did the agent assume, and were the assumptions reasonable?
  3. Which requirements turned out to be ambiguous only when something else read them?
  4. Which tests caught a wrong assumption, and which wrong assumptions slipped through?
  5. Could a second model produce a compatible implementation from the same artifacts?
  6. How much of the result did you have to hand-modify, and why?

The experiment is valuable when it fails, which it probably will. The list of things the agent got wrong is an exact inventory of what your current documentation omits — and it is the same list the next hire discovers in month three, and the same list an incident finds at 3am.

What is the actual asset?

One specification and its verification suite branching into four candidate implementations: the current C# service of twenty-five thousand lines, a Go implementation, a Rust implementation, and a runtime that does not exist yet. Every branch must pass the same contract tests, property tests and policy checks. The point is not language portability but that behavioural knowledge survives the implementation. Shown as a thought experiment, not a demonstrated production capability.

Suppose a service is 25,000 lines of C#, and suppose its behavior is genuinely captured by eighty pages of specification, three hundred contract tests, a hundred and twenty property tests, its schemas, and its policies. If an agent can produce a Go implementation that passes all of it, which of the two was the intellectual asset?

I do not think that question is settled. But notice that it is now askable, and that ten years ago it was not. The same reframing applies to what an organization should be investing in: domain knowledge, evaluation systems, operational history, and architectural policy are all harder to reproduce than application code, and they are the parts most teams currently keep the worst records of.

It also makes "source code" a slightly strange phrase. If the code is generated, the source is somewhere else — and we do not yet have good words for the layers:

intent source -> implementation source -> executable artifact

Where this leaves us

None of this is inevitable, and I want to be careful not to write as though it is. Models may plateau. Regeneration may stay too expensive. Specification debt may turn out to be worse than code debt. The honest position is that intent-first development is currently a plausible direction with a small amount of shipping tooling, a lot of commercial enthusiasm, and almost no production evidence at scale.

But the direction is worth acting on in its weak form today, because the weak form is just good engineering. Write down the invariants. Make the constraints machine-checkable. Treat the specification as something that can fail a build. Those things pay off whether or not the strong form ever arrives.

The most durable artifact in software may eventually be neither the code nor the binary, but the structured representation of human intent from which both can be recreated.

Some questions worth sitting with:

  • What percentage of your codebase is genuinely irreducible knowledge, and how much of it exists only because a human currently has to spell everything out?
  • What information would be required to regenerate your most important service from scratch?
  • Are your tests strong enough to distinguish a correct implementation from a merely plausible one?
  • If the specifications disappeared but the code survived, would you still understand why the system exists?

And the one I would start with:

If an agent could regenerate your entire codebase tomorrow, what would you wish you had preserved today?

Book a discovery call