Decision Engineering

What we engineer when the code writes itself.

TLDR: AI made code cheap, so specification became the bottleneck, and specification works best as several small, structured views of a system.

The Decisions Moved

If you've built anything with a coding agent lately, you know where the time goes.
Not typing - no, the agent types faster than any of us.
The time goes into explaining. What the feature should do, what it shouldn't. What we decided last week, and why.
Then correcting the LLM when it guesses wrong.
Then explaining everything again next session, because it forgot.

Typing code used to be where a lot of decisions got made. Product may have decided the feature direction in advance, but every line we wrote also proposed and settled more questions: what happens on a timeout, which value wins when two sources disagree, what does an empty list look like. We made those calls at the keyboard, as we typed. The typing gave us time to reflect, time to knead the clay as we poked and prodded at the problem until the shape of the solution emerged. Now the typing happens elsewhere, but those answers still have to exist before the code does.

The code was a sprawling repository of all our thousand different decisions. It wasn’t always an obvious place for them, but it was what we had, and we found ways to make it work. Conventions, comments, naming, careful stewardship of this crystallized decision lattice. We’d warn new devs about the most fragile parts, and create system upon system to help keep everything consistent - protecting our future selves from our past, knowing that we’d forget something important six months from now and guessing how to best remind ourselves before we accidentally trampled the complex ecosystem we oversaw.

But that’s not what code is for anymore. Code just runs.

Whatever we may lose in this new paradigm, we do also gain an opportunity. Now that code no longer serves as that final decision point, we’re free to adjust our process to work more organically with our sources of truth and their representations. So - how do we want to work? How will we make all those tiny decisions, day in and day out?

Spec Driven Development: One View Isn’t Enough

There’s a lot of buzz lately around spec-driven development. It makes sense - we need to tell the LLMs what to build. But there's an old joke that a sufficiently detailed specification is called code.
When we try to write out the full details of a system using specs, haven’t we essentially created another programming language? And wait, the LLMs are writing the code… so are they writing these specs too? How do we spec the specs? Help, I’m trapped in a loop! Someone throw me a ++!

We're already feeling our way toward this.
CLAUDE.md, AGENTS.md, ARCHITECTURE.md, cursor rules: we're all writing down the system so the agent stops forgetting it.
But these files grow into a pile.
Architecture, conventions, gotchas, and last week's decision all share one document that quickly becomes unwieldy.

Spec-driven development is more organized, but it asks humans for just one viewpoint (behavior - essentially what we used to call BDD) and derives everything else from it. But the decisions we used to make at the keyboard aren't all behavioral. Which value wins when two sources disagree is an information question. What happens on a timeout is a functional one. Where it runs is a deployment one.

It should tell us something that we already have a ton of ways to describe systems, and we’ve been using them all along. We sketch a state diagram on a whiteboard to settle a disagreement, draw boxes and arrows to onboard a new hire, write an ADR when a decision feels big enough. Sequence diagrams, wireframes, entity diagrams, schemas, jobs to be done, user stories, and dozens more. These evolved for a reason - each solved a specific problem in a specific situation, that the others couldn’t.

But we never lingered. Nobody spent a day on a state machine or a week on a physical architecture diagram - we needed the rough directional answer quickly so we could get back to the code, where the "real work" happened. The artifacts were just scaffolding, and we tore them down as soon as the code was standing. But now the code builds itself, and the scaffolding is where the work is. Any approach that collapses it all into a single format to appease an LLM is throwing the baby out with the bathwater.

What we need is a way to organize this information; a more structured approach to these tools that have historically been very helpful but also very ad hoc. And, we need a methodical way to turn these disparate artifacts into a cohesive codebase.

We need a fundamentally new approach to software development.

The Structural Solution: Multiple Viewpoints

I have great news: there’s been decades of work on this very topic and it’s ready to go.

If there's a gopher in your cornfield and it's behind a rock, the best binoculars in the world won't find it. But walk around to the other side and you don’t need binoculars - the gopher is immediately obvious, it’s just sitting there next to the rock. The same idea applies to system descriptions; there’s no one diagram to rule them all.

In 1995, Philippe Kruchten's 4+1 strategy described software architecture through four views (logical, process, development, physical), plus scenarios that cut across all four to show how they work together. Rozanski and Woods later turned the idea into a full method in Software Systems Architecture. These methods share a core insight: no diagram can show everything, nor should it.

Rozanski and Woods' argument starts with people. A system has many stakeholders (users, developers, testers, operators, buyers), and each cares about different things. One artifact that tries to answer all of them becomes too complex to understand and answers none of them well. So you split the description into views, each built from a viewpoint: a reusable template naming the concerns it covers, who it's for, and how to describe it.

The concepts underneath, which come from IEEE 1471 and are now standardized as ISO/IEC/IEEE 42010, are worth borrowing.

  • A viewpoint is a reusable template for one set of concerns; Rozanski and Woods' include Functional, Information, Deployment, and Operational.
  • A view is your system described from one viewpoint, e.g. an “Information View”.
  • A view is made up of artifacts, individual documents like a schema or a state machine, and each artifact is an instance of an artifact type template with its own conventions.
  • The same artifact type can appear in several views: a state machine for a record's lifecycle belongs in the Information view, while one for a background job belongs in the Functional view.
  • A perspective is a cross-cutting concern like security or performance.

NB: ISO 42010 calls them models instead of artifacts, but I’ll use artifact to disambiguate from large language models

The structure here isn't mainly for the LLM, since they read prose just fine. It's for us. It makes gaps visible while we write (a decision table shows you its empty cells). It makes it obvious what information belongs where, giving every idea a home.

If this all feels very familiar, that’s because it’s all derived from our existing experience as developers. Code isn’t for the compiler, it’s for us. Separation of concerns, DRY, coupling, careful object design, clear interfaces - those are for us too. And these code concepts translate directly into our system descriptions.

A clear set of artifacts can be combined into a strong view, with good information density, consistent internal checks we can assert automatically, and a solid basis for gap detection. A muddy set of artifacts can only be combined into a mush and a prayer. Only one of them can effectively describe a substantial system.

The Generative Solution: LLM Workflows

It's all well and good to create a variety of artifacts, but eventually they have to turn into code.
Lucky for us that LLMs are great at taking dense, structured, slightly ambiguous inputs and turning them into code!

The workflow has four parts: design, build, steering, and feedback.
All of them rest on one prerequisite: a shared catalog of viewpoints and artifact types, so everyone, human or LLM, knows what each artifact is for and what a complete one contains.

Design - Decisions We Make On Purpose

Humans write artifacts, and we don't have to write them perfectly.
Sketch a state machine, jot down a scenario, spitball a wireframe.
Let an LLM snap it into place, cleaning it up against the artifact type template the way a whiteboard app turns a wobbly square into a rectangle.
It suggests the canonical form of what we wrote, points out slots the artifact type expects but we left empty, and flags information that belongs in a different artifact entirely. The artifact type template helps us detect internal inconsistencies in the artifact early.
Then the LLM reviews the artifacts against each other, infers how everything connects, and highlights gaps in the overall system.

Say our checkout sequence diagram reserves stock as soon as an order is placed, so we never oversell. Our order state machine says that if payment fails, we retry for 24 hours and then cancel. Both are perfectly reasonable on their own. But walk the "payment fails" scenario through both, and something's missing: nobody said when the reserved stock gets released. Unpaid orders sit on inventory for a full day, and a popular item shows "sold out" while real customers get turned away. Maybe that stock never comes back at all. Neither artifact is wrong. We just never decided whether a failed payment should get to keep holding stock, and that's a business call, not a technical one.

Every artifact has its strengths and weaknesses that make it uniquely useful in certain situations. The traditional Swiss cheese approach to development takes advantage of this to help us get code out the door quickly and safely. The unit test catches the bug the type checker missed, and vice versa. We use both because we could never make a type checker powerful enough to catch every logic error, because that’s not what it’s for.

Multiple artifacts across multiple views is Swiss cheese development for the AI era.

Build - Decisions We Oversee

Like any build, this one produces intermediate artifacts along the way: maybe a task breakdown or a requirements list, definitely tests.
Nobody edits these by hand (see MDA below), but we keep them, because the intermediates help us understand and debug the build process.

When the LLM hits something the description doesn't cover, it doesn't guess silently.
It makes a reasonable call and leaves a receipt: "Assumed lists are private to one user.”
Receipts go into a queue, and as we review them, we adjust the system description if and when appropriate so the next build makes more informed decisions.

Because code is cheap, we can build as many times as we need to.
Generate a few versions under different assumptions. Try a silly one if you want!
Ultimately pick the one that looks best for now, knowing anything underspecified might change in a future build anyway.

Day to day, though, nobody regenerates the world.
A PR starts with a small design change: a new transition, a revised policy, iterated to stability.
The LLM reads the diff and makes delta changes to the code.
The commit holds both the design change a human made and the code change the LLM made, plus perhaps the build artifacts that connect them.
Reviewers read the intent, and automation confirms the code matches.

Steering - Decisions On The Fly

I couldn't ship an app today without significant prompting, and I don't expect that to disappear overnight.
The question is where the steering lives.

Stewart Brand made a point about buildings in How Buildings Learn: the site almost never changes, the structure rarely, the floor plan often, the furniture constantly.
System descriptions work the same way.
Architecture views change slowly. Feature specs change per feature. Examples change per bug.
Below those sits a fast, local layer: directives.
"Debounce search by 300 milliseconds."
"Cluster map pins above 50."
Most of the prompting we do today belongs here, and that's fine.
Directives live in the design repo next to the views, scoped to a feature or module, each with a one-line reason, and every build reads them.
Regeneration respects them instead of clobbering them.

When something's wrong, fix it at the highest layer where the fix is true.
When several directives say the same thing, that's a missing policy, so promote it.
When one stops applying, delete it.
Occasionally you'll want to freeze code outright, like a function you've reviewed and don't want regenerated.
Pin it, give it a reason, and treat it like vendored code.
An honest exception beats pretending there are none.

Feedback - Decisions Surfaced By Reality

It’s the joy and the bane of software - users take paths through the product no one could have imagined. Production surfaces bugs, new use cases, new opportunities. Without artifacts, LLMs watching production can only look for anomalies. With artifacts, LLMs have vastly more understanding of the intention of our systems. They know what’s illegal, what to look for, what’s important - without having to infer it from the code. Anomalies can be automatically classified into bugs (code diverges from system description) or gaps (the system description never covered this).

As always, these simply flow back into the queue to be designed and built against according to their importance and urgency.

A lot of that is pretty similar to what's already happening.
The biggest differences are toward the top.

But Wait...

Isn’t this just more fragmentation? Knowledge is already scattered across code, Jira, Figma, wikis, and Slack.

True, and plenty of people are working on pulling requirements out of slack! But what these channels lack right now is structure, boundaries, and adherence. A Jira ticket can hold anything, so it ends up holding everything, and nobody can tell what's really there or what’s missing.

The value of a view isn't that it automatically agrees with the others. Nothing about views forces agreement. The value is that each one is specific. A specific view holds more of its kind of information, more comfortably, and its constraints make gaps and disagreement noticeable. Constraints don't guarantee agreement; they encourage it, by making disagreement visible.

Haven’t we tried this?

Yes.
Model-Driven Architecture, CASE tools, and UML round-tripping all generated code from diagrams, and they mostly failed.
The usual story: anything the diagrams couldn't express got hand-edited into the generated code, regeneration clobbered the edits, and teams fell back to treating the code as the truth.
The diagrams rotted.

Architecture description has its own version of this story.
IEEE 1471 standardized viewpoints and views in 2000, and its successor, ISO/IEC/IEEE 42010, refined them into the vocabulary this piece uses.
Few teams outside enterprise architecture ever adopted it.
Not because the ideas were wrong; the descriptions just never had a job.

Now they do, because three things are different today.

  • First, rigid generators needed complete specs, while LLMs tolerate incompleteness by filling the long tail with sensible guesses.
With receipts, those guesses become visible.
  • Second, connecting artifacts lazily removes the maintenance overhead that made documentation a full-time job.
  • Third, and most important, the escape hatch moves.
In MDA, underspecified meant custom: when the diagram couldn't say something, you edited the output.
Here, you change an input and re-run the generation.

When the built system is structurally predicated on the design artifacts, the design artifacts become load-bearing and the associated processes and cadences keep them up to date.

Does this work?

Let me know :) I’ve had good preliminary results and I’ll report back soon with more details. Worst case, where code generation never becomes fully view-driven, the views are still far better context for today's coding agents than a sprawling instructions file or a pile of tickets.
You can start small, with a few artifacts for the parts of the system that matter most.

I’m still iterating on this myself, and I think it will take some experimentation for individuals, teams, companies, and the industry to settle on a path.

The New Inner Loop

The exciting thing about this approach is that so many artifact types already exist, and many even have tooling.
E.g. there are VS Code extensions for draw.io, so we can already create machine-readable diagrams inside common editors.
How many of us are doing it? I didn't even know it existed until recently.
Templates and tooling may evolve to support this workflow. Today, a shared catalog and a handful of Claude skills can get us started.

Put it all together and the developer's inner loop changes. It used to be plan first, then type, build, test, repeat. Now it's write loosely, snap, critique, decide, generate with receipts, and observe, then back into the views. Humans and LLMs work together at every level, checking each other; Swiss cheese all the way down.

The goal isn't more AI in the development process. It's making more of a system's knowledge explicit, checkable, and reusable, and making that cheap enough that we actually do it. We'll never finish the spec, but we can build an evolving process that keeps finding what's missing.

Developers have always held a model of the system in their heads. We just get to interact with it directly now.