Greenfield, brownfield, and where I work with AI
How I decide what to delegate to AI across greenfield, brownfield and harnessed systems.
I’ve been trying to put a clearer shape around how I actually work with AI.
Not which tools I use, or how many agents I can have running at once. More where I put myself in the development process, what I’m comfortable handing over, and what has to be true before I do.
I think there are three slightly different modes.
Greenfield: explore first
On a greenfield project I’m usually trying to understand the problem before I’m trying to build the system.
That might mean talking to customers, looking at how something is currently done, making rough prototypes, changing my mind, planning a bit, implementing something, putting it in front of somebody and looping back around.
It isn’t particularly linear.
Understand → shape → build → learn → repeat.
AI fits quite naturally into this because there are usually lots of reasonably isolated things I can send away.
I can shape a feature, give Claude or Codex enough context to work on it in the cloud, then move onto something else. I might be exploring another part of the product, talking to a customer, or working through the next problem while that implementation is happening.
The important bit is that I’m not trying to remove myself from the loop.
When something needs more precision, I’m back in the IDE. That’s where I’m reviewing the implementation, tracing something properly, making architectural decisions, looking at security, understanding why something behaves the way it does or getting another agent to critique what has just been built.
The cloud is useful for throughput.
The IDE is where I tend to slow down.
Brownfield: diagnose first
Brownfield work feels different.
Walking into an existing system and immediately asking an agent to start changing things feels backwards to me.
My first job is normally diagnostic.
What is this system doing? Where are the boundaries? What assumptions have accumulated? What talks to what? Where is the data? What breaks if this changes? What did the original developers know that isn't written down anywhere?
Only once I understand enough of that can I decide what I’m comfortable delegating.
And quite often, the answer is that the work needs a better harness around it first.
By harness, I mean the layer that turns intent into bounded execution.
In my own setup, Claude routines take my direction, turn it into tickets, work through them and return the results for review. The harness coordinates that loop: what needs doing, what context is available, what can be changed, what state needs preserving and where I need to come back in.
Tests, fixtures, types, seed data, logging, evals, feature flags, documentation, acceptance criteria and staging environments aren't necessarily the harness by themselves. They are the constraints and feedback the harness can use to keep the work bounded and tell us whether it has done the right thing.
So the loop is slightly different:
Diagnose → map → harness → change → verify.
Once that exists, the workflow starts looking much more like the greenfield one. Work can move into the cloud, get checked against the harness, come back to me for careful review, then get merged and observed.
Eventually greenfield becomes brownfield
This is probably the bit I hadn't properly articulated before.
A greenfield project doesn't stay greenfield.
The thing I'm happily changing very quickly in week one becomes a system with customers, data, integrations and consequences.
At some point it needs the same treatment.
The prototype gets tests. The fuzzy architectural decisions become explicit boundaries. Important behaviour gets documented. Evals appear around AI behaviour. Deployments get safer. Observability improves.
The project becomes harnessed: not just wrapped in an agent workflow, but equipped with enough constraints and feedback for that workflow to be dependable.
And once that happens, something interesting changes in how I work.
I can give agents more autonomy because I'm less dependent on the agent understanding everything perfectly.
The system itself can tell us when something is wrong.
So perhaps there are really three modes:
Explore when we don't yet know what should exist.
Diagnose when something already exists and we need to understand it.
Operate once we've built enough constraints and feedback around the system to move quickly without being careless.
All three still involve me writing code. All three still involve AI.
The balance just moves around.
More bounded implementation can happen away from me. More of my time goes into intent, product judgement, architecture, reviewing, debugging and deciding what deserves to become part of the harness next.
I'm finding that this distinction matters more to me than whether something is “AI-written” or “human-written”.
The question I keep coming back to is simpler:
What would need to be true for me to safely hand this piece of work away?
Sometimes the answer is a good prompt.
Sometimes it’s a test.
Sometimes it’s three hours understanding the system first.
And sometimes the answer is that I should just open the IDE and do it myself.