# How I work with AI

> The loop, what the model sees, what gets reviewed, what the agents can touch, and what happens when a run dies halfway.

By Abhishek Kolge. Canonical page: https://abhishekkolge.dev/ai-workflow

## The loop

### Plan first, then diff

The agent writes the plan before it touches a file. I approve it, then it implements against that. If the plan and the diff disagree, the diff is wrong.

### I stay on the reading side

On a big change I set the boundaries, workers make the edits, and I check the result against the repo myself. Architecture and the final merge stay with me.

### A goal needs an end state and a budget

A run that continues on its own gets a bounded objective, the command that proves it’s met, and a token ceiling. When it hits the ceiling it stops and shows me what it has instead of carrying on quietly.

### Fan out only where work is disjoint

I give each package or feature one worker in its own checkout, so two never edit the same file. Anything with ordering stays a single thread. Every extra worker is another full model session being paid for.

### Code instead of a dozen tool calls

When a step needs five lookups and some arithmetic, the agent writes one sandboxed function that calls the tools and returns the answer. The filtering and the sums happen in real code, so the totals come back right.

### Strong model plans, cheap model finishes

The expensive model investigates, names the risks and makes the first edit. The repetitive rest goes to a faster one in the same session. I don’t split it this way on security work, or where every step is still a design decision.

### Branch instead of arguing

Correcting a session that has gone down a bad path costs more than going back. I return to the turn where it went wrong, rewrite that instruction and take a new branch. The old one stays in case it turns out to have been right.

### Workflow where the path is known, agent where it isn’t

A known sequence gets written as explicit steps with typed inputs and outputs, so it replays and I can see which step failed. An agent loop is for parts where the next move depends on what the last one found. Most production work I see is the first kind, even when it gets pitched as the second.

### Done means verified

A task is finished when I’ve checked its deliverables against the current repo. The agent saying so doesn’t count, so I ask for the command output instead of a summary.

### Risky attempts get their own branch

Abandoning one then costs nothing. Some of what I built this way ended up reverted, and each of those dead ends cost me a branch instead of a rollback.

## What the model gets to see

### Most bad output is a bad context window

When an agent goes sideways, I look at what was in front of it before I blame the model. Usually it had the wrong file, a stale convention or instructions that contradicted each other. Fixing the input fixes more than swapping the model.

### Each kind of context has its own channel

Standing rules go in a file the agent always loads. Facts my code already has get interpolated into the request. Anything that changes mid-run comes through a tool call, so the agent never works from a paragraph written an hour ago that is now wrong.

### Keep the front of the window still

Instructions rewritten every turn throw away the cached prefix, and the whole window gets paid for again. The stable part goes first and stays put. Anything that moves during the run arrives later, as a message or a tool result.

### The rules live in the repo

I commit conventions, contracts that must not drift and playbooks for the parts that are easy to get subtly wrong next to the code, and they load automatically. In a monorepo the nearest file wins. When an agent learns something the hard way, it goes in the file.

### Procedures load when they are needed

Playbooks for the fiddly parts live in the repo as skills the agent can search and open on demand, rather than pasted into every system prompt. Base instructions stay short and the agent still knows where to look.

### Retrieve narrow, then read wide

I use search and symbol lookup to find the three files that matter, then read those properly. Dumping a directory into the window pushes out the part that mattered. Semantic search over old conversations costs an embedding call on every turn, so it stays off until a feature needs to remember last week.

### Trim what tools put back

A tool returning forty rows of JSON per call fills the window with things nobody reads again. Verbose results get summarised to the fields the model needs before they land in the transcript, and stale ones get dropped.

### Hand off on purpose

At a phase change a session becomes a brief: goal, what’s done, what was decided, constraints, what was tried and rejected, next step. Recent turns stay verbatim behind it. Letting a window compact itself the moment it fills is how a rejected approach comes back an hour later.

### Long transcripts become a log of observations

On work that runs for days the transcript gets compressed in the background into a dense log of observations, recent turns still intact in front of it. That keeps the order of events readable, which a single flattened summary written under pressure loses.

### A worker inherits what I choose

A delegated task gets the brief and the files it needs, not the whole parent conversation. Everything that leaks down is context the worker has to read past, and half of it is about something else.

### When memory and the repo disagree

Cross-session memory is scoped per project and holds decisions and preferences, nothing sensitive. When it disagrees with the working tree, the working tree wins, and I check anything it claims about a file before acting on it.

### Retrieval quality is a product decision

On retrieval features, chunking, metadata filters and reranking have moved answer quality further than any prompt rewrite did. Filter by what you already know about the user before you ask a model to be clever.

## Reviews, tests and traces

### A second model watches while the first works

On a high-stakes change a separate model reads along as the session runs and can flag a missed requirement or a dangerous API before the work is finished. Small notes stay out of the way. A real risk steers the run, and something broken stops it.

### A reviewer pass before I read the diff

Every risky change gets a read-only review that reports findings and changes nothing. The failure modes that keep coming back are written down in the repo, so the reviewer looks for them by name instead of rediscovering them.

### A review has to name a defect

Review output has to name a provable defect, where it triggers and what it costs, with a priority and a confidence. Lock files, generated code and style preferences are excluded up front, because nobody finishes reading a report full of noise.

### Security gets its own pass

Vulnerability review runs as a separate scoped, read-only sweep rather than a line in a general review. Each finding ends up marked fixed, accepted or false positive, so the same one doesn’t get re-argued next month.

### LLM features ship with tests

An LLM feature ships with a test set and a metric, the way an API ships with tests. For anything that generates a query or an action, that means accuracy measured on execution against real data, re-run every time the prompts move.

### Scores have to be able to block

Some checks are gates and have to pass outright, like calling the right tool without a tool error. The rest have thresholds pointed the right way, because hallucination and toxicity should stay under a ceiling rather than clear a floor. The checks produce one verdict, and that verdict can block the merge.

### Whole conversations get tested

Plenty of failures only appear on the fourth turn, when the agent has forgotten the first or reaches for tools in the wrong order. Those run as a sequence against one thread with assertions per turn. A single score for the whole chat hides the turn that broke.

### Every failure becomes a case

Eval cases live in a versioned dataset, so a run can be reproduced against the exact set it was scored on. Anything that goes wrong in production gets added before the fix ships, so the same mistake can’t slip through twice.

### Every run is traceable and priced

Prompts, tool calls, tokens, cost and latency are recorded per run, and a trace id comes back with the response, so the feedback button and the complaint land on the exact run. Sensitive fields are redacted on the way out, and production samples a slice, because tracing every call costs money too.

### One commit per idea

Work is split into atomic, signed commits, each readable on its own. Pre-commit checks run for real, and a refused hook stops the commit instead of being worked around.

## What the agents get, and what stops them

### Tiered permissions, no static keys

Reads are free, and writes and commands are gated. The dangerous ones get blocked in code before they run. Nothing an agent touches holds a long-lived cloud credential.

### Delegation is authority, so scope it

A background worker can’t stop and ask me a question, so anything needing an answer fails there instead of waiting. I want it that way round, and it’s why a worker gets a narrow task and a narrow permission set.

### Internal systems get typed tools

Where an agent reaches an internal system it gets a tool with a schema and a permission boundary, so the call is validated before it runs. I trust a tool that can only do the safe thing more than a prompt asking the model to behave.

### A human approves the irreversible ones

When a tool would delete records, spend money or send anything to a customer, it pauses with its arguments on screen and waits. Approval is bound to those exact arguments and that one call, so approving it once doesn’t approve the rest of the session.

### Anything the model reads is untrusted

I treat documents, web pages and ticket comments as data, and nothing in them counts as an instruction. Untrusted text gets checked for injection and stripped of anything that reads like a command, and a tool call that arrives out of nowhere doesn’t get to run because the text asked for it.

### Nothing sensitive leaves without a reason

Personal data is redacted on the way into a prompt and scanned for on the way out. Secrets never sit in a context window at all. On regulated data I build that boundary in from the start instead of adding it in a hardening pass at the end.

### Structured output, then validate it

Anything a system consumes answers in a schema, parsed and validated before it touches a database or a downstream call. When validation fails I would rather the call fail loudly than let a half-parsed object travel.

### Every guard declares what it does when it trips

Each check says up front whether it blocks, redacts, rewrites, or notes the problem and carries on, so nobody decides in the moment. A blocked call comes back marked as blocked, and the violation is logged either way.

### Spend limits stop the run

Token and cost limits are enforced inside the loop, so a run going in circles is stopped rather than reported at the end of the month. Retries are bounded too, because a third identical attempt against a broken service only burns money.

### Language server first, text edits last

A rename goes through the language server, so imports follow the symbol rather than the spelling. Shape changes go through the syntax tree, previewed against current files. I read the match count first, because matching shape is not the same as knowing which function I meant.

### Stop the known mistake mid-sentence

Some errors are recognisable while they are being written, like reaching for a helper removed a release ago. Those get caught as they happen and the agent is corrected and restarted, which is cheaper than reviewing the finished version of a wrong edit.

### Hooks fail closed

The checks between a proposed command and a real one run locally and block by default when they error. A guard that lets the call through when it breaks protects nothing.

### A model per job, and no lock-in

A strong model plans, a cheaper one does the routine typing, a small one handles utility work. Every provider is swappable, and anything sensitive can run against a local model.

## When the work runs long

### Long runs save their state and resume

Anything waiting on a person or an external system saves its state and resumes from there. A crash, or an approval that comes back the next morning, costs only the time spent waiting.

### New input wakes an idle run

New input to an idle run arrives on the same thread and either wakes it or waits for its next turn, so it continues instead of rebuilding its context from nothing. State that keeps changing gets its own lane rather than being re-pasted every turn.

### Recovery replays, so tools are idempotent

A resumed run picks up from the last checkpoint, which can mean a call being made twice. Every tool that writes carries a key and is safe to repeat. Without that, a replay can charge someone twice.

### A failed run says where it failed

A run comes back with which step failed and what it was holding. Transient errors get a bounded retry, and a step with nothing to do exits cleanly instead of pretending to work. The client can drop and reconnect without losing the stream.

## Where I don’t use them

### The rule that keeps someone safe

Where a wrong answer can hurt someone, the rule is plain code. A model may add to the list of things to avoid, but it can’t take anything off. I tried giving a second model a veto over that kind of check, and it spent the veto on risks that were not there.

### The number a client signs off on

A model can read the documents and pull the figures out, but the categories it can pick from are a fixed list in its schema, a person approves the result before it counts, and the calculation that follows runs with no model involved.

### Debugging the running system

Asking a model why production is broken gets you a plausible theory. I would rather attach a debugger, stop on the line that matters and read the real values, or go measure the timeouts, locks and load myself.

### Deciding what to build

A model is good at the fifty ways to implement something and no help at all on which one the business needs. That conversation stays with the people paying for it.
