Your team has the tools. Nobody has shown them what good looks like.

Three days, hands on, in your codebase. The team leaves with agents doing real work, a number to track it by, and the scaffolding still sitting in your repository.

[ask about a workshop]

The gap isn’t the budget.

Most organisations have already bought the licences. Usage is up, output is up, and nobody can say whether the work is getting better. The gap isn’t training spend and it isn’t effort — it’s that no one has told the team what good looks like, in their codebase, on their work.

It is also getting more expensive to leave. Habits set in the first few months of agent use, and they set whether or not anyone designs them. Correcting a practice is slower than forming one.

What actually happens.

  • Where the team actually is
    Day one

    We measure before we change anything. Intervention rate — how often a human had to step in before a task was done — on real tickets from your backlog. Most teams discover the number is not what they assumed, and that different squads are in completely different places.

  • Building the harness
    Day two

    The scaffolding, guardrails and feedback loops. Research and planning as written artefacts rather than chat. Specifications that compile into tests. Hooks that enforce quality deterministically, so nothing depends on anyone remembering. All of it in your repository, on your CI.

  • Trusting the output
    Day three

    What review is for when the agent wrote the code, and why clean code with passing tests can still be wrong. Turning each failure into a skill with tests behind it, so it cannot recur. The team leaves with the loop running and the habits installed.

Not a course. Not a keynote with exercises bolted on. Three days of your team’s real work, with the practice built while they do it.

What you have on Monday.

  • A number you can track

    Intervention rate, measured on your own work, with a baseline taken on day one and a way to keep measuring after we finish.

  • Working scaffolding in your repo

    Not slides. Plans, specs, hooks and tests that were written during the three days and are still there on Monday.

  • A shared vocabulary

    The whole team describing the same problems the same way, which is most of what makes a practice stick.

  • Habits formed early

    The patterns get set in the first few months of agent use. Setting them deliberately is far cheaper than correcting them later.

The questions you actually need answered.

  • How long is it, and can it be shorter?

    Three days is the full version, and it is the one I recommend, because day three is where the habits actually set. I also run a three-hour hands-on session and a half-day version for larger groups — those work well as a way to build the internal case before committing a team for three days.

  • Is it remote or in person?

    Both. In person is better for a single co-located team, because the side conversations are half the value. Remote works well for distributed teams and for running the same session across time zones.

  • Do you use our codebase, or examples?

    Yours. We work against your repository and your backlog, because the failures that matter are the ones your code produces, not the ones a sample project produces. A toy example teaches the idea and leaves the transfer to chance.

  • What do people need to know beforehand?

    Working engineers who can read and review code in your stack. No prior agent experience is required, and mixed levels are usually an advantage — the teams who are already ahead end up teaching the rest, which is exactly the transfer you want.

  • Who should attend?

    The engineers doing the work, plus whoever will own the practice afterwards. An engineering manager or staff engineer in the room makes the difference between a good three days and a change that survives.

  • Can it be customised?

    It has to be. The shape stays the same; the content depends on your stack, your review culture and where the team currently sits. We agree that on a call before anything is scheduled.

  • How do we know it worked?

    The baseline from day one. Intervention rate is the measure, and it is deliberately unflattering: it counts how often a human had to rescue the work. Velocity will look good whether or not anything improved, which is why it is the wrong measure.

Tell me where your team sits.

Team size, stack, and roughly how people are using AI today. I will tell you which version fits and what it would cost — and if a workshop is the wrong intervention, I will say so. The first conversation is free.

[write me a short note]

If you would rather read first: what actually works with AI coding tools, why AI assistance can cut coding mastery by 17%, and code review when the agent wrote the code.