The technical workshop. For people who already write skills and run them on a schedule, and now need the architecture, the gates and the evals that let something run when nobody is watching it. Four hours, or seven if you take the afternoon.
We grade AI capability on a five-rung ladder borrowed from martial arts — white, yellow, green, brown, black — and we walk every room through it in the first five minutes, because the honest answer decides which day is worth your money. This page is the top end of that ladder. If your team is anywhere left of brown, the four-hour workshop is a better use of the day, and a cheaper way to find out.
The entry test is brown, and we hold it. If your team has not written and scheduled skills, day one here will be spent teaching what the four-hour workshop teaches better — and you will pay this workshop's rate for it. We would rather tell you that now than at nine in the morning.
Getting one agent to run is a weekend. Getting a fleet of them to run safely, survive a model change, and be owned by somebody with a name is the actual discipline. Here is how we grade it, and where this day puts you.
It watches something real, it drafts, and it stops at a gate for a person. Every run is logged, every action is reversible, and one named human owns it. Not a demo on a laptop — a thing that ran last night.
Several agents across more than one function, each with scoped credentials, an eval suite built from real failures, and a deploy you can roll back. You can change one without holding your breath about the others.
Digital labor sits on the org chart with a budget and an owner. New processes get designed for agents rather than retrofitted onto them, and the question stops being whether to automate and becomes who approves it.
Every one of these is the thing that breaks a working prototype three weeks later. We teach them in the order you hit them, on your own systems, with your own credentials, in your own repository.
The loop, the tools, the stopping conditions. Why most things called agents are a prompt in a while-loop, exactly how that fails, and what separates one that recovers from one that spirals.
Writing skills that hold when someone else runs them, versioning them, and a shared library the team actually maintains — plus the judgment call of when a job is a skill and when it needs to become an agent.
What goes in the window, what gets retrieved, what gets compacted, what gets remembered between runs. The difference between an agent that stays sharp for four hours and one that forgets its own instructions.
Tokens are the meter that is always running. Caching, compaction, budgets per run, and knowing your cost per task to the cent — before finance asks, not after.
The frontier model where judgment matters, the small fast one for mechanical steps. Routing between them, falling back when one is down, and upgrading when a new one ships — with your evals as the safety net.
Connecting the system that has no integration and never will. You write one against your own internal API in the room — resources, tools, auth — and point an agent at it.
The single largest determinant of whether an agent works, and the one nobody spends time on. Naming, schemas, defaults, and writing error messages the model can actually act on rather than apologize about.
Where the gates go and what an approval actually locks. Making "it drafted it and waited" the default, so the interesting question becomes which gates you are confident enough to remove.
How you know it still works after the model changes underneath you. Building a small suite out of three real failures from your own logs, and wiring it so a regression fails loudly instead of quietly.
Least privilege per agent, secret handling, audit trails, the accounts it must never reach — and the prompt-injection attacks your inputs will eventually carry. Written down, scoped, and demonstrated.
Traces, retries, failure modes and alerts. What to watch, what to ignore, and how to tell a bad day from a broken system when nobody was in the loop overnight.
Scheduled, triggered and on demand. Version control for prompts and skills, staging versus production, and rolling back at nine on a Sunday night without a heroic effort.
The first four hours are the $5,000 workshop and end with one agent in production. The afternoon is a $2,500 add-on — $7,500 all-in — and builds the platform. Hands on keyboards throughout: one machine per person, your stack, not ours.
| Time | What we do | Why it is there |
|---|---|---|
| 0:00 – 0:15 | Belt check | What is already running, what broke last time, and which process we are putting an agent on today. Chosen before the day, confirmed in the room. |
| 0:15 – 1:00 | Tools and MCP | Write an MCP server against one of your systems. This is where the capability stops being generic — an agent is only as good as what it can reach. |
| 1:00 – 1:45 | The loop | Build the agent: what it watches, what it drafts, where it stops. Then break it deliberately and watch how it recovers. |
| 1:45 – 2:00 | Break | The one where the security questions get asked properly. |
| 2:00 – 2:40 | Gates, permissions, secrets | Scope its credentials down to the minimum, put the human gate in the right place, and prove what it cannot reach. |
| 2:40 – 3:20 | Evals from real failures | Take three things it got wrong this morning and turn them into a suite that runs on every change. This is the habit that outlives the workshop. |
| 3:20 – 4:00 | Ship it | Deploy, schedule, wire up traces and alerts, and roll it back once so you have done it before you need to. |
|
The four-hour workshop ends here — one agent in
production, owned and reversible.$5,000 · per company
| ||
|
The afternoon — the platform.
+$2,500 add-on · $7,500 all-in
| ||
| 4:00 – 4:20 | Lunch, and the pick | Platform or internal application — chosen at booking, confirmed now that the morning has shown where the friction actually is. |
| 4:20 – 5:30 | The substrate | Registry, shared auth, secrets and audit — or the application's real schema and real authentication. The boring part that makes the next ten quick. |
| 5:30 – 6:30 | Build on top | Two agents running on the substrate, or agents wired into the application — with the shared eval harness running in CI. |
| 6:30 – 7:00 | The handoff | The pattern written down, owners named per agent, and the rollback rehearsed one more time on the day's real deploy. |
Chosen when you book. The morning proves one agent can run. The afternoon is about the tenth one taking a day instead of a month.
The substrate underneath the fleet. An MCP server against your core system of record, a shared place agents and skills live, common guardrails, and one eval harness they all run through. Boring, and the reason the next ten are quick.
Some of what your team wants is not an agent — it is the tool nobody has had six weeks to build. Real auth, real data, deployed where your team owns it, in your repository with your pipeline. Built in the afternoon, in front of you.
Both options end the same way: something running that you did not have at breakfast.
This one has prerequisites, and they are not negotiable. A technical day spent waiting on a credential is an expensive way to do nothing.
Confirmed in the two-hour engineering session three to five days ahead, which is included.
Up to ten seats. The mix matters more than the count — a room of only engineers builds the wrong thing very efficiently.
Per company, on site, for your technical team — same price as the four-hour workshop, and most companies eventually buy both, in that order.
| 4 hours $5,000 | 7 hours $7,500 | |
|---|---|---|
| Two hours of AI engineering before the day | Included | Included |
| MCP server written against one of your systems | Included | Included |
| Agent built, gated and deployed to your environment | Included | Included |
| Eval suite built from your own failures | Included | Included |
| Permissions and secrets scoped in writing | Included | Included |
| Everything committed to your repository | Included | Included |
| Afternoon build — platform or application | — | Included |
| Shared eval harness and CI job | — | Included |
| Fractional AI engineering after the day | On request · one day a week upward | |
Add the three-hour afternoon for $2,500 at checkout, for $7,500 all-in. Card, bank transfer, Affirm and Klarna accepted.
Not sure yet? The note below gets you a straight answer first, at no charge.