Fifteen statements about how your delivery organization actually runs. Tick the ones that are true today. It takes about three minutes, and the arithmetic at the bottom tells you whether to start a conversation with us this year or spend the next two quarters on something else.
Nothing on this page is submitted anywhere. There's no form, no email box and no follow-up. We can't see your answers and wouldn't know you had been here.
Tick what is true today, on evidence you could put in front of someone this afternoon. An in-flight program that will make it true next quarter counts as untrue. The four groups are different sizes because the risks in them are different sizes.
Agents raise the volume of change moving through your pipeline. Whatever is already wrong with that pipeline gets louder.
If builds are red more often than green, agents generate failures faster. This is the single most common reason we tell an organization to wait.
Where each release needs a human escort, extra throughput just queues in front of the escort. Your change board sets the ceiling long before the tooling does.
Agents iterate. A ninety-minute pipeline turns iteration into a batch job and most of the gain drains away into waiting.
Generated code needs a test seam to be safe. But quality automation is at its weakest on legacy code that has none, and it's better to say so now.
This is less about memory than about whether root cause gets written down anywhere. An operations agent has to read something.
Reversibility is what makes agent-authored changes tolerable in a regulated shop. A rollback runbook nobody has exercised is a hope with a document number.
The shortest group. It kills more engagements than the other three together.
Someone has to accept what the agents produce. Without that person the work stalls in review and the pilot dies quietly around month four.
When acceptance needs a quorum, your cycle time is set by that quorum's calendar, and no amount of automation upstream moves it.
Policy-as-code with no accountable owner drifts toward default-allow inside two quarters, which is worse than having no policy layer at all, because it still looks like one.
Retrieval inherits whatever access control already exists upstream. It never improves it.
Retrieval makes accidental openness discoverable in a way a search box never did. When that surfaces in month two, the architecture gets blamed for what the file server was already doing.
A classification living in a compliance spreadsheet can't be evaluated at query time, which is the only place it counts.
Entitlements have to be enforced inside the query, from metadata attached at ingest. Filter in application code after retrieval and one application bug becomes one data breach.
Entitlement sync lag is how a leaver keeps retrieving for hours. Worth measuring before anyone tells an auditor otherwise.
Two statements, both about whether a second year exists.
The gateway and the audit trail often pay back inside the first quarter. Knowledge graphs and fine-tuning pipelines don't, and those are the parts most often built too early.
If every workload is read-only question answering over documents, buy a licensed product and put it behind your SSO. Durable execution, compensation and an action ledger are a large share of the build and they earn nothing here.
The advice differs by band, and one of the three bands tells you to go away and come back later. That band is the reason this page exists.
Start the conversation, then argue about scope rather than about readiness.
Hold whoever you hire to a baseline measured from your own git history in week one, before anyone commits to a percentage. Pick the single value stream that irritates you most and refuse to let it grow while the work is running; a pilot covering four streams is four pilots and it will finish none of them.
Two questions worth asking us or anyone else: what does year two cost in licenses you buy directly, and how many of your own engineers are expected to be running this after the engagement ends. If either answer is vague, the handover isn't real.
Three or four untrue statements is ordinary. Almost nobody scores fifteen.
So tell us which ones before we quote, because the plan will spend its first fortnight on them and you should be able to see that line in the price. Hiding them costs you the fortnight anyway and buys nothing.
One caution. If your untrue statements sit mostly in data and permissions, treat the score as worse than it looks. Permission repair has a long tail and it's owned by identity, records and compliance teams who don't report to you and who plan in quarters.
If you answered no to more than four of these, do not hire us yet. Fix the build first — it is cheaper than paying us to discover it, and we will tell you the same thing in week one for money.
Where your no answers cluster tells you which repair comes first. They are not interchangeable and they run on very different clocks.
Come back and score yourself again when two of those are fixed. We'd rather lose the deal than lose your first six months. And if what you actually need is somebody to tell your board that AI is happening, hire a consultancy. They're genuinely better at that than we are.
This page is a filter, and a coarse one. The two-week assessment covers one value stream and returns four things.
We sell no products and hold no reseller agreements. Every tool involved is licensed by you, in your own name, at your own rate, and you can replace any layer of it without asking us. That is the only reason our opinion on tooling is worth listening to.
Describe the value stream that frustrates you. We come back with a baseline, a target and a price, usually within two working days. A real engineer reads these.
Book a readiness assessment