Keelson · Blog
AI Ops vs Build: which one is this problem?
01 · Intent signals
What you are feeling this Monday.
The same operator reads the dashboard two ways depending on which kind of work the queue is asking for. The first read is the build read — drift that compounds week to week, a metric your CFO can already verify, and a queue whose exceptions all look the same. The second read is the ops read — a system that runs and the team quietly overrides the 4% of cases it does not yet ship cleanly.
Claims cycle time is up three weeks in a row. Contracts drafted per week is trending down. Exceptions cleared per shift lives above the threshold the team set in month one and never came back below. Anything that drifts week over week and the operator cannot name the cause on a Monday morning is a build cue, not an ops cue.
When the exceptions queue is full of cases the operator has already seen yesterday and the day before, the queue is not noise. It is the labels the system is missing. Same shape, same field, same upstream feed dropping the same record — that is the build cue the ops retainer can never fix on its own.
02 · Scope triggers
When the work has crossed from ops into build.
Two scope triggers move the work across the line from a retainer into an engagement. They are simple to read off the operator data and they are the only two that matter for a mid-market budget. If neither applies, the answer is ops and the invoice should match.
Single-workflow scope
If the conversation ends with “claims triage,” “contract drafting,” “settlement matching,” or any one workflow the team already runs by name, the work is small enough to scope as a 90-day build. If the conversation ends with “the operations function” or “the AI stack,” the work is too big for a build and the answer is AI Ops.
Single-owned metric
A build is the right shape when the operator can name the metric that will move if the work ships. The metric is locked to the quote and the metric is the number the CFO will read on a Monday dashboard three months later. If no single metric will move, the work is an ops retainer and not a build.
03 · TCO math
The three-year cost of picking the wrong one.
The expensive error is calling a build an ops problem and paying a retainer to keep patching what a 90-day build would have shipped in quarter one. The cheaper error is calling an ops problem a build and paying a 90-day invoice for work the operator could have done with the vendor on the phone a year later. Both directions cost real money across three years; the cost is just front-loaded differently.
A mid-market AI Ops retainer runs between fifteen and forty thousand a month. Over three years that is between half a million and a million and a half of spend on a system whose exceptions queue is the same shape in month thirty-six as it was in month one. The total includes the metric the operator never moved and the runbook the team never owned. That is the cost of the wrong call.
A 90-day build begins between sixty and a hundred and twenty thousand for a mid-market scope. Against retaining AI Ops for the same period — six to twelve months of retainer at fifteen to twenty-five thousand a month — the build is several times more expensive than the work needed. The downside is real and one-directional; you cannot un-spend the build invoice. The upside is small and conditional; a phone call would have shipped the equivalent change in two weeks.
04 · The 90-day decision rule
The Monday test an operator can run alone.
Pick one Monday, open the operator dashboard, and walk through four checks. If three out of four are yes, the work is a 90-day build. If one out of four is yes, the work is AI Ops. If two yes answers land on each side, the answer is the hold call below and the engagement is not yet ready to be priced.
Can you name the workflow in one sentence? “Claims triage,” “contract drafting,” “settlement matching.” If yes, it is shaped for a build.
If the work ships, does one number on the operator dashboard move? The CFO should be able to read the number without your help before and after.
Last quarter, did the exceptions queue carry the same shape two weeks in a row? Repeated shape is the labels the system is missing — build work, not ops noise.
Can the operator name the cause of the drift without opening a ticket? If you cannot, the work is build. If you can, the work is ops wear and tear.
05 · When neither answer is right
What we recommend when the test comes back inconclusive.
Two yes answers on each side is the most common answer the test returns at a mid-market operator we have not worked with before. The work is real but the shape is not yet visible. The right next step is a two-week audit, not a build and not an ops retainer. The audit ends in one document — the workflow, the metric, the exception taxonomy, the locked quote. The audit costs less than a quarter of either wrong call and it is the only procedure that recovers the cost of being unsure.
Continue reading
Related posts
More working notes for operators building AI systems that hold up after handoff.
AI Ops / Strategy
How mid-market teams scale AI workflows after the first build: strengthen governance, change operator behavior, assign durable ownership, write runbooks and escalation paths, and expand a portfolio without losing metric discipline.
Governance / AI Ops
A practical operating checklist for mid-market teams: inventory every AI workflow, assign owners, set risk tiers and approval gates, control data and vendors, define human review and exception escalation, log what matters, and review the system monthly and quarterly before drift becomes an incident.
How mid-market operators choose one workflow, establish a baseline they can defend, measure implementation ROI, set pilot exit gates, and hand a production system to the person who will run it on Monday morning.
The practical AI ops digest
Working notes on outcome-locked pricing, the 90-day build, and what AI Ops actually does after handoff. No marketing — just the operator notes we send to the people who already read this blog.
Bring us the workflow
Read every post, then bring us the workflow.
The audit ends in one document, target selection ends in one quote, and the metric ships in 90 days. No NDA required to start the conversation.
Replies within one business day.