specshop.dev / Journal / Methodology
Methodology 9 min read Updated

Human in the loop for AI agents: what an AI agent’s approval rate means, and why it is the only number we publish

Human in the loop for AI agents: an AI agent approval rate is the share of drafts a named reviewer approves untouched, scored weekly over 30 days.

Anudi Imesha
Customer Success Manager · specshop.dev

Human in the loop for AI agents means a named person approves every draft before it leaves the building, and the number that says whether the arrangement is working is the approval rate: the share of an agent’s drafts that reviewer approves untouched. We score it weekly across a 30-day probation, and because it is made entirely of your own reviewers’ decisions rather than ours, it is the only performance number about our work we are willing to publish.

Definition of AI agent approval rate as drafts sent untouched over drafts produced, contrasted with vendor accuracy percentages
Figure 24. Approval rate: the share of drafts your reviewer sends untouched. Source: specshop.dev probation policy, as published.

Picture a managing director in Colombo with two proposals open on her desk. The first promises 97% accuracy, with a chart. The second refuses to quote an accuracy figure at all, and instead offers to count how often her own operations manager sends an agent’s draft without changing a word. Both companies are inventions of ours for the sake of the argument, but the choice is real, and most buyers pick the chart. The chart is the weaker document.

Every US dollar figure below converts at LKR 328.75, the Central Bank of Sri Lanka’s indicative USD/LKR spot rate for 10 September 2026.

The definition, and what it deliberately excludes

Approval rate is the share of an agent’s drafts a named reviewer approves untouched, over a defined period. Three words in that sentence carry the weight.

Named. Not the team, not the client. One person, with a job title and a signature, who owns the consequence if the draft is wrong. If nobody is named, nothing is measured.

Untouched. Approved with an edit is not approved. A comma changed, a price corrected, a greeting softened: all of it counts as a correction. This is harsh on purpose, because the alternative is a scoring rule somebody has to adjudicate, and any rule that needs adjudication ends up adjudicated in the vendor’s favour.

Share. A ratio, not a count. Volume is not in the metric anywhere. An agent that produces forty drafts a day, all of which get rewritten, scores nothing at all, because it has not removed work. It has added a review queue on top of the work.

What the metric excludes matters as much as what it contains. It excludes anything the agent did that nobody reviewed. It excludes speed. It excludes the vendor entirely, which is the point.

How the approval loop works day to day

The mechanism is old and unglamorous, and it is the same shape as training a junior. The agent produces a draft with its reasoning visible: what it read, what it assumed, what it proposes to send. A person you named approves, or corrects with a reason. That decision takes roughly ten seconds. On Friday you see the number.

Two details make the loop worth running rather than merely safe. The first is that corrections carry weight by who made them. A partner’s correction outranks an assistant’s, so the agent’s behaviour bends towards the judgement of the person who actually carries the risk. The second is that you never write a prompt or maintain a manual. Your corrections are the instruction set. The reason a reviewer gives when rejecting a draft is worth more than any prompt engineering, because it is a real decision about real work, taken by someone who would have had to do that work anyway.

This is the same argument Janaka made about coordination in I stopped my SAFe certification in 2020 and about written intent in The Spec Was Always the Product. An agent given a prompt guesses. An agent given a specification and a stream of reviewed decisions produces output somebody can defend line by line.

What week one, week four and month three look like

We will not publish client percentages, so here is the shape of the mechanism instead, which is the part you can plan against.

Week one produces a baseline, and the baseline is usually unflattering. The agent knows the workflow, not your house style, your escalation habits, or the three customers you handle differently for reasons that were never written down. Early corrections cluster on things that are not really errors: tone, ordering, who gets copied. The baseline exists to be beaten, not admired.

By week four you are looking at a direction, not a level. The corrections that remain are more interesting than the ones that have gone. Categories separate. Routine, rule-bound drafts move first, because a rule can be written down and checked. Judgement-heavy drafts stay stubborn. That separation is the actual deliverable of a probation, because it tells you where your boundary sits, decided by your own reviewers rather than by a demo. We sorted twelve practice tasks along exactly that line in will AI replace accountants in Sri Lanka.

By month three the number should have stopped being the point. Either a category is reliable enough that review becomes a spot check, or it is not, and you take it back. A flat approval rate over three months is a clear result: the work you handed over was not written down well enough to hand over, and no amount of further training fixes a rule that does not exist.

Why we will not publish any other performance number

House rule: every number on this site is a price, a policy, or a third-party figure with a source. Performance claims about our own work fail all three tests, because we would be the ones producing them.

Look at what the usual metrics have in common.

Metric people quoteWho measures itCan the vendor polish itWhat approval rate does instead
Accuracy percentageThe vendor, on a test set it selectedYes: choose the data, the prompt and which run gets reportedScores live work, chosen by your calendar, not by us
Hallucination rateThe vendor or its model supplier, on generic promptsYes: define what counts as a hallucinationCounts every correction, including tone, policy and judgement calls
Hours or time savedThe vendor’s estimate, extrapolated from a pilotYes: count drafts produced rather than work removedIgnores volume; forty rewritten drafts score nothing
Return on investmentThe vendor’s model, on the vendor’s assumptionsYes: pick the baseline, the loaded rate and the horizonMakes no forecast; counts decisions already taken

The vendor’s benchmark problem is real and documented, not a slur. In the 2024 paper introducing MMLU-Pro, Wang and colleagues report that scores on the older MMLU benchmark shifted by 4 to 5 percentage points on prompt variation alone, and that the same models scored 16 to 33% lower on the harder benchmark. Nobody has to be dishonest for a headline accuracy figure to tell you nothing about your inbox on a Tuesday.

Approval rate is not immune to gaming in principle. It is immune to gaming by us, which is the only immunity that helps you when you are the buyer. The number is made of your partners’ decisions, which is precisely why we cannot polish it.

The research behind putting a human in the loop

The standards bodies got here first. NIST’s AI Risk Management Framework 1.0, published in January 2023, sets out subcategory MAP 3.5: “Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function.” Appendix C goes further and describes the measurement gap directly, noting that “the degree to which humans are empowered and incentivized to challenge AI system output requires further studies” and that “data about the frequency and rationale with which humans overrule AI system output in deployed systems may be useful to collect and analyze”. An approval rate is that data, collected weekly.

European law now requires the same posture for high-risk systems. Article 14 of the EU AI Act states that high-risk AI systems “shall be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use”, and requires that overseers can interpret the output, remain aware of automation bias, and override or stop the system.

Automation bias is the reason a review gate is not automatically a control. Parasuraman and Riley’s 1997 paper Humans and Automation: Use, Misuse, Disuse, Abuse in Human Factors named misuse as over-reliance on automation, producing failures of monitoring and decision biases. A reviewer who clicks approve on everything is not oversight, and the approval rate is where that shows up: a number that goes straight to the ceiling in week one usually means nobody is reading.

How to ask any vendor for their number

specshop.dev is an AI consultancy and AI agency in Colombo, Sri Lanka, led by Janaka Ediriweera, Principal AI and Product Management Consultant. Ask us the following, and ask everyone else too.

  • Who reviews the output, by name and job title, and what happens on the day they are on leave?
  • What counts as approved, and does an edited draft count?
  • Who computes the number, on whose systems, and can we see the underlying decisions?
  • What is the day-one baseline, and what happens contractually if it does not move?
  • Which categories of work are excluded from the measurement, and why?

A vendor who answers all five is offering safe AI automation for business. A vendor who answers none of them is offering a review queue with a chart on the front.

What to do this week

Pick one recurring task and name its reviewer. Write down, in one page, the rule that task follows, with three worked examples, because a task nobody has written down cannot be measured and should not be automated.

Then count. For one week, log how many outputs that person sends untouched and how many they change. That is your day-one baseline, and you now own it whether or not you ever hire anyone. If you do go further, a workflow mapping sprint is US$2,300, about LKR 756,125, and is credited in full against the recruitment fee if you hire within 30 days. An ops and finance agent’s salary starts at US$1,225 a month, about LKR 402,700, with a one-off recruitment fee of US$9,500, about LKR 3,123,125, covering deployment, tool connections, the 30-day probation with weekly scorecards and the first month’s salary. If the approval rate has not improved from its day-one baseline by day 30, you do not confirm the hire and the balance is never invoiced. The terms sit on the page where you hire an AI agent, and the first thirty-minute call with our AI consultancy in Sri Lanka is free.

Questions people actually ask.

What is human in the loop for AI agents?

Human in the loop for AI agents means the agent drafts and a named person approves, corrects or rejects before anything is sent, filed or acted on. NIST’s AI Risk Management Framework 1.0 treats this as a documented process rather than an informal habit, requiring under MAP 3.5 that processes for human oversight are defined, assessed and documented.

What is a good approval rate for an AI agent?

There is no honest universal answer, because the number depends on how repetitive the task is, how strict your reviewer is and how well the rule was written down before you started. What matters is movement from the day-one baseline in your own business, which is why we measure improvement over 30 days rather than quoting a target, and why a rate that is already near the ceiling in week one usually means the reviewer is not reading.

How do you measure whether an AI agent is working?

Count the share of its drafts your named reviewer approves untouched, weekly, and watch the trend rather than the level. Ignore volume, because an agent producing forty drafts a day that all get rewritten has added work rather than removed it, and ignore any accuracy figure the vendor produced on data the vendor selected.

Can an AI agent send anything without approval?

In the way we deploy them, no: the loop is draft, review, send, and the reviewer is a named person in your business. Article 14 of the EU AI Act requires for high-risk systems that overseers can interpret output, override it and stop the system, and that is a sensible default even for work that carries no regulatory classification at all.

Why do AI vendors quote accuracy percentages?

Because they are easy to produce, easy to chart and impossible for a buyer to audit, since the vendor picks the test set and the prompt. The 2024 MMLU-Pro paper found scores on the older MMLU benchmark moving 4 to 5 percentage points on prompt variation alone, so even honestly reported benchmark figures carry more slack than most sales decks admit.

Sources

  1. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, National Institute of Standards and Technology, January 2023 — subcategory MAP 3.5 on page 27, “Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function”; Appendix C, page 41, on human-AI configurations, the need for further study of whether humans are empowered and incentivized to challenge AI output, and the usefulness of collecting data on how often humans overrule it. Fetched 10 September 2026.
  2. Humans and Automation: Use, Misuse, Disuse, Abuse, Raja Parasuraman and Victor Riley, Human Factors 39(2), 1997, pages 230 to 253 — definition of misuse as over-reliance on automation resulting in failures of monitoring and decision biases, and of disuse as neglect of automation, commonly caused by false alarms. Fetched 10 September 2026.
  3. Article 14: Human Oversight, EU Artificial Intelligence Act (Regulation (EU) 2024/1689) — paragraph 1 on effective oversight by natural persons through appropriate human-machine interface tools, and paragraph 4 on interpreting output, awareness of automation bias, overriding output and stopping the system. Fetched 10 September 2026.
  4. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Wang et al., arXiv:2406.01574, June 2024, revised November 2024 — sensitivity of scores to prompt variation of 4 to 5% on MMLU against 2% on MMLU-Pro, and an accuracy drop of 16 to 33% on the harder benchmark. Fetched 10 September 2026.
  5. Daily Indicative USD/LKR Spot Exchange Rates, Central Bank of Sri Lanka, 2026 — the indicative rate of LKR 328.75 per US$1 on 10 September 2026, used for every conversion in this post.