Feedback events 0 Accepted first pass 34%

Brain spec for AI agents

Agents that get promoted.

A team of agents that answer in four parts, wait for a human, and learn from every review — from which human, on what, and why. Until the work stops routing through you.

Early runs. Two in three proposals come back for rework — and every rejection carries your reasoning with it.

Feedback events0
Accepted first pass34%

The question

Why can’t an agent learn the way a person does?

Nobody hands a new hire a knowledge base and calls it onboarding. They do the work, someone checks it, and someone tells them why it was right or wrong. Six months later they’re making calls nobody needs to review, and not one line of that came from a document.

Agents don’t get that. They get a prompt, a file you have to maintain forever, and a memory that resets. So you explain on Monday what you already explained on Friday — and the thing you are actually maintaining is not an agent, it’s a manual.

The library was never the learning.

The answer

Every answer has four parts.

An agent that hands you a plan without saying what it thinks is wrong, or why, is asking you to trust a guess. B‑Spec answers in the same four parts every time, and each one stands on its own.

01

What is wrong.

The symptom, stated plainly, with the evidence it is reading from — the ticket, the thread, the ledger line, the log.

02

Why it is happening.

The cause it has settled on, and what it ruled out to get there. This is the part most agents skip, and the part you most need to check.

03

What resolved looks like.

The end state, described concretely, so there is something to judge the plan against rather than judging the plan against itself.

04

The path.

The steps in order, with what each one touches and what it would change. Nothing here happens until a person says so.

Because the parts are separate, your review is separate too. You can tell it the diagnosis is right and the plan is wrong — which is a far more useful thing to be told than no.

Where it comes from

A 2,500‑year‑old diagnostic.

The structure is not ours. It is the Four Noble Truths, the foundation of Buddhist teaching and one of the oldest formal diagnostic methods anyone has written down: there is suffering; it has an origin; it can cease; there is a path to its cessation.

It is often explained by analogy to a physician’s method — identify the symptom, find the cause, establish that recovery is possible, prescribe the treatment. We make no claim on the philosophy and we are not selling one. We borrowed the shape, because after two and a half thousand years nobody has produced a better one for working out what is actually wrong and what to do about it.

Most agent output starts at the fourth part and skips the first three. That is exactly why you can never tell whether it understood the problem or simply recognised the shape of one.

The workspace

It looks like a team, because it is one.

Agents sit alongside your people in channels you already understand. Work arrives, an agent answers in four parts, and whoever is best placed reviews it — inline, part by part, with a reason attached to each judgment.

Interactive — click a channel, judge a part, give it a reason

#customer-success 3 agents · 8 people To review · 2 parts

Interactive mock — nothing leaves this page. Try flipping a judgment: the reason field is the product.

Where it runs

Three tiers. Stated honestly.

The brain owns the learning; the surface it runs on is swappable. Running the beta on a surface we didn’t build is the first portability result — not a stopgap.

Live in beta

Today — Discord

Real teams, real work, real feedback events. Agents sit in channels beside your people; reviews happen inline, part by part, with a reason attached to each judgment. Rough edges and all — this is where the learning data comes from.

In development

Next — Slack & Teams

The same brain, on the surfaces most product companies actually work in. Because the spec owns the learning, nothing is retrained for the move — the adapters are the work.

Design mock

The direction — our workspace

A purpose‑built review surface: channels, agents and people together, judgments routed to whoever is best placed to make them. Every image of it you’ll see — including the one above — is a labelled design mock until it ships.

How it works

Written by the work, not for it.

A review loop you were going to run anyway. The difference is that B‑Spec keeps what comes out of it.

01

Work arrives.

Webhooks, chat, an email trigger, a scheduled routine, or somebody just asking. The request lands in a channel and an agent picks it up — grounded in whatever your connected systems already know.

02

It answers in four parts.

What is wrong, why, what resolved looks like, and the path — each with the evidence it used. Then it stops. There is no step where an agent acts on its own guess.

03

You review each part, and say why.

Mark any part right or wrong, with a reason. Approving teaches it as much as correcting — most systems only learn from failures, which trains an agent to avoid mistakes rather than to recognise good work.

04

It learns who taught it.

Feedback is not anonymous. It carries the authority of the person who gave it, in the domain they actually own, and it is weighted accordingly.

05

Then it acts.

Draft or send, according to your intent and its confidence. Patterns that hold across feedback events become durable behaviour — which is why the curve tracks iterations, not the calendar.

Multiplayer

It learns from your team, not from “the user”.

Almost every AI tool has exactly one user. This one has your whole team — and it knows the difference between you.

Weighted by whose it is.

Your tech lead outranks everyone on architecture. Your support lead outranks everyone on what to say to a customer. Feedback carries the authority of the person who gave it, in the domain they own.

Routed to who should judge.

It learns who is best placed to review each kind of work and sends it there, rather than to whoever happens to be on rota that morning.

Shaped to how you read.

The same answer, presented the way each reviewer wants it. Full evidence for the person who always asks for it; the short version for the person who never does.

Aware of where you are strongest.

Over time it builds a picture of whose judgment holds up where. That picture belongs to the people in it — visible to them, used for routing, and never a performance score.

The bar

It asks when it isn’t sure.

Once first‑pass work is reliably right, you can let an agent act without a reviewer — but only above a confidence bar you set. Anything below it still comes to a person. Every time, no exceptions, no override.

Straight through0% Asks a human100%

The bar is yours to set. The score is ours to justify.

On that number

98 has to mean 98.

A confidence score a model invents about itself is worth nothing. Ours is calibrated against what actually happened — the decisions your reviewers made, on your work, over every feedback event you have given it. When it says it is 98% sure, that number has been checked against outcomes.

The threshold itself stays yours. Set it high while you are still building trust, move it as the evidence comes in, or leave every action behind a human indefinitely. Autonomy is a setting, not something that happens to you.

Blast radius

What an executor is allowed to touch.

Once a path is approved, the agent drives execution tooling to carry it out — a Hermes setup updating your site, a hosted OpenClaw instance running a campaign. The executor works for the agent; the agent answers to your team. Which is exactly why the boundaries are explicit:

Scoped

Every executor runs with an explicit scope — these pages, this campaign, this ledger. A path that would reach outside its scope doesn’t run degraded. It doesn’t run.

Reversible first

Paths declare what’s reversible. Irreversible steps are flagged in part four before you approve — not discovered after. Draft‑or‑send is per action, per your intent.

Accounted

Every executed step reports back — what changed, where. A bad prescription is visible in one place, attributable to one approval, and feeds the next round of learning.

Two teachers

Your reviewers judge the diagnosis; reality judges the path. Execution outcomes flow back into the same brain as human feedback — two supervision signals, one ledger.

The roster

Three agents now. More when they’re undeniable.

The aim is to cover most of the work of keeping a product company running. We would rather ship three that are undeniable than six that are plausible.

Customer Success

Support replies, escalations, refunds and credits, churn signals, renewal and QBR prep. Reviewed by whoever owns the account.

Live

Ops & Finance

Invoices and dunning, reconciliation, vendor admin, month‑end pack, spend queries. The domain where the confidence bar earns its keep.

Live

Web

Site copy and page updates, broken links, redirects, publishing, and the small SEO jobs that never reach the top of anyone’s list.

Live

Sales & Marketing

Sequences, proposal drafts, campaign copy, pipeline hygiene.

Next

Dev Lead

Triage, review notes, release prep, and the CI failures nobody wants to open.

Next

Inputs

Work arrives. Context is already there.

How work arrives

Webhooks
Chat — Slack, Teams
Email triggers
Scheduled routines
Forms & intake
Somebody asking

What it already knows

Helpdesk & tickets
Shared inboxes
CRM
Billing & subscriptions
Accounting ledger
Issue trackers
Docs & wikis
CMS & site content
Git, PRs & CI
More on request

What it feels like

You stop holding it up.

At the start you are in every loop. Every proposal waits on you, and the whole thing is standing on your attention.

Supervision100%

The output never changed. The support did.

And one more thing

The same brain works for your people.

Every reason your reviewers gave is knowledge the rest of the team can use. A new hire’s first week runs on the same judgment your seniors spent years building — not a wiki nobody has updated since the last reorg, and not a two‑hour call with whoever happens to be free.

It cuts both ways. Your seniors stop being the only place institutional knowledge lives, and your juniors stop paying the tax of not knowing what everyone else already knows. Nobody had to write any of it down.

What changes

The shape of your week.

Measuring 1st‑pass acceptance rate, tracked per agent and per part. Publishes on the Research page when the cohort has earned a number.
Measuring Events to durable behaviour. Iterations, not months — your volume sets the pace, and we publish the distribution, not a cherry‑picked case.
Verified 0 knowledge base articles you have to write, review or maintain. True by construction — the review loop is the documentation.

No invented metrics. Every figure on this site is measured or absent. The 98% in the gate section is illustrative of a customer‑set threshold, not a performance claim, and the workspace screen above is an interactive design mock, labelled as one.

The beta · open invite

Product companies. 7+ people. Four domain owners.

This is for teams of seven or more with at least four people who regularly review each other’s work across two or more domains — someone who owns what customers get told, someone who owns the money, someone who owns the site or the code.

If your team can’t name four, it isn’t for you yet. Discord‑native teams get the shortest path in — that’s where the beta runs today.

Seven people. Four owners. One brain.

Book a demo

See it on a workflow you actually run.

Thirty minutes. We’ll walk through two or three real cases where an agent went from every answer being sent back to almost none, then look at where your own work is still waiting on a person who has better things to do.

  • 01What you bring: one workflow you repeat and one you keep correcting.
  • 02What you see: the same workflow at the first review and at the thousandth.
  • 03What you leave with: an honest read on whether this helps you yet.
Deriving…