A research programme in self‑learning agents
There is no such person as “the user”.
Every system that learns from feedback assumes one teacher. Real work has many — different authority, different domains, disagreeing for good reasons. We’re building agents that learn from a team, and a brain spec that keeps what they learn.
The question
Why can’t an agent learn the way a person does?
Nobody hands a new hire a knowledge base and calls it onboarding. They do the work, someone checks it, and someone tells them why it was right or wrong. Six months later they’re making calls nobody needs to review — and not one line of that came from a document.
Agents don’t get that. They get a prompt, a file you maintain forever, and a memory that resets. So you explain on Monday what you explained on Friday — and the thing you’re actually maintaining isn’t an agent. It’s a manual.
The library was never the learning.
The thesis
Three claims we intend to prove.
Feedback needs something smaller than an answer to attach to.
A thumbs‑up on a whole answer is a scalar — it can’t say the diagnosis was right and the plan was wrong. Our agents answer in four separable parts, so every judgment lands on the part it was about.
Disagreement between reviewers is signal, not noise.
The field averages annotators into one preference. We think feedback should carry the authority of the person who gave it, in the domain they own — your tech lead on architecture, your support lead on what customers get told.
What an agent learns should outlive the tools it runs on.
The learning lives in a versioned brain spec — not in any chat surface or executor. Swap Discord for Teams, swap one executor for another, and the six months of judgment come along.
Measurement note. We track first‑pass acceptance, feedback events to durable behaviour, and calibration error against reviewer decisions. We publish results when they are results — not before. They land on the Research page as they accrue.
Where the programme stands
In beta. In public. In progress.
Three doors
Who we’re looking for.
We’re proposing a standard, not a feature: a portable learning layer between the interfaces people use and the executors that do the work. If category‑defining infrastructure is your thesis, we should talk.
Write to us →Multi‑principal feedback is an open problem and we’re running it live — on real teams, with real stakes. Research‑assistant positions are open; the problems are stated precisely, with what a contribution looks like.
See the open problems →Seven or more people, with at least four who review each other’s work across different domains. You bring one workflow you repeat and one you keep correcting; we put an agent in the loop.
See the harness →