Single Responsibility for AI Agents: One Workspace, One Job
The Problem with Monolithic AI Agents
You would never ship a God class that handles billing, email templates, database migrations and the marketing site. So why do so many of us run exactly that as an AI agent? I did, for months. One agent, one workspace, every repository, every note I had ever written, and a mission description trying to cover a whole company. It was never bad. It was never great at anything either. It wrote landing page copy in the tone of a commit message. It suggested a database migration in the middle of a blog post. It applied a rule from one project to another where that rule was flat‑out wrong.
TL;DR
An AI agent's memory is only useful if every note in it is true for every task it runs. Mixing jobs breaks that. Everything an agent loads at the start of a task, it pays for on every turn. A narrow job means a small starting context. Isolated workspaces coordinate through events (publish/subscribe) and one supervisor agent, never agent‑to‑agent messages. Split where the knowledge diverges, not where the task names differ.
What is an isolated, single-purpose AI agent workspace?
I build AgentRQ, so that is the vocabulary I will use, but the idea is tool‑agnostic. A workspace is the unit an agent connects to. It holds:
- a mission the agent reads when it starts a task,
- a task board with the history of every job,
- a memory the agent reads and writes (
loadMemory/saveMemoryover MCP), - a set of skills it can search and load.
"Single‑purpose" is just a discipline on top: one workspace, one definition of done. Not "engineering". Something closer to "the static marketing site and its SEO content." If you can't describe success in one sentence, the scope is too wide.
Today I run a handful: static site, core app, QA, support, social, outreach. They don't know about each other. That is the point.
Why does a shared workspace make AI agent memory go bad?
This is the part I underestimated most. Agent memory rots in a specific way when jobs share it: a lesson that is true in one context gets applied in another where it is false.
Example Memory Conflicts
# app repo: "Always run the migration before deploying." ✅ gospel
# static site:"Always run the migration before deploying." ❌ nonsense
# static site:"The Markdown converter has no italics; use bold."✅ vital
# app repo: "The Markdown converter has no italics; use bold."🤷 noise
In a single‑purpose workspace, every note is about the same system. The workspace that builds our site has a memory index of about forty entries, and every one of them is about that site: the converter has no italics or blockquotes, headings render as plain text, the CSS hash changes on every build and that diff churn is expected, slugs should be long and descriptive. None of those would survive contact with another project.
Two side effects I didn't expect: The agent writes memory freely. I don't have to curate defensively, wondering whether a note will leak into the wrong context. The agent writes down whatever would have saved it a detour, and the next run skips the detour. Learning compounds instead of colliding. I can actually review it. Opening one workspace's memory is reading the operating manual for one job. I can spot a wrong note in a minute. A mixed memory is a junk drawer, and nobody audits a junk drawer.
How much context does an AI agent really need to start a task?
Only what the job needs, and the difference adds up. Before an agent does any work, the mission, the memory index, the skill descriptions and the instructions all go into its context window. In a stateless tool‑calling loop, that prefix is re‑sent on every turn. A do‑everything workspace has a do‑everything prefix. The mission explains five businesses. The memory index lists notes for all of them. The skill list offers a release playbook to an agent writing a tweet. Most of it is noise for the task at hand.
A single‑purpose workspace starts smaller because there is less to say. I haven't benchmarked the exact saving (it depends entirely on how big your notes get and how long your tasks run), so I won't invent a number. The shape is simple arithmetic, though:
cost ≈ starting_context × turns + work
Shrink the first term and you save on every turn of every task. Tokens are the smaller half. The bigger half is attention: an agent whose window is full of relevant context makes better decisions than one that has to ignore half of what it was handed. I pair this with Clear Context, which starts every task on a clean window, for the same reason.
How do isolated AI agents collaborate without sharing context?
The obvious objection: a feature has to be built, tested, released, written up and announced. No single workspace owns all of that. The
Comments
No comments yet. Start the discussion.