When I set out to build my own AI operating environment, the architecture seemed obvious: model it like a company. A scrum master. An agile coach. A solution architect. A security engineer. A proposal writer. I ended up with a roster of more than thirty specialized agents, each with a name, a charter, and a carefully written role definition.
It looked impressive. It was also mostly bloat.
What running it revealed
The problem was not that the agents were bad. It was that a role definition is not the same thing as delivered work. Most of those thirty agents accumulated definitions faster than run records. Each one added a coordination surface — another charter to maintain, another handoff to reason about, another place for context to go stale — without adding output I could point to.
The signal was in the run logs. A small core of agents did real, repeated, trackable work. The rest were organizational theater: an org chart I had built because org charts are what companies look like.
The consolidation
I cut the roster to six. Not because six is a magic number, but because six was what the evidence supported. Every agent that stayed had a track record; every agent that went was a definition in search of a purpose.
The lesson generalizes beyond my own setup. Agent value does not come from specialization on paper. It comes from a clear task, trustworthy context, an enforced review path, and evidence you can inspect afterward. One agent with those four things outperforms a department of agents without them.
What I would tell a team starting out
Start with fewer agents than you think you need. Give each one real work and track it — task IDs, costs, outcomes. Let the roster grow only when the run records justify it. If a definition sits idle for a month, delete it; you can always write it again the day you actually need it.
The org chart is not the system. The work is the system.