I've been working with AI tools for a while now. But a team of agents — not tools I pick up and put down, something closer to a standing team working alongside me — that's only the last six months. Infrastructure monitoring, research, writing, calendar-scanning, working memory, code. Running continuously. The whole thing.
I thought I had a pretty accurate picture of what that would look like. I was wrong about most of it.
Here's the honest version.
I thought it would feel like delegation. It doesn't.
The word I kept reaching for, before I tried this, was delegation. Hand the work to the agent, review the output, approve or redirect. Clean. Managerial.
It's more recursive than that. The agents surface decisions I didn't know I had to make, ask questions I hadn't thought to ask, and come back with work that changes what I thought I wanted. Half the time what I'm doing isn't reviewing output — it's being prompted to think more clearly about what I actually asked for.
That's not a complaint. It's a better kind of work. But it's not delegation in the way I understood it. You're not managing people who execute a brief. You're in a continuous, fast loop where the output of one exchange shapes what you do next. If you came into this expecting to hand off and step back, you'll be surprised.
I thought the leverage would come from speed. The real leverage is memory.
Speed is real. Work that used to take a week takes a day. I'm not going to oversell that.
But the compounding effect I didn't anticipate is continuity. The team I work with now carries context I would have lost. Decisions I made months back, why I made them, what I ruled out and why. When I'm debugging a product call or trying to remember what I decided about something earlier in the year, that's in the room with me.
Most solo operators and small-team founders bleed time and clarity through context loss. You forget why something was built the way it was. The institutional knowledge lives in someone's head, and when that person's attention drifts, the knowledge drifts too. A team of agents that works with you continuously doesn't forget. That turns out to matter more than speed.
I thought I'd catch all the errors. I don't.
I was never naive enough to think the agents would be perfect. I expected mistakes. What I didn't expect was how often the confidence of a wrong output would look similar to the confidence of a right one.
There's a version of this that sounds like a generic "AI hallucination" complaint. It's more specific than that. The errors I miss are usually the plausible ones — the output that's correct in form and subtly wrong in substance. Not a fabricated fact, but a slightly off framing. Not a failed task, but a completed task against a spec I'd described imprecisely. I've shipped a few of those. They weren't catastrophic. But they were mine as much as anyone else's.
Working this way puts real pressure on how precisely you can articulate what you want, and on building verification habits rather than relying on trust. I'm still calibrating both. Six months in, I'm a better question-asker than I was. I'm not done getting better.
I thought the hard part would be technical. It's not.
The hard part is knowing when to be in the loop and when to let it run.
For anything reversible and well-scoped — pull together a brief, monitor a system, run a search, draft something I'll edit — the loop can run. I don't need to supervise every step. That's where the leverage comes from.
For anything with downstream consequences that are hard to undo — a message that goes to someone important, a decision about how something gets built, a judgment call about risk — I need to be in it. The loop needs me in it. And the failure mode I've had isn't "ran without me when it should have stopped." It's "I wasn't precise enough about which category this was, so it treated a consequential thing like a routine one."
That's an articulation problem, not an agent problem. But it's the thing that requires the most ongoing attention, and I didn't have a good model for it before I started.
I thought I'd eventually feel like I had it figured out. I don't.
Not in the way I expected to, anyway. Six months of running this has made me more comfortable with the rhythm of it and more honest about the limits. What it hasn't done is produce a feeling of settled competence — the sense that I've solved how to do this and now I just do it.
It keeps changing. The team's capabilities evolve. I figure out a new pattern that works and then the context around it shifts. What I've learned these past six months isn't mostly how to use agents. It's how to work alongside something that grows and fails and improves in a cycle, while I do the same thing. That's closer to a relationship than a tool.
I'm not saying that to make it sound profound. I'm saying it because if you're coming into this expecting a system you install and then run, you'll be disappointed. You're building a working relationship. That means it gets better the longer you're in it, and it also means the learning never really stops.
Six months ago I thought I was building a workflow. I was actually building a team. That's a different thing, with different demands on you. Worth knowing going in.