Coding agents are my daily build tool, not a demo I ran once. That changes what you notice about them.

The thing I notice most: an agent handles a well-specified task well and a vague one badly. The bottleneck is almost never the model. It is the shape of the work you hand it.

Four agents, one pipeline

At Dubizzle Labs I built an autonomous SDLC pipeline with four agents. A ticket-creation agent turns a request into a ticket. A plan-creation agent turns the ticket into a plan. A coding agent implements the plan. A testing agent checks the result.

It is tempting to describe that as a four-agent system and stop there, as if the number were the design. It is not. You could build the same four agents and get nothing useful out of them.

The architecture is in the decomposition, not the agent count.

Contracts at every handoff

What makes the pipeline work is that every handoff has a contract. Each stage has explicit completion criteria: what its output must contain before the next stage is allowed to start. Each agent has its own guardrails, scoped to the one job it does.

That turns a vague request into a sequence of small, checkable tasks. The planning agent never gets "build the feature". It gets a ticket that already passed a definition of done. The coding agent never gets an idea. It gets a plan.

Without those contracts you have four agents passing ambiguity down a line, and each one makes the ambiguity worse.

Distrust the confident handoff

The failure mode that matters in agent systems is an agent that looks like it succeeded. It returns something well formatted, plausible and wrong, and the next stage builds on it.

Completion criteria are how you catch that at the boundary instead of three stages later. The testing agent at the end is not there because tests are good practice. It is there because the earlier agents will sometimes be confidently wrong, and something has to be designed to disbelieve them.

What I take from it

When I design an agent system now, I start from the handoffs. What does each stage promise the next? How would the next stage know if that promise was broken? The agents fall out of those answers. Choosing how many to have is the last decision, not the first.