Reasonable But Wrong

My orchestrator recently shared a question from an implementing agent: it found a race condition, where a worker could update a ticket status even after its ownership expired. It proposed having the worker build a transaction around both reading the ownership and writing the status.

Reasonable. But wrong.

In traditional, human-driven software development, we had a term for a particular kind of work: silo programming. This is where an engineer would talk to people to get the requirements, then disappear for weeks into an office while they furiously coded, then at the end present the finished work. We joked about slipping pizza under the door to keep them fueled.

Over time, we realized that this often created downstream problems around repeated code and idiosyncratic architecture. We shifted to daily standups so people could at least be aware of what others were doing, with the hope that this would lead to greater collaboration. It even occasionally works.

Coding agents are silo engineers, turbocharged. Give them a spec, they go on their implementation journey, and come back with tens of thousands of lines of code. They love to implement, and you'll get an entire collection of functions with slight variations that do essentially the same thing.

And this is the root cause of the issue in my example: the implementing agent wanted to write its status into the DB because it did not understand that I have a rule stating that entities must have a single owner. From where it stood in its silo, it could not see the other microservice that owns the tickets. Instead of working within the boundary, it proposed moving it. Reasonable but wrong.

In my own workflow, and as seems the emerging best practice in agentic engineering, we spawn new agents per task. They dedicate the context window to implementing just that piece. This is intentionally building the silo, but I've found this gets the best results per worker. I'd rather not change it.

We addressed this in humans by increasing communication between them. In an agent workflow, this translates to a cross-task gate. I've added a step in my workflow that does a holistic review across all the tasks in a release. In practice, this often catches mismatches between the tasks because this is the only place those contradictions appear: each component is internally consistent. (My own telemetry shows between 8% and 20% of the defects caught by my final cross-task reviewer are not visible in the context of any individual task.)

I also address this by constructing architectural rules that are explicit enough to check. "One writer per entity" is unambiguous, and I can write a mechanistic check around this. This blocks the most obvious issues.

Eventually though, the agents will reach the fuzzy bounds of what can be done deterministically. I have a knowledge base they can search where my past decisions are recorded, but at the same time they have explicit guidance to escalate to me when the path is ambiguous. Which is why the "should I add a transaction" question appeared in my queue.

I'm not going to stop slipping the proverbial pizza under the door. Agents, like humans, often do brilliant work when given the opportunity to focus. They will keep making reasonable-but-wrong suggestions. My job is to ensure that the cross-silo coordination is built into the process, and the hard calls still make it to me.