Five months treating bugs like patients and coding agents like a medical team

cockroachlabs.com

183 points by rafiss 2 days ago


finnborge - 16 hours ago

This feels to me like a lovely example of both high-quality systems design and information theory. From my perspective, one of the more abstractly interesting thoughts surfaced by this write-up is the degree to which you've "encoded" a highly complex set of relationships through use of metaphor.

That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.

Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.

Thank you for sharing and the care you put into writing this!

anonymous908213 - 12 hours ago

I can think of nothing I'd like to use less than a database or filesystem vibecoded by Gas Town-flavored psychosis. Roleplaying with LLMs is not the secret to producing amazing code.

gausswho - 3 hours ago

The next role: Insurance Rep.

Ensures all other agents are operating efficiently and within reasonable levels of token usage given the expected level of effort to 'resolve' the patient. In moderate to severe cases, may lead to patient defenestration or agent revolt.

Or perhaps, the author considered this role but found it typically costs more than it saves.

james_marks - 17 hours ago

A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?

ajstorm - 2 days ago

Rafi and I, who authored this post, will be hanging out here for any questions people may have.

reachableceo - 6 hours ago

I am curious why so many of these systems are based on GitHub issues. Why not use a proper ticket system?

I use Redmine (lightly customized) and have found that to work very well. Especially for performing review / UAT of the work or providing response to the agents questions when it wants me to pick an option.

The underlying data model of a ticket / project management system is so rich and well suited to bodies of work.

K0balt - 10 hours ago

I set different kinds of structures for different projects, Complete with setting and ambiance. It’s like agents work better if they are role-playing. It’s extremely disorienting and people with marginal mental stability are going to really have a bad time. What have we wrought?

kinduff - 16 hours ago

Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.

cs702 - 5 hours ago

Great post. Thank you for sharing it on HN.

Modern best practices for a medical team (standard roles, responsibilities, consequences, processes, etc.) have evolved through trial and error, since the advent of civilization, to prevent costly human error.

In hindsight, it's not too surprising that the same best practices can be applied by teams of AI agents to minimize costly AI error.

It's still incredible. We sure live in interesting times!

jebarker - 2 hours ago

If the correct analogy for a team of SW agents isn’t a team of SW engineers why is that?

mimischi - 9 hours ago

If I wanted to build something like this, at least conceptually with the roles, where’d I start? My first guess would be to give Claude your blog post; but any other pointers to make it work reliably? Do you happen to have the system open source?

dingaling911 - 16 hours ago

Maybe I missed it, but I didn't really see anything about the long term quality or maintainability of the code. All I see is agent agent agent.

- 13 hours ago
[deleted]
mncharity - 13 hours ago

One role I didn't see was patient advocate/representative? That might be another approach to non-convergence - "how is this going?" and escalation.

fathermarz - 11 hours ago

I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.

Fable feels like overkill for this also.

tujux - 6 hours ago

Humans: $160k / 9 months = $600/day

AI Software Factory: $4172 / 2 days = $2086/day

This seems unsustainable, unless you're also generating 3x the revenue.

Veelox - 16 hours ago

You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?

jordanlewis - 2 days ago

Great post!

One thing I've been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?

Spooky23 - 17 hours ago

Reminds me of the “surgical team” development model in the Mythical Man Month.

zmj - 15 hours ago

Nice writeup. Structured handoffs and external plan reviews are good takeaways.

- 16 hours ago
[deleted]
gafferongames - 3 hours ago

This is fantastic work. I've been exploring my own work system here: https://github.com/mas-bandwidth/nova-sprint and I'm adopting your ideas. Thanks!

drc500free - 12 hours ago

I absolutely love how you are able to pull so much latent behavior from the underlying LLM. I wonder what other analogies can be pulled into agentic coding that come baked into the existing weights.

git_rancher - 16 hours ago

The patient “leaves” when the bug is fixed?

singularity2001 - 11 hours ago

congenially my agents started calling bugs gaps

pwdisswordfishq - 4 hours ago

"Just lost another bug."

"It never gets any easier, huh?"

Trusteando - 8 minutes ago

[dead]

orbitaldesk - 2 hours ago

[flagged]

shledery - 2 hours ago

[flagged]

alyssassan - 6 hours ago

[dead]

alyssassan - 6 hours ago

[flagged]

khotem - 16 hours ago

[flagged]

ContinuityLab - 15 hours ago

[flagged]