Fences, Not Sandboxes

yegge.ai

58 points by tosh 10 hours ago


nepthar - 7 hours ago

(Asking sincerely in good faith - I read this note along with skimming a few of the linked ones, and I'm familiar with gastown)

Can someone explain the value calculation to me? This seems like someone has truly let AI VASTLY expand the codebase of what seems to be a medium sized hobby game to ~10-100x the amount of engineering required. Especially with statements about how wheelhouse, his software factory, has grown to nearly the size of the actual software he's writing. He also mentions several times that wheelhouse is specific to developing the game. Then, he discusses "pulling in beads" which itself seems on my read to be enormously engineered. (note that he also says that you burn a lot of tokens with agents "keeping your beads in sync" and I didn't have the time to figure out what that meant, but it looks like as complicated as it is, beads can't reconcile itself without burning $$$).

It seems like if you answer "yes" at every time you have the question of "can I make AI do this", you end up burning $120k/month in tokens on your side project.

Again, I am not disparaging this- but I feel like I am genuinely missing something and would welcome help understanding it.

stillpointlab - 28 minutes ago

What Yegge describes in this post is something I have seen to a lesser extent in my own agent use.

But one specific thing I noticed, which his inline comic lampoons, is Fable (especially) explicitly stating "[User] has ruled", or "The rulings are in".

I was confused by this until I saw a post on X about how someone came to his agents in the morning after they ran all night, and the sub-agents had refused the requests from the orchestrating agents because they thought the decisions being made were not inline with the user. They thought the decisions were from the ochestrating agent.

This, as well as the governance things that Fable and Codex both often request, made me thing of provenance, especially related to decisions an authority. As we get deeper and deeper hierarchies of agents (as we likely will) this idea of authority, who has it and where does it come from, feels like it will be a key component to agental systems.

xnx - 6 hours ago

> First, my secret: I see the future by living in it. I am spending the equivalent of $122k/month of API token spend, or about $4,000 per day, using 21 Claude Max accounts, a number that has been growing steadily at 2 per week.

Good to know that together, Yegge and Zitron bring balance to the force.

mccoyb - 7 hours ago

> I am running an organization of around 50-60 agents, five of whom are interfacing with around 10 humans in the outside world: myself, my 5-person core game design team, my accountant, my chief of staff, and a few others. Only Fable is allowed to talk to humans, via Slack and email.

RIP to those poor humans. I can't imagine having your brain melted by Fablespeak as your FTJ.

Animats - an hour ago

This scrolled off HN too fast.

The concept is like Gas Town - AI as a organization, not an emulated human. Yes, it's inefficient. But it scales.

(A very long time ago, I got a tour of Xerox PARC, even before Steve Jobs did. Alan Kay explained that they were building the future of computing, accepting that it cost far too much to be cost-effective. They assumed the hardware would catch up. It did. Took about ten years.)

CoolestBeans - 7 hours ago

What is the overhead to have agents play model UN? Why is the coordination so elaborate? It sounds like a deeply complex and expensive emergent behavior that maybe looks comprehensible but could be nonsensical. Also like any complex system, can you actually predict the outcomes?

Waste is a failure case. How do we know the code factory is actually productive or just agents filling up the computation resource cap because they can?

jshaqaw - 6 hours ago

The proof is (or isn't) in the game. If the game is something genuinely great then this all worked and is important. If the game is a bloviated, boring, and derivative mess then it didn't. Without any proof point on the output it is hard to say if the article has any value.

AgentOrange1234 - 5 hours ago

"Wheelhouse is about 600k lines of code and tests (mostly bash)"

This article is a joke right? Please tell me you are joking. Jesus christ.

grebc - 7 hours ago

Yegge’s writing has declined proportionately to his AI usage.

sneilan1 - 6 hours ago

I'm glad someone is trying this and exploring what's possible. It's easy to deride his work by saying what's the value but are you running 50 fable agents and letting them run wild to see what the future looks like?

maherbeg - 2 hours ago

I'll say that Gastown sounded absolutely crazy, but the idea of having orchestrator threads to manage your work and keep tabs on it, having validators to validate the other work etc. were generally the right shape. I think GasTown probably could have been really successful if there was a pared down version with more obvious names rather than the fun names.

I'm going to be thinking about this blog post for a while though because if you squint and tear it apart, there are probably really good generalizable pieces in here to take home for future models.

zbentley - 3 hours ago

Congratulations, you have reinvented RFCs, architecture reviews, linters, and integration tests. From first principles. With cutesy names.

nvader - an hour ago

https://store.steampowered.com/app/1541710/Wyvern/ Wyvern on Steam.

Looks like Runescape

vishalrad - 3 hours ago

100% This is EXACTLY my experience and I didn't even spend 1% of what Steve has. Fable loves to create rules around the evaluations and decisions that I make. This is good, but this is also scary because 1/ I am fallible 2/ I don't have time to digest every detail and make a careful decision. And if I do spend that time, the system will return in 10 mins and give me another massive set of decisions to make .. thus creating an unending loop, resulting in decision fatigue... which leads to #2 again.

So AI is optimizing in some ways for an AI as the judge, not human as a judge. It needs this input, this steer, but humans aren't built to support this.

pianopatrick - 3 hours ago

If the agents are really doing politics now, we can test and see which method of organizing agents politically works best.

danpalmer - 2 hours ago

Lots of talk about engineering his game, lots of talk about how this game is exactly what he wants, no talk of play testing. Is there any world in which this is going to be a good game?

nzoschke - 5 hours ago

Seems like this is the year of the sandbox and governance layers.

I recently wrote how we set all this up for our agent computers that help people with their email:

https://housecat.com/blog/agent-computer-101

I do agree even more rules help keep things on track, but I do these as linters, specifically as Golang `go vet` and `go fix` commands. That works 100x better than any SKILL.md or team of agents in my experience.

perrygeo - 4 hours ago

> So it's not $120k/month of real money, but it's still crazy spend. I would guess I'm one of a handful of top individuals on Earth outside the frontier labs, in terms of my experience with top-end models.

We get it bro, you buy lots of tokens. My son buys lots of pokemon cards. Still haven't seen shit for return on that investment either, but as long as its all in good fun, "you do you".

> That's how I'm able to tell the future. I'm living in a world that will not become cost-effective for most people for another year.

That's pretty easy to verify. #remindme August 24, 2027 - are $120k worth of tokens in 2026 generally "affordable" by 2027? By what mechanism have the economics shifted?

Theory5 - 4 hours ago

I don't even like the idea of humans operating without guardrails where I work, much less AI.

coder-pm - 7 hours ago

How do you find the bad decisions with so many subscriptions active? isn't that looking for needle in a haystack:)?

kibwen - 5 hours ago

In the glorious future where I can conjure a game with a snap of my fingers for the cost of a pack of gum, why would I ever choose to play Wyvern, the game tailored to Steve Yegge's tastes, when I could just make a game tailored to my tastes? Perhaps Steve will be content to play his MUD with agents pretending to be human players?

fwlr - 4 hours ago

> the subject matter is too complicated to explain … All I can do is walk you around like an excited tour guide, one who has unearthed an ancient alien civilization.

Now this is AI psychosis.

SwellJoe - 3 hours ago

It kinda feels like Yegge is trying to start a cult. And, I just don't find that very interesting.

AWebOfBrown - 3 hours ago

I honestly think Yegge has reached a point where he knows he isn't interested in doing the leg work of today's software engineering, but wants his career to ride the wave. The lowest effort, maximum value to extract in his position is abusing the crap out LLMs to an extent that is genuinely novel to secure thought leadership, but on close inspection it's just pointless token spend driving blog posts and publicity.

To me it is absolutely farcical that he was being being paid by a harness company (Amp) as a staff software engineer, slop-coded a solution to keeping LLMs on track (Beads), and having paid for all of this Amp and Steve...parted ways. Beads never went into Amp. Someone cottoned on to the value Steve was providing, for my money.

Avshalom - 6 hours ago

>This tier, or the one just after it, will power hundreds to thousands of new AI employees at every company.

>And companies are in no way, shape, or form prepared for this transition.

so, uh, what's the business model here?

BeetleB - 7 hours ago

> I didn't want this to be a long post, and I think I've succeeded.

Classic Steve Yegge!

throwaw12 - 3 hours ago

I think Steve somewhat predicted the future well. He might be slightly off, because he is overly optimistic and operates at the edge, but look at Gas Town, when it was released it felt like dystopian, today it looks not too far away, I am sure most of your orgs are already running some kind of agent to triage the tickets and in some cases automatically open the PR in your git repo.

I think he is onto something this time as well

groby_b - 5 hours ago

"I would guess I'm one of a handful of top individuals on Earth outside the frontier labs, in terms of my experience with top-end models."

I'd argue that number of tokens burned does not translate into "experience".

And if you needed to spend $300k or so to find out that process helps manage larger efforts (process is what "fences" are outside Yegge-land), you sure spent a lot.

I mean, I'm glad the guy's got a a hobby he enjoys, but there's less insight than you'd hope for.

- 6 hours ago
[deleted]
- 7 hours ago
[deleted]