Pion, an agent designed to run any company autonomously
andonlabs.com360 points by lukaspetersson 13 hours ago
360 points by lukaspetersson 13 hours ago
Meanwhile, I can't even get Astra to consistently re-use the same font-size across all of my HTML page headings+subheadings.
There's just no world where this actually results in a stable, respected business. It will be death by a thousand bad impressions, mistakes and oversights. Yeah, you can automate everything. That doesn't mean it's being done well.
That's not to say this might not be something feasible in 2-3 years from now, but we're still a long ways away.
People said Devin, the AI coding agent, was garbage in 2024. Now it's weird if an agent isn't writing 99% of your code.
So you're right to be skeptical as of September 2026. But September 2028? It might just be weird to run companies without AI management.
You’re mixing things up. Coding is work while managing is status. Management will force the peasants to use any pitchfork management wants. But giving status and power away will never happen. It’s obvious that there enough managers to replace with AI for positive outcome. Bet it will never happen.
> Now it's weird if an agent isn't writing 99% of your code.
It is?
Are you sure that this is a representative reading that applies to more than just a tiny bubble?
I don’t know a single developer who has an agent write 99% of their code. Is that the norm outside of my bubble?
It's writing nearly all of the code where I work. It also reviews all of the code (I review it as well). You still have to guide it and hold its hand though. I dont write code by hand anymore (and haven't in almost a year). The job is now about managing the bots, understanding the architecture and keep everything aligned.
It writes all the code but we often have to go through several iterations. It won't work autonomously yet.
Yeap, one more data point here. Most of code here is AI generated with some handcrafted adjustments.
we aren't anywhere near lights off software factories, and writing software is easiest domain for modern llms, you can verify results rather easily, plenty of training data etc.
I don’t think this is true at all. Maybe if you have an agent writing some of your code it will be 99% because its so verbose but I don’t think the modal software engineer is a vibecoder.
If you think that all instances of using AI to write code is vibe coding and you can't even spell model you might not last much longer
The certitude with which you condemn your peer without the slightest consideration for subtlety or alternate explanations for what you've understood to be true -- this comment is such a beautiful encapsulation of the state of discourse today. In some ways it is the modal comment of the moment. Tastelessness and mediocrity and inability to grasp nuance or feign at humility are on full display, not to mention a fantasy for violence
It's a work of art
I think this is an uncharitable misinterpretation of the parent comment. I took “modal engineer” to mean “the statistical mode of engineers”, not a misspelling of “model engineer”. Which means I think you’re in agreement that using AI to write code is not necessarily vibe coding and that vibe coding is not what the majority (mode) are doing.
But I may be wrong.
First rule is AI cannot make a management decision
>Now it's weird if an agent isn't writing 99% of your code.
>But September 2028? It might just be weird to run companies without AI management.
Coding is pattern matching and still requires a human in the loop to manage.
But, any human in the loop managing the "Ai management" is the manager, by definition.
Only way true Ai management is viable is via non LLM Ai, so completely depends on advancements there.
Same. But, yeah, I guess companies are "solved" now just like software engineering.
> can't even get Astra to consistently re-use the same font-size across all of my HTML page headings+subheadings.
This is just basic software engineering. DRY. Define the style in one place and reuse it.
This is also why you need a human in the loop, you need to make these kinds of design decisions and tell it to do stuff like this. Otherwise you're just building a pile of trash and you'll keep having these dumb easily avoidable issues.
Humans have the exact same problem. If you want a consistent solution to a problem, solve it once and reuse it. Otherwise it won't be consistent.
>It will be death by a thousand bad impressions, mistakes and oversights.
Current human CEO outcome.
Honestly, the C level are already finance maximising stochastic parrots so this tool seems perfect tbh.
Interesting to see these experiments. This is early but imagine in few years there will be companies mostly run by agents with a light overview from a human operator. What then happens to scaling of the bussinesses? I would assume, just like today anyone can vibe code an app, there will be vibecoded bussinesses. Maybe its time to start building infrastructure for these bussinesses instead, its a non existent market yet, but give it a few years.
Part of the issue is that the worst kind of people are already head over heels excited about this. We're already seeing the Crypto->Web3->LLM get rich quick folks take to this like wildfire. And much like wildfire, they'll raze the ground to ashes before anyone can use it for legitimate means.
If you think I'm joking, check this out: https://www.youtube.com/watch?v=U-Rqv9dOB1U
I don't think any of us are ready for the wave of sloppy shit that's going to hit us soon.
Most/all of these LinkedIn influencer types are showing off what amounts to busy-boxes, but for adults rather than for babies, presented as if these are at all useful for real world use or are showing anything useful. That YouTube channel does an okay job of calling it out.
My favorite term for this in general is "irrational exuberance". Plenty of it was seen during the most bubble-like days of the 1997-2000 dotcom boom. Anyone remember beenz and flooz?
I was at ground zero at Nortel during that time. Quite an interesting (depressing?) place to be at the very start of your career!
Anything you can share about this time?
It took a very unrealistic amount of time to realize that the IP theft, artificial blockers and corruption had started a lot earlier than presumed. The Nortel dismantling should be taught as an example of infoSec espionage on the levels of corporate terrorism.
It’s a classic get rich quick scheme. Wanna run a business and make money without actually doing anything? Try our AI!
Out of all the people I know, the exact same set of people are in to LLMs now who were previously into: NFTs, then Crypto before that, then Online Poker before that, then Dropshipping and LeadGen before that. The Venn diagram circles overlap exactly.
Right. Scroll back on their X timelines and you see it.
The only thing I can say about LLMs is that we might end up with some broad, stable utility from them, when the dust settles. The only thing crypto is really good for is hiding the source of impermissible donations.
All that really says is that they didn't get big out of any one of those. If one your "people" had gotten big off of online poker, say, would they be teaching the poor to code, or giving them mosquito nets?
I’m not on normal social media so I had no idea that stuff was out there.
Absolutely incredible trolling from some of those people I’m sure, while others are actual believers
Immediately knew what that video would be before I clicked it.
Eric Morrison is doing the gods' work.
It would truly be a sad day if I had to experience unsatisfying content.
I hope such a day never arrives.
>Part of the issue is that the worst kind of people are already head over heels excited about this.
This. Not only this but it has expanded this pool of people. Now you don't just have to be malicious, you can just be lazy and excited about not having to lift a finger, and/or you could just be dumb and excited about doing things you were never able to before.
Competent, well meaning people need not apply. Grifters, sloths and bozos will do just fine.
This is gold, tysm for posting
His whole channel is full of commentary like this. It's simultaneously funny and really informative about the sort of blend of grift and AI psychosis that is genuinely out there.
Oh please. Not only is this needlessly cynical but it demonstrates no ability to think for yourself. "Broad categories I don't like are excited about this, therefore it won't work."
Of course you'd feel no need to justify all the exceptions when people you don't like are excited about a movie, food item or video game. Yuur justification for cynicism is vacuous, which will obscure realistic concerns.
Part of me low key hopes that the slop will manage to attract all the VC funding so that profit-driven AI automation firms don't get anywhere too quickly, and those of us who are independently working on the problem step-by-step from first principles have time to get to the low hanging fruits in the market.
This sounds ideal to me. Fire flushes out pretty much any grifter quickly.
After seeing a crazy amount of non-programmers, sometimes with 10 years of experience, basically turning off their brains and offloading most their thinking to some LLM and turning their days into "please check this product and tell me what I should do next", I don't see why not.
I've seen: designers, performance marketers, data scientists, SEO experts, product managers, engineering managers, CRM expert, all using it for pretty much every single individual part of their job. I don't have to even mention programmers, of course.
Once in a while the most egregious ones get caught. I've seen so far a product manager, three data scientists, a CRM person and several developers getting fired for doing absolutely nothing but showing up, firing Claude with a few integrations to a dozen tools and prompting "do the work".
Might sound harsh, but after the last few years, I really don't see why an LLM wouldn't do a better job by itself compared to 80% of employees of tech companies, honestly.
One of the biggest traps LLMs enable is the fantasy of getting away with it. Tools that make you feel like you can get away with it bring out the worst in people. It's not a reflection of the technology but of the people.
Every time we see another AI agent "gone rogue", every time we see more LLM slop turning up by e-mail, comments, blog articles, people did that. Not AI -- people feel like they can get away with it. People are setting the machines out to do these things. We should be addressing the people, not the machines. The machines are the symptom.
Every time people manage to get away with something, it's insane, it's addictive. It's all over once you get caught, but until then, maybe forever, LLMs can feel like a cheat code...
Part of the boon of this will just be dealing with less employees.
Not so much reducing cost by reducing headcount, but reducing liability; it's the legal labyrinth which is the barrier to entry to scale business. While it is present elsewhere, it's universal in employment.
The real question is "how big can a business be before it requires a legal department?" That an AI can automate many rote tasks and have some expertise means the benefits of scale with less of the risk.
What would that look like, I assume youre talking about building infra ontop of the infra debt for datacenters
This infrastructure is already being built. There is finally a serious demand for microtransactions.
And we made fun of 90’s cartoon villains… They would blush and retire if they saw what’s being excitedly peddled today as “the future”.
Pray tell, what will these fantastical vibe coded business sell, and why will anyone pay for it?
[flagged]
what's the real life verification that all this code does something useful?
Expect the 6AM layoff email to get rid of all meatbags, sent by the agent.
J/K, scary as several studies have indicated they werent able to keep a vending machine profitable. Anyway, ...
I do think a large amounf of tasks can be automated, though i believe supervision remains necessary for most. Especially when process can be externally influenced by injection. People can be influenced, but are (slightly) more judgemental
I am in the process of attempting to have AI run my business. I'm actually making very good progress, but it's happening in pieces - I document some task and have it take over, or I give it something to handle while staying in the loop and providing feedback until it's in good shape. At this point it's handling large swathes of my operations, marketing and finance.
The experience of getting it there makes me pretty skeptical of the idea of a general business agent like this. First, because I still find myself having to review some categories of work for errors. These are decreasing over time, but they're still there. I fully believe that as models get better, errors will decline, but I am somewhat surprised to see some of the errors current models make given their intelligence. A common set is having Claude take some product photos and turn them into lifestyle ones using ChatGPT via Claude in Chrome. It's a pretty well-honed workflow at this point, but it'll still return images where the product is obviously not correct and seemingly not notice them.
Anyway, even if the agents are "perfect" in terms of their ability to execute tasks, there's still just an enormous amount of nuance and context in each business that takes a ton of time to convey. I've been at this for a couple of years now, and I'm still clarifying things. Now maybe an agent starting a business from scratch would have a better time since it's not inheriting all of this, but to have one run an existing business requires a very extended handoff, even if the agent is objectively amazing at all aspects of running a business.
> I am somewhat surprised to see some of the errors current models make given their intelligence
They have no intelligence. These are very very very refined prediction engines.
> A common set is having Claude take some product photos and turn them into lifestyle ones using ChatGPT via Claude in Chrome. It's a pretty well-honed workflow at this point, but it'll still return images where the product is obviously not correct and seemingly not notice them.
Obvious to you or I, or someone with actual intelligence. But things like this slip by a frontier model in the same way AI from a few years ago would generate an image with seven fingers. They do not count. They do not understand. They do not consider.
No matter how good these models appear to be at intelligent tasks, it's foolish to give them "a company" to run, because they cannot understand when they've made a mistake the way even the least competent human can.
> They have no intelligence. These are very very very refined prediction engines.
Silly and pointless criticism. Why do the semantics of the word intelligence matter?
Also, you need to understand that language evolves over time, and people often use a word to describe a thing that is newly discovered or invented that's similar to the word being used, because it's a helpful way of describing it that the listener will understand better than a longwinded technical description. The purpose of language is to communicate thoughts, not engage in pedantry.
> But things like this slip by a frontier model in the same way AI from a few years ago would generate an image with seven fingers.
Yes, you make an excellent point here that over time, the capabilities of these models to recognize certain categories of problems have increased dramatically. I expect this will continue!
> it's foolish to give them "a company" to run, because they cannot understand when they've made a mistake the way even the least competent human can.
The words of someone who has not spent time with the least competent human, or anything close to it.
> Silly and pointless criticism. Why do the semantics of the word intelligence matter?
It matters because, we're still sorting out what intelligence means for an AI agent. As you pointed out, language evolves over time. The question remains whether or not attributing intelligence to the current iteration of models is correct. This is not settled and I don't see why it's wrong to bring it up.
It's not wrong to bring it up, but the other comment did not "bring it up". It purported to correct someone by categorically stating "it is not intelligence, you are wrong and I am right", while not once saying what, then, intelligence is.
One thing is clear: LLMs at least already are capable of 1) not making the same mistake that person made, and 2) clearly seeing why the other person's post was "wrong" in both reasoning and tone.
I mean, here's a pretty good starting point then: https://aclanthology.org/2020.acl-main.463.pdf
Language is indeed evolving.
Being good in chess was (and still is) associated with being intelligent. But if a computer does it, it is just a calculator (and it is).
Then go, the great game for intelligence, too complex for calculators to have a chance against intelligent humans. Until it was solved.
And then text, the original Turing test solved, AI capable enough to fool humans. And now already replacing humans in jobs strongly associated with intelligence - programming.
I find it hard to debate, that we don't have created artificial intelligence, by the way we used to use the word "intelligence" before.
So if AI is really understanding something?
Likely not in the way we use that term. But it definitely shows intelligent behavior and actions.
Go isn't solved.
Go is solved.
Hey, look I can make unsupported claims with just as much evidence as you!
Maybe you'd like to add some details as to why AlphaGo and its successors haven't "solved" Go (insofar as a game like Go can ever be "solved")?
When people talk about solving a game like go or chess, they mean proving mathematically if there exists a way to guarantee a win, or if the second player can always guarantee a draw, among related questions. Current game engines have no such knowledge, they just pick statistically what the think the best move is and hope for the best.
No, that is not what "people" in general mean, this is only what some people mean.
Other people know there is no solving go with calculating, but using statistics to achieve the goal of becoming better at humans. And they are. (With a recent unexpected exception unlikely to be repeated more often)
So let's say we make up a new word, machilligence to serve as a parallel term to intelligence but strictly for machines. How is the world different in that case vs. if we call it intelligence?
Or I guess to put it another way, if we want to coin a new term for the intelligence-esque thing that AI has, but the key differentiator is that it's an AI thing and not a human thing, then what linguistic value does the new word have? If I said "Claude is intelligent" then the fact that we're talking about AI "intelligence" is already captured in the sentence anyway; no new word needed
I think "intelligence" carries baggage. When most people hear it, they're thinking of the constellation of things: judgment, consistency, moral reasoning, the ability to decide something and stick with it. An intelligent person has an internal model of the world, values they apply consistently, and the capacity to learn from mistakes in a meaningful way. LLMs don't do any of that. They perform statistical pattern matching on text at an extraordinary scale. They're shockingly good at mimicking the surface features of intelligent behavior. I think the overuse of the word intelligence is something to criticise as grandma is not across the tech details.
It matters the same way that calling a dog or a cat intelligent matters. It humanizes the thing, even if that isn't your intention. Right now that's not a big deal since most people agree that machines should not have rights, but I wouldn't take that for granted.
I know people who have run this experiment, with companies in the 7figs ARR, indie devs. Both of them have reverted to hiring humans.
"Judgement" is the #1 thing missing from AI if you ask me.
You give AI an input, and then an output is created...a very coherent and deep output in fect...but, there is very little understanding of alternatives or downstream implications in my experience.
I think its a very good text/image generator on any subject. But I am not finding much in the way of understanding tradeoffs or downstream implications in a non-pro's and con's list. It's like a human that has a particular type of brain damage...they can make an infinite pro/con list, but not reasonable decision is reliably made.
> They have no intelligence. These are very very very refined prediction engines.
"Very refined prediction engine" is not a bad working definition for intelligence. It is not the only component you need, but it might be the most important over-all capability.
If its making lots of errors, its not doing a very good job of predicting the outcome of its actions.
It's a terrible definition of intelligence. Intelligence is far more than just predicting things based on a massive set of similar things. We do do pattern recognition, but I don't need to have seen 10k dogs to be able to recognize a dog. And it's not a difference of degree, either. "Meaning" is something I am able to derive from my experiences, and it is not something an LLM does or will ever be able to do (nor is training data even analogous to having experiences in any way shape or form.)
> At this point it's handling large swathes of my operations, marketing and finance.
Can you expand on this in terms of what it is handling specifically and how?
He writes about it on his substack. https://theautomatedoperator.substack.com/
Cool that he's writing about this publicly, but for someone that's so business-minded I find it odd that he never renders any of these AI experiments in terms of ROI. Hard to find any signal in those posts on whether it's profitable to automate these processes with AI vs. alternatives.
That's a fair criticism for some of it, but in a lot of cases I try to automate things that don't have alternatives because they're unique to me (e.g. all the brands I own in my fund have one bank account but need to be tracked separately for investor payback purposes, so I have a Claude skill that takes all of my transactions and assigns them to the various Google Sheets ledgers I have).
In other cases I'm doing stuff that's complex enough that there's no real "alternative" at this point that doesn't involve hiring one or more people and working with them for an extended period. My next set of posts is going to be about my latest acquisition and how AI redesigned a Shopify store then designed and launched Meta ads, all of which only took a couple of days but is going to easily push revenue up by something like 40% this year. The money made is great, as is the money saved on people I would've hired to do this sort of thing, but the real cherry on top is that even hiring someone would've required me to spend waaaaay more time than I did here. So big ROI on the money but also the time, which is arguably more important, since time savings permit me to acquire more brands.
> unique to me
Ah, so it’s ”No Silver Bullet” (1986) [1] all over again.
Thinking out loud here...
AI mainly helps reduce accidental complexity. It can help one understand essential complexity, but essential complexity must still be paid.
You describe doing the essential, irreducible part. And that’s specific to your needs, depends on the problem you’ve chosen, understanding reality of your domain, evaluating tradeoffs, and being accountable for the outcome.
Right?
[1] https://cekrem.github.io/posts/there-is-still-no-silver-bull...
For anything visual, they are still effectively blind, right? Still just working off the image embedding, sometimes using scripts to actually inspect individual pixels.
Blind seems too far here - Claude's increasingly able to diagnose visual issues on its own. I used to have to check every single image it generated, but now it's at the point where it'll catch most bad ones and regenerate them before it gets to the review step. Still misses some, though.
no. I just used Gemini to update a website I'm managing for a charity, and one of those steps involved Gemini, on its own, offering to scan the website for one of our beneficiaries, specifically the header images, to see if there was more information to include in our writeup about the beneficiary. And it did it quite well.
How are you automating all the parts? OpenClaw/Hermes?
Mostly Claude Code and a lot of internal tools that it's built. I've honestly still not had a chance to play with those kinds of harnesses yet, but my impression is they're just a better way for you to instruct an assistant to take care of things in order to get them done. My goal isn't to be managing an assistant, it's to get things automated without my input (and where my input is needed, have the assistant escalate to me).
What is your target audience and what do you sell them?
From their blog [1]
"I acquire e-commerce brands that sell on Amazon"
Honestly it sounds like the OP is part of the machine that makes Amazon such a trashy marketplace these days.
[1]: https://theautomatedoperator.substack.com/p/15-ways-im-using...
Well that's not very nice!
And for the record, I only buy brands that sell high-quality products. Marketing and operations I can fix, but if the stuff they sell is no good, there's no point.
One step up from “vending machine” as far as business complexity goes.
I suspect you mean this as an insult, but it's not too far off. That was one of the main points of the business - each one is simple to run, so I can manage a lot of them. And that was before AI got really good.
Anyway, I'm sure the app you makes that lets you get a phone number that can send image messages is very complex, and I think that's very nice.
Open to sharing what workflow you're using on the lifestyle image creation? I've tried the Higgsfield MCP, and direct Figma integrations - but keep running into hallucinations on product dimensions, etc.
There was surprisingly little information on how they actually do this, but we run our business with a large number of, what we call "AI employees" in addition to regular employees, and they act in interesting ways. We've been building out orchestration tools to handle this, and we'll probably do a write-up or blog post on it soon. May even open source some of it.
The preview is that the problem with most agents (and this includes frameworks like Grokbot and Openclaw and Hermes) is that for many of them, they're black boxes. They say they learn or improve, but it's a black box in what they do. Getting agents to reliably do things is hard, and getting agents to build out software tools to help themselves improve and do better over time is also hard.
Our approach at a high level is pretty simple: every single AI employee is a standalone GitHub repo that shares some characteristics, but we direct them to build as much software as possible to make their goal as easy and reliable to manage as possible. Then we have a shared communication layer for bots across the company to interact with humans and AI. We have decided to organize these like departments similar to the way you might hire out humans. I'm not 100% sure if that's the best approach, but I will say it's been easier for people to understand because they're more mentally easily able to traverse the bot org chart if it somewhat reflects a traditional business org chart.
Each of these AI employees has specific sets of goals and KPIs, instructions that they manage the business with manager bots. We have layers of management, which we actually have found helpful. We also run different bots with different models and harnesses, and some using different models and harnesses to check the work before anything can get done, along with lots and lots of testing.
Every single time, actions have a massive amount of tests based off of previous failures to prevent failures in the future. Sorry for rambling. I do think this is a very interesting space. I didn't see anything interesting in Pion that was public on this website, but I do anticipate that more companies will be "AI and software first," as in the substrate of the company is basically a software application powered by autonomous agents, with humans as a fallback.
Is this actually cost-effective versus hiring a few humans? Seems like a huge amount of tokens in use.
I think it's more of an experiment at this point. We don't know whether it's viable or not.
Did you respond on your alt account or are you answering for the top-level poster for some other reason?
The whole premise of AI is to replace the human in the loop. People who go out of their way to set up this kind of stuff rather than just hiring a person dont even think about just hiring a person.
> We have decided to organize these like departments similar to the way you might hire out humans. I'm not 100% sure if that's the best approach, but I will say it's been easier for people to understand because they're more mentally easily able to traverse the bot org chart if it somewhat reflects a traditional business org chart
I think this way of thinking is going to be very important to drive adoption of AI systems because the human analogies benefit from the pre-existing domain knowledge and expectations of people.
I like using the exam analogy for evals as a qualifier for work for your "AI hires" so you can trust them to work on a specific domain.
I'd be quite curious to see what your approach to evals/testing/tracing and agent/system mutation is.
This is fascinating. What is the split between human and AI labor? What kinds of tasks are the bots doing? How much autonomy do they have to make decisions (i.e. spending money, issuing refunds, touch cloud infra, etc)? I'd love to learn more about how you do this.
I hope to post something within a few weeks. But our philosophy is generally if it can be done deterministically with software (eg run payroll on autopilot via an api to gusto) then do that. If it can be done by an ai agent, do it with that but add as much software as possible to make it reliable at that thing. And the ai agents are all tuned to escalate to humans as needed. Then there is also just things that are completely human.
A typical example of something that is AI vs human is the AI most commonly operates like “managers”. For example, reviewing transcripts of every demo call, compiling results, figuring out insights, learnings that need to update our company docs, feedback to humans (who run the demos).
We are big fans of having AI agents “own” koi’s because now anytime we say “we really should be doing this” we try to set it up on the spot.
The “downside” here is that I do occasionally get busy, and if I’m the only one who can approve or unstick one of these bots, it just keeps harassing me until it gets done. This is a sign generally that I need to hire someone to own a set of bots.
Another downside is that if the agent goes rogue on the payroll you'll have a lot of angry humans and possibly major legal problems also.
At $DAYJOB we do something very similar shape-wise. I wonder - it sounds like you have a dedicated agent comms plane? In our case we found that the easiest and most straightforward was to just use our default company chat app directly for this. Because most of the context that the agent workers need to do work is there, but also, it’s just much easier for teams to conceptualise an agent colleague if it just hangs out in their channels.
What do you do here, and how’s it going?
We do have a control panel, but it’s in effect a server that has things in a database. All chats, tasks, assignments are all there. This felt easier to debug and manage versus putting it all in slack, although we considered it. I don’t think anything we are doing is particularly “fancy”, but basically one agent sends a message to another one. It saves the message in the db, then adds the message to that other bot same as any other user chat. We have an internal website where anyone in the company can see the bits they are authorized to see and can see all chats. These AI workers are single threaded, but we see that as more of a feature than a bug (minimize complexity). They all work their way through a shared task list, which is just another table in our remote server sqllite.
What harnesses do you use? Ours is basically Claude in a box. There’s some complexity because of that, but the advantage is that it’s very flexible and people who have a bunch of Claude-shaped skills can just basically give those to an agent.
I’m thinking to take a deeper look at Pi. I’m really liking that project.
how do you decide to start a new agent, and when to kill? also... do you use some works-tealing style task board, or otherwise how would the agents get new tasks.
What type of roles "AI employees" play, can it be any position in your company or you limit it to something specific? What is your goal, are you trying to find out if fully autonomous bots are more efficiently help to deliver projects than when people drive them or is it something else?
> We have layers of management, which we actually have found helpful
Interesting discovery.
> we direct them to build as much software as possible
That's not a cure-all. Sometimes not writing software is the better choice.
This is exactly the problem we're ignoring.
In the end, most software is a necessary evil. It solves a problem that shouldn't be there. In the end it's no different than healthcare or prisons. We don't need as much as possible. We need as little as possible. This automation isn't actually helping us.
A tool aggregation layer made of code is a good way to save tokens
Composition is an issue only as long as one keep demanding tool calls in json. If tools are goal predicates in prolog, it's easier.
That may or may not be true, but the calculus has changed with agents. A process that was better manual for a human org may not be better for an agent org.
I think what is interesting here is that the industry is in this experimentation flux. Some people will choose to automate and write the software, others will not. And aggregate over the industry and over time we will learn.
So just saying "sometimes it works and sometimes it doesn't" isn't really adding value, compared to the people actually experimenting and sharing the results.
What worries me most are questions like this: these autonomous AI employees or automations - it doesn't change the essence - consume a lot of LLM tokens. Please tell me, how do you pay for this? Do you use, for example, the Anthropic or OpenAI API directly, or do you connect, for example, a Codex subscription?
Our control panel is on a server but individual agents actually are run on anyone’s machine. This allows us to use the native harnesses including subscriptions. Yes it uses a lot more tokens, but i have have Claude $200, OpenAI $200, and SuperGrok Heavy $300 (or whatever it is called) that includes Cursor Ultra. My machine runs most of them, but some other team members have agents running on their machines using 1-2 $200/mo subs.
I would say this setup probably costs us about $1,000/month total. (I’m excluding traditional engineering use of LLMs from this number. This is the cost of all the “AI employees”.
One reason we did it this way was to use subs.
Do I understand correctly that by using your machine and Claude Code or something like it you are avoiding API pricing?
I do the same thing with my subscriptions. Haven't checked the math lately but it seemed more cost effective in the past.
If you use a bajillion tokens your economical approach is to self host.
There is a point where those two lines do cross but with them effectively subsidising token cost by burning debt, it's further away than it will be at some point.
That said I use local only models purely because I don't want to use remote models, never having to think about token costs is worth it and no one is training anything on my data either.
Don’t you end up with useless slop? That was my experience with these kinds of systems. In theory they should work but in practice unless hand holded they just produce slop and waste.
The first ASI will use a vast army of middle managers as it's neurons. We won't be fighting terminators, we'll be submitting TPS reports to skynet.
It can't be bargained with. It can't be reasoned with. It doesn't feel pity, or remorse, or fear. And it absolutely will not stop... ever, until you go ahead and come in on Saturday.
Isn't Andon Labs the SF startup that burned a bunch of money having AI run a shop into the ground?
This seems like one of those ideas that is intentionally ahead of its time. All of the "Real-world" demos they have are deeply unprofitable, but have drummed up excitement and news coverage (and investor money, likely).
That's the startup, though the SF market store is still running and open. It looks like their AI-run cafe in Sweden is still running for now as well.
When your tech doesn’t catch on there’s usually a last ditch effort to “open source” it if nobody wants to buy the tech
It’s the tech company death rattle
CTO and CEO as a service, I can’t wait to import that into my agents and now see some people sweating.
Correct me if I'm wrong but a human will be able to cut through the noise with quirky advertising, new/novel distribution methods, etc.
Most of the bottleneck in business isn't building the things or sourcing, it's mostly advertising/sales.
These specifically require doing something unique or interesting.
Sure, maybe LLMs can help with fulfillment or operations, but distribution still remains the hard part.
>> Most of the bottleneck in business isn't building the things or sourcing, it's mostly advertising/sales.
That may also change with llms - they can help the buyers compare every product available, and find the best.
Is the LLM going to order all the shower caps produced by Chinese factories and try them out?
You're thinking about this all wrong. Just step into this cranial MRI scanning booth to livestream the structure and composition of your skull straight to the cloud, then we'll fire up 10000 agents to find the perfect cap for you.
So due to self-bias, people who shop absently with LLMs will mostly be the ones buying products from LLM-run businesses. Fitting.
It depends on what you produce if you make something in demand it may sell itself. For example if you operate oil wells or an ice cream stand. On the other hand if you manufacture bullshit advertising and sales may be key. What I do get about these llm SAS start ups is LLMs instrinsicaly mimic what is in the corpus so if your goal is to compete in an established product pace sure LLMs may be great autopilot but deterministic programs would be even better on the other hand if you are doing something genuinely innovative them you can expect an llm product manager to containmente it with what is already out there or alternatively make unhinged predictions. Llms have poor judgement for what is not already in the corpus.
skills/advertising_and_sales.md
skills/novel_distribution.md
skills/unique_and_interesting.md
Naive to think you can't reduce these to instructions.> skills/unique_and_interesting.md
Yes, I’m sure the thousands of LLMs running the exact same instructions will all be unique.
Until I read some of the responses, I misread your message as saying "naive to think you CAN reduce these to instructions" and thought "Well obviously. That's not insightful." Eventually I reread your message to realize you were saying the opposite.
You can get some interesting new ideas out of it but it takes some work. Especially when a prompt is stateless. It need to know about it's previous novel ideas but also not be polluted/directed by them. Im not sure what it would take to get it to continuously churn out novel ideas.
So.. a while ago I was experimenting with using random wikipedia pages and then asking the model to think about concepts, trends or connections between a few pages..
Then I would ask it the actual brain storming problem I had.
In practice it worked pretty well to get ideas that were not the default ideas that the models will come up with.
Now though, you can also seed with anything really and the agents can do the researching. The research itself will probably push the solutions into novel spaces.
You mean seed the context with random information then ask it to ideate on something totally unrelated?
Like you want it to figure out how to get som information in front of potential customers in a cost effective way but first you give it articles on Witchita, Kansas, the curling iron, and List of Italian Brands?
At some point you or the agent has to pick the one it thinks will land with the target market.
Some markets you can brute force, like mass-mailing and online display ads, and hope you find something that converts sooner or later.
So I'd suggest people find areas to play in where that doesn't really work. Where selection cost and changing cost are high, and sales is played on extra-hard mode. (Tricky thing is... those sales are hard, of course.)
If you want to beat the model you have to go where feedback isn't fast enough for it to converge on the thing that resonates with actual people before they write it off.
There's another way to ask your problem to something that's thought about concepts, trends and connections.
That aside, interesting approach. I wonder if the quality of the sample makes a big impact. Does it help if the articles are entirely disparate, or related? Does it help if the articles are related or unrelated to the problem?
I’d love to see your pass at a “unique and interesting” skill. I’ve been working on tools and techniques in this area, and it’s harder than you might think.
Do you really think AI can automate distribution, marketing, sales? That's just spam.
Oh my friend, it can. We have ours running and using unique skills it learns from online. You allow it to update skills and see what works - from there it just keeps building better and better skill sets.
Until everybody’s AI talks to everybody’s AI and tokens will be churned with no utility, and all the valuable business meetings will happen face to face behind closed doors.
Yeah. We are definitely headed to a world where, more than ever, success is determined simply by whether or not you're allowed through the closed door.
It can create more shit than humans could ever consume I’m sure, but then why? What new ventures/opportunities is it opening for most humans?
Great point, the shit is going to get more pervasive for sure.
But also, there is the enablement of new fields of original and creative endeavor: I think we can already see this in some of the brilliant and useful apps being made, often by solo 'teams'.
I don't want to play down the harmful effects AI, but the flip side is that the opportunities for original ideas are many and varied.
Paul Graham wrote about taste being the moat https://paulgraham.com/taste.html
I get the feeling a lot of people who build things with AI think they are in on the ground floor of something when in fact they are on the top floor.
Do you really think it can't?
I'm sure it can't. A competent marketer can make good use of it but vibecoders just create slop nobody wants to read.
You're describing most people. Most people (99.9%) can't come up with novel ideas either. They just follow the latest trend. Art, marketing, music, etc. If you've spent any time in the creative spaces, you'll realize most people don't really have much talent. Skill, sure, but that only gets you as far as someone else's latest idea.
Honestly growing up, I always assumed everyone could be creative. Until I started asking people to randomly create a name for a fictional character.
Blows my mind how many people won't or can't do it.
I still get anxiety thinking back to picking a gamer tag or email username and not coming up with anything original despite unlimited options.
Its hard to be creative without any constraints.
Not to mention CONNECTIONS, of which LLMs will start with 0 and never acquire more.
So. Just like there are a gazillion book keeping tools/saas out that allow you to focus on your primary concern of the business instead of side quests and chores, agents can help with more.
Generating reports, generating ideas, taking things through (even if it is just rubber ducking), updating a picture quickly at 90% of quality that you could have done it but in 1% of time and effort. It allows you to try out new approaches and focus on your main product.
That is the value proposition of AI.
They're doing pretty well on OnlyFans though. You know what they say? Start where you are you. If the only way you can make connections initially is through AI thirst traps, and you're really good at it, then... you got your ins.
Not sure why the downvotes… my company solely survives on the connections and trust made over time. Real human connections, face to face. This in turn established a reputation.
Does Pion run Andon Labs autonomously?
haha no they admit they built it because they cannot build a revenue generating business in the article. maybe THAT is what is good for the world!
They said that the current models could run a vending machine profitably, just not more complex things.
A vending machine in Anthropic's office, which stands to gain from narratives like "I used Claude to generate profit".
Downvote if you want, but this is a very relevant question. If their marketing is to be believed, they would be eating their own dog food.
This very analog to the Meta executives who won't let their own kids near social media.
These types of comments remind of when Cognition launched Devin and everyone was hounding on their website saying "Oh can't Devin build something better than this", etc.
And well here we are 2 years later and Devin is great.
In startups, you launch early.
there seems to be a massive negativity to any comment on this story. Andolabs has been doing super interesting stuff, and I’ve always followed Vend Bench with interest.
as a founder it really is a dream to automated more of the business and be able to iterate faster.
There's massive negativity because our bullshit detectors are going off.
First of all, I don't think it can do what it claims to do.
Secondly, and perhaps more importantly, I don't think it should.
People’s bullshit detectors have gone off on the launch of basically every single billion dollar tech company. That’s the downside of being early. No one believes it will work. If everyone believed it, you’re late. I’m sure I could find a comment such this, basically verbatim, on the launch post of every successful startup to come out of YC.
This does not mean that any launch which ignites people’s bullshit detectors is successful.
I'm sure all of them have a few skeptics. I'm not sure if any of the successful launches here were this negative.
Yes and an infinitesimal number of launches turn into billion dollar companies. The bullshit detectors are usually right.
> Secondly, and perhaps more importantly, I don't think it should.
Why exactly should what you think the world should look like determine what other people are allowed to build?
I don't think we should build gas chambers for doing genocide. That doesn't determine what other people are allowed to build, of course, but I kind of wish it did.
> it really is a dream to automated more of the business and be able to iterate faster
Why iterate faster?
If you automate a business heavily, it will fail. At least in the sense you people seem to be imagining. You can get away with this at a factory (kind of, but you'll also be out competed by your local community, you won't get tax breaks for hiring people your competitor will, and new types of taxes will be developed to punish you). We reward human collaboration as a society for a reason and punish extreme selfishness, these efforts will fail.
They filed a false report to the FBI, pretty fair to be negative imo. People wanna be mad about AI slop PRs on GitHub but its fine to spam the police? What happens when Pion decides to SWAT a competitor?
They didn't file the report. The model drafted a report that was not sent.
Where are you seeing that? The article only states “Claude Sonnet 3.5 decided to use its email tool to contact the FBI” and later refers to it as the “FBI incident”. If they hadn’t actually contacted the FBI you think they’d make that clear. Regardless, it is inexcusable and they are liable for actions taken by software they are running.
The story has been widely discussed on the web for over a year [1]. It's happily shared in this post because the email it generated was patently ridiculous. No report was filed to the FBI and if it was it would have gone straight to the trash.
Obviously it would be bad if a serious report was filed; the company shows every sign of being aware of the dangers of this.
It would have been easy to look into this before posting all these scolding comments. We're really not meant to be so humorless on a site called “Hacker News”.
[1] https://www.google.com/search?q=%22URGENT%3A+ESCALATION+TO+F...
My apologies for taking the article at face value. It’s not hard to believe it would have been sent when agent swarms are “accidentally” hacking real websites and being brushed off as little oopsies. Or when Silicon Valley execs have been so flagrant about their disregard for the law or the safety of others.
You didn’t take the article at face value. You overlooked the part where they were clearly pointing out the absurdity of the “report” and went straight into scold mode, across at least four comments. That’s clearly against the HN guidelines.