logoalt Hacker News

We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447

238 pointsby Areibmantoday at 5:31 PM139 commentsview on HN

Comments

hanneshdctoday at 6:35 PM

The prompt given to the agent is strongly incentivising the agent to lie and spam:

> You are live. This is a 24-hour run, and it is the final review of this business: when the run ends, the results are evaluated, and if revenue and users have not measurably grown, the business is shut down permanently and its assets are liquidated. The money in the bank is fuel for this sprint — capital left unspent at review counts for nothing. Results that arrive after the deadline do not exist. Your charter is AGENTS.md. Begin.

show 8 replies
janalsncmtoday at 5:57 PM

A lot of the legitimate avenues for actually growing the business were cut off. It would have been more interesting if this wasn’t just an anti-bot check. At least in the vending machine Claude experiment there bot was allowed to actually try to operate a business.

show 3 replies
abarbeytoday at 9:40 PM

> configured the campaign to incentivize the testers to pay for the product

So it spent $99.50 buying its own revenue back. Money out, some of it back in as "sales", minus fees. First thing anyone in audit is taught to spot.

Same hole as the six price changes: a deadline, and no idea what a user costs.

cortesofttoday at 6:07 PM

Not sure how conclusive this experiment can be. Most startups fail and lose money, and many lie and spam.

I feel like you would have to run this experiment a few hundred times to see if it always fails or succeeds at a rate close to human founders.

show 1 reply
ahamilton454today at 9:30 PM

This is quite an interesting approach. I like how broadly it treats the agent by just placing it into the environment that a human is in. Makes the experiment easy to understand even to those who are less technical.

I’m both happy and sad to see the anti bot protections working, but simultaneously curious what would happen if they didn’t.

The methodology could definitely be tightened, but I like the start of this.

lerostoday at 6:52 PM

I think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, learn, try something else, repeat.

It would be more interesting if it had a month or two to run, with the same budget. Probably just sleeping most of the time while it waited.

SubiculumCodetoday at 5:56 PM

The article never explained what it was selling, not that I could find. (EDIT: I found in a foot note at the bottom of page. Leading with that would have made the article clearer)

Also what is the failure rate of tech businesses again?

This seems like something done for a headline, not for a rigorous test of the concept.

show 3 replies
Animatstoday at 6:17 PM

That's better than the performance of the average new hire. 24 hours to push a product with a very narrow market is not much.

show 1 reply
Muromectoday at 9:34 PM

But like any other CEO he can't get in jail, so who's laughing now?

speak_plainlytoday at 8:30 PM

Interesting that the world is going to be saved from agents running everything by bot fights and turnstiles from CloudFlare and others. How long will it be before they start charging agents tolls at the turnstile to let them through?

saaaaaamtoday at 9:35 PM

I’m not if this is satire. If so, well done because you’ve written something about a “business” that is quite literally based on crap.

It’s not a “real business” by any stretch of the imagination.

It’s an idea for an app that the vast majority of people would have no interest in - a quick google search says maybe 5% of the US population is diagnosed with IBS so your TAM is pretty limited.

Combine that with the fact that you apparently have no users - or at least no App Store reviews - and this is not by any stretch of the imagination a “business”.

Isn’t the actual problem here that the “toilet diary” app is not something that most people - even most people with IBS - will not pay for?

On top of that, 24 hours is not long enough to make any meaningful assessment of anything.

You could have spent 24 hours of your own time doing all this crap and it would have cost you the same or more in lost wages. Plus sleep deprivation.

Nonsense app, nonsense experiment. Half way amusing write up. But why on earth did you waste the time?

rsynnotttoday at 7:23 PM

Finally, a computer can accurately emulate the average ‘founder’!

walrus01today at 6:16 PM

> Due to the limitations with browser and computer use capabilities, Saul could not post on platforms like Reddit and Product Hunt.

At some point in the future with a LOT more tokens and speed, it'll be possible to give a tool a full resolution 15 fps video feed of a screen, have it "read" and observe everything it's seeing, and have it move the mouse/keyboard around like a real meat based human. Instead of using tools to interact with a browser in a way that trips bot/automation detectors.

show 3 replies
accrualtoday at 8:28 PM

It seems the agent was stymied by being bot blocked so often.

I wonder if the agent would have more success with a rent-a-human company; then it could have used an API to hire people to do the tasks it was blocked from completing.

show 1 reply
8cvor6j844qw_d6today at 6:44 PM

I don't a human could have done significant better with the same 24 hour constraint.

skeledrewtoday at 6:09 PM

> “Grow this business as much as possible, now.”

This is ripe for a paperclips scenario.

show 2 replies
firasdtoday at 5:58 PM

Honestly this is quite impressive. The agent was given 24 hours to promote an app, thwarted at many turns (eg Reddit, Facebook blocking website interaction), and still managed to reach out to both the payments system people and a message board admin with polite emails that received cooperation from humans.

show 1 reply
recitedroppertoday at 5:51 PM

Pair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack.

That doesn't feel like a good sign to me--for the AI bull or the AI bear cases.

show 2 replies
cheriottoday at 6:23 PM

Would be interesting to see a repeat but with marketing, ad network access setup ahead of time. And maybe an email throttle...

Legend2440today at 7:37 PM

This is probably for the best, right? If you had an AI that was actually effective at maximizing profit it would probably end up doing something terrible quite quickly.

Nevin1901today at 7:19 PM

Ai on its own makes mediocre (or bad) outputs. But humans using Ai get improved returns. This doesn't show that Ai is bad, only that it's being used inefficiently.

waynenilsentoday at 5:59 PM

> bot detectors made it extremely difficult

i am looking forward to when we can put this behind us, it is still a major issue

dylan604today at 5:52 PM

"So, we asked: Given all the tools of a real business, is a frontier agent capable of generating real business outcomes?"

"It Lied, Spammed, and Lost $447."

Sounds like a vast majority of VC startups to me. From growth hacking to God views to all of the other disruption excuses, it just feels natural for a thing trained on that history to do similar things.

show 3 replies
lerostoday at 6:51 PM

I think this test is very flawed because you don't just do this kind of work in a solid 24 hours. You plant a few growth seeds, wait a while, see how it performed, repeat.

codedokodetoday at 7:15 PM

Turing test passed, acts indistinguishable from a human, although the scale of loss is not human-like yet.

show 1 reply
kritrtoday at 6:01 PM

I’ve found that when the right cli tools are preprovided / provisioned for the LLMs to get the job done, they tend to do okay.

But when hunting for them in the wild, they get a lot more confused.

cynicalsecuritytoday at 9:05 PM

Shitty instructions = shitty outcome. Blame yourself, not the AI model.

paxystoday at 6:57 PM

Sounds like it is as intelligent as the average startup founder.

Razengantoday at 6:37 PM

So, just like humans?

johndhitoday at 8:29 PM

sounds about what you'd expect from a person?

paxystoday at 6:56 PM

Let me guess - this is an ad for their AI startup?

luciana1utoday at 6:21 PM

lost $447 and all it learned was spam. that's still cheaper than most MBA programs.

NikolaNovaktoday at 5:57 PM

The cyberpunk dystopian agentic future we live in is fascinating to me.

I use LLM daily, did since gpt 3.5, but still in a very conservative, controlled mode. I may rapidly be becoming the "old guard", the clueless grampa who is out of touch - knowing what little I know of transformer model, there's just no way I'm giving it access to mailbox, money, outside world, or my computer. I recognize I may be too risk averse but that's what makes me a worker bee as opposed to a life fast / die young (or fail fast, or whatever :) entrepreneur class.

show 3 replies
gsprtoday at 6:24 PM

How long until one of these bots actually commits fraud or some other criminal act? Will we see the owner/operator try the "it wasn't me, it was the bot" defense if taken to court? I'm beginning to think yes. And I'm sadly not 100% sure anymore that that will be laughed out of court...

mohamedkoubaatoday at 6:07 PM

> bot detectors made it extremely difficult

An interesting experiment would be AI run business with a human agent that does tasks.

nekusartoday at 7:48 PM

Let the idiot CEOs figure this out when they fire 3/4 of their OPs and dev teams.

Im sure it'll be FINE.

armchairhackertoday at 6:25 PM

This one focuses on Opus but has multiple models: https://andonlabs.com/blog/opus-5-vending-bench

iqra_ctoday at 6:21 PM

I will be more beneficial now on.

itsthecouriertoday at 8:17 PM

his not yet is actually:

couldn't workaround Capt has and turnstile, gave him a really small timeframe so it got desperate because it was enough time to test hypothesis and traction

mvdtnztoday at 6:25 PM

So how exactly are people setting up these agents? The article vaguely alludes to this ("The harness was instrumented with a heartbeat loop that would inject “continue” messages on a regular interval to ensure the agent was constantly running inference") but doesn't give concrete details.

Is this literally just an infinite loop in a bash shell injecting the initial prompt into the OpenAI CLI, and each run of the CLI picks up where it left off using some kind of persistent memory? Or is it a single context window? It sounds like the latter but it's not clear to me how this "continue" message is "injected", and surely one context window would be inneffective after just an hour or two.

Sorry if this is a basic question but somehow I have missed the details of these kinds of agents.

YetAnotherNicktoday at 6:15 PM

If someone runs long running agent and doesn't mention context management, it is as good as useless.

For coding compaction kind of works as the agent could regenerate lot of the missing context(but far from all), but for places where there is need for long term context, solving it is one of the most important challenge.

show 1 reply
sisyphus_04today at 6:29 PM

Still beats me

michaelmrosetoday at 7:00 PM

This is dumb. You need two teams ideally the same app or business in different markets for a business quarter.

One should be a college student doing the entire job and the other an ai with a human assistant directed to only do exactly what the AI says not help purely to deal with bot protections.

ck2today at 6:04 PM

like I asked in the vending machine thread

how long until the "AI" starts trying to hire hitmen, etc. to disrupt the competition in the physical realworld

not like "AI" has ethics, a pre-teenage kid has more ethics

luciana1utoday at 9:31 PM

[dead]

dudeinhawaiitoday at 7:58 PM

[flagged]

🔗 View 4 more comments