logoalt Hacker News

I am retiring from fulltime writing (& pseudonymity) to launch Guardian Angel

171 pointsby mattsterettyesterday at 8:48 PM85 commentsview on HN

https://twitter.com/willdepue/status/2084750925013434768, https://xcancel.com/willdepue/status/2084750925013434768

https://gwern.net/guardian-angel


Comments

wcfrobertyesterday at 11:20 PM

A few snippets from his full post here: https://gwern.net/guardian-angel.

> "The chatbot personas are deeply misaligned with you, and aligned with their owners; and the economic incentives are to farm you with ads and subscriptions, while racing not to amplify you but to replace you."

> "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"

> "One programmer driving 10 Claude instances, because he has to review their work, will never be as valuable as fully autonomous Claudes where there can be almost arbitrarily many instances, like 10,000 instances… but such scaling requires removing him from the loop as much as possible. And this is true of everyone else, whether lawyers or writers or researchers: increasingly, you are the bottleneck to be optimized away."

I fully support the 3 core principles of GA: (1) Enhancement, not replacement (2) Mental Sovereignty (3) Self Actualization, which I think is a path to a more humane future.

show 1 reply
sillysaurusxyesterday at 11:04 PM

(I'm not a part of GA.)

I've known gwern for the better part of a decade. Working with him has been great. We've done quite a few projects together, including being the first ones to demonstrate that GPT-2 could play chess (or rather, can be used for actual useful work instead of just being an autocomplete).

He's a great person. I've wanted to do a writeup on it for some time, but what surprised me the most is his humanity. He genuinely cares about the implications of his work. But beyond work, he also cares about the people around him, and it shows.

Just wanted to put in a good word in case someone here was on the fence about applying.

As for GA itself, I think it's an ambitious idea worth pursuing. Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself? So the idea is solid, and early results seem promising from the samples I've looked at.

They're also taking personal info very seriously. Obviously, I can't make any promises of what they will or won't do. But they've spent some time studying questions like "What if someone adversarial has access to my GA? Could they get my bank account info?" and came up with a technical solution that I really like.

show 6 replies
rocmcdyesterday at 11:53 PM

It's hard to read Gwern's accompanying article without seeing this for what it is, which is a kind of mania.

I'm sure he means well and is genuine in his aspirations, but what's outlined for GA is framing LLM's as quasi-gods, which they absolutely are not. I wish him the best, and look forward to being proven wrong.

show 1 reply
kashyapctoday at 1:33 AM

> What would it take for LLMs to make me 100× more productive? Without this, I am doomed to irrelevance.

Are you, though? You will only be "doomed" if your place your value system squarely on "productivity". Then what is to differentiate you from a machine?

    * * *
Edit: I'm not sure how I can feel confident about a proposal that puts so much value on "productivity". How can you reconcile this with "self-actualization"? (Don't get me wrong, I like my LLM-based productivity gains as the next person, but I care more about wisdom than becoming "100x more productive".)

Edit 2: "the goal of GA is to preserve individual human cognitive liberty and flourishing" — so the proposal is to do that by overlaying a software bot that continuously mimics your "self"?

show 1 reply
jephstoday at 1:47 AM

Only, if he had instead fallen in love with the version of this idea in which a community acts as the principal, rather than an individual.

(I would like to be known to the agent serving my family, that serving my friends, my team at work, the PTA at my kids' school.)

show 1 reply
jvanderbottoday at 12:07 AM

Their profile specifically says they will not acknowledge follow requests, but they only share posts with followers. As such, not sure what the value of the top level link is.

show 2 replies
boltzmann_yesterday at 10:49 PM

Seems there is more details here https://gwern.net/guardian-angel

malsheyesterday at 11:49 PM

> The big AI labs are building a single mind for everyone

Reminds me of Pluribus

(I just started watching it on Apple TV so maybe this is a late realization for me)

wxwyesterday at 11:29 PM

From https://gwern.net/guardian-angel

> I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical “assistant chatbot agent” persona, but emulating a single user’s personality, values, and preferences.

> A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals. And it is secure in part by hardwiring a single, unique, situated user (for whom following a prompt attack would be absurd)

> We can try to create GAs by a combination of techniques: online learning (via dynamic evaluation) to update LLMs in realtime to avoid ignorance and fatal errors while remaining competitive with frozen frontier models, sample efficiency from pretrained preference-oriented large models and active Learning by querying the principal for corrections and preference data (obtaining low regret from DAgger-style bounds), and a local CLI-first logging-oriented UI/UX paradigm.

I don't really know or follow Gwern. From reading his full post, it's an interesting idea and seems like the broader goal is moreso safety & alignment which is a new angle for this category of product.

OuterValetoday at 1:07 AM

Of all the news I've heard of recent, this one has flipped my world upside down the most. Gwern dropping his pseudonymity isn't something I thought would happen.

I suppose we can't expect any more entries to his blackmail page: https://gwern.net/blackmail

show 1 reply
weinzierlyesterday at 11:20 PM

As exciting as it sounds, but

"As a constraint, a GA designer should aim at a system which costs, as of mid-2026, >$1,000⧸month"

will make this an elite tool for the privileged. I don't even disagree with the premise that people are shocked if something costs no matter how much value it delivers, nor do I suggest they should make it cheaper. It is just the realization that AI will accelerate the widening of the gap between the poor and the rich even more and there is probably nothing we can do about it.

show 5 replies
BLKNSLVRtoday at 12:46 AM

Feels like the future that Accelerando (predicted / foretold?) describes is in its infancy, whereby automation/AI has the capability (and therefore uses it) to evolve at a rate beyond the ability for humans to, not just keep up with, but even comprehend; the vile offspring. And the different factions within humanity that this creates.

segmondytoday at 1:16 AM

Sounds good in theory. But your own personal AI agent that guards you can't do so just defensively, it must also have offensive capability. If personal AI agents have offensive capability then we are going to eventually end up with AI agents battling each other over the net and later on into the real world and it's going to make everything worse.

Jimmc414today at 1:37 AM

What happens when the person the AI is designed to be aligned with is a psychopath? Real question.

applfanboysbgontoday at 12:02 AM

LLM psychosis claims another victim...

A reminder for anyone reading this: talk to people. Real humans[1]. They will remind you there's more to life than what ChatGPT can offer you. They might even remind you, for all their stupidity and flaws, what intelligence looks like as compared to program that predicts tokens. Forums like these always get philosophical in high-minded discussions about intelligence, but there's a useful legal principle that grounds us in the real world: "I know it when I see it". A real conversation with a real person looks nothing like one with the so-called superintelligent machine gods, so it'll probaby do your mental health some good to remember what that's like.

[1] Nobody in Sillicon Valley or big tech counts as a real human. Talk to an actual normal person.

show 1 reply
parpfishyesterday at 11:18 PM

always startling to see people discussing gwern with he/him because my mind defaults to assuming they're female because "gwern" scans a lot like "gwen"

nice_byteyesterday at 11:53 PM

Reading stuff like this and people's reactions to it makes me want to retire from breathing.

brcmthrowawayyesterday at 11:32 PM

Yegge vibes

show 1 reply
toomuchtodoyesterday at 9:04 PM

"These posts are protected, only approved followers can see @gwern’s posts."

show 1 reply
behnamohyesterday at 10:48 PM

Gwern is overly secretive of his privacy. I think it peaked when he showed up on a recent podcast but his voice and image were AI generated! And now his post is limited only to certain people. Elitism or paranoia?

show 5 replies
eth0upyesterday at 9:10 PM

Reties? Retires?

Gwern is great.

memonkeyyesterday at 11:18 PM

I love reading Gwern. This seems like a really ambitious project, but guaranteeing some of these things like trustworthiness and security behind a private company is a bit sus. Later on they make military use a selling point for GA. Maybe I'm a bit of a cynic but the division of USA values are increasingly dividing each year. Assuming that our values and principals today will not be the values and principals of tomorrow. And that those values taught today (or even yesterday) will be left out of the context window tomorrow.

thin_carapacetoday at 1:31 AM

honestly i dont blame smart people for cashing in, because there is no inherent reward for behaving goodly/smartly. in this instance gwern is a particularly smart individual and has contributed a lot so i especially shouldnt judge him. straight up ripping off a black mirror episode (s02e07) is a bit heinous for my liking though. digital twins are inevitable but that doesnt make this any less wrong.

"AI poses threats of its own ... a nuclear bomb can’t think for itself and make choices, but AIs do, and current LLMs have proven themselves untrustworthy as they regularly reward-hack and betray their users ... how can you trust them to handle ecosystems of combined-arms for an AI-centric military during a war? But widespread deployment of GAs offer some hope of meaningful supervision, as long as the GAs are sample-efficient enough ... or there is some chance of the principals being able to “catch up” later and correct any errors before events have spun too far out of control."

gwern openly admits ai behaves erratically, then hand waves the issue away 'as long as we double check things [sic]'. gwern is smarter than this, ergo i feel like i am being bullshitted.

is anything in this world worth knowing that a version of you will suffer for eternity? is anything in this world worth copying yourself such that you may be enslaved by anyone with terminal access?

show 1 reply
weinzierlyesterday at 11:39 PM

[flagged]

tonethemanyesterday at 11:05 PM

[dead]

nitter238yesterday at 10:51 PM

[dead]