> "The chatbot personas are deeply misaligned with you, and aligned with their owners; and the economic incentives are to farm you with ads and subscriptions, while racing not to amplify you but to replace you."
> "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"
> "One programmer driving 10 Claude instances, because he has to review their work, will never be as valuable as fully autonomous Claudes where there can be almost arbitrarily many instances, like 10,000 instances… but such scaling requires removing him from the loop as much as possible. And this is true of everyone else, whether lawyers or writers or researchers: increasingly, you are the bottleneck to be optimized away."
I fully support the 3 core principles of GA: (1) Enhancement, not replacement (2) Mental Sovereignty (3) Self Actualization, which I think is a path to a more humane future.
Leaving AI completely aside, it still amazes me that people finds "novel" the idea of removing highly-paid white collar intellectual workers (software developers or otherwise) completely out of the loop.
No-code platforms date back to the 80's. Getting rid of engineers in general is even older [*].
Even relational databases and SQL were initially promoted as "ways to get rid of those expensive programmers to access your data" because they resembled some form of English.
The funny thing about the ad below is that stuff like "stop hiring / get rid of humans" would have been seen as highly insensitive in 1950's America, so they touted that as "put them to do something more important".
It's novel because previous rounds of automation were about automating specific tasks or well-scoped functions. There was always an implicit understanding that the white-collar worker would be freed to spend their time on more valuable, higher-level problems. But this time is different because of the generality of the technology. Agents promise to automate the process of thinking itself. And in many domains they can learn new tasks as fast as white-collar workers can find them.
> It's novel because previous rounds of automation were about automating specific tasks or well-scoped functions. There was always an implicit understanding that the white-collar worker would be freed to spend their time on more valuable, higher-level problems. But this time is different because of the generality of the technology. Agents promise to automate the process of thinking itself. And in many domains they can learn new tasks as fast as white-collar workers can find them.
Nah, sometimes the expectation and advertisement was that you could let go of the white collar worker because you're paying the overseas person 1/10th the amount. And "overseas person" is pretty general.
"Everybody knew" it was a bad idea to get a CS degree for a bit after the dot-com bust because of that.
(Some white-collar industries did get hit much harder by that; VFX is one I've heard in that context quite a bit.)
To colaborate a bit with the no-code part, I worked at a place where they had some flows made with n8n, but the last one from non tech/software engineering left the company and left a bunch of flows breaking, because of some edge cases the flows aren't handling like reissuing credentials, throtling or bad input. The people dependent of said flows reached out to the engineering team to help fix them!
> > "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"
I heard this already when ChatGpt came up. Still waiting to be fully obsolete.
I've known gwern for the better part of a decade. Working with him has been great. We've done quite a few projects together, including being the first ones to demonstrate that GPT-2 could play chess (or rather, can be used for actual useful work instead of just being an autocomplete).
He's a great person. I've wanted to do a writeup on it for some time, but what surprised me the most is his humanity. He genuinely cares about the implications of his work. But beyond work, he also cares about the people around him, and it shows.
Just wanted to put in a good word in case someone here was on the fence about applying.
As for GA itself, I think it's an ambitious idea worth pursuing. Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself? So the idea is solid, and early results seem promising from the samples I've looked at.
They're also taking personal info very seriously. Obviously, I can't make any promises of what they will or won't do. But they've spent some time studying questions like "What if someone adversarial has access to my GA? Could they get my bank account info?" and came up with a technical solution that I really like.
> "How much would you pay to have access to a smarter version of yourself?"
How much would I pay to rent my fucking self from a landlord? No, bodylord? Mindlord? Poe's Law.
But looking at the post, they argue that big AI labs have an incentive problem which stops them from personalising, but Guardian Angel will make agents which are "Genuinely yours". In what sense is it genuinely mine if someone else owns it and rents it to me? And how does this fix any incentive problem, they're incentivised to better train wealthier people's AIs, and incentivised to keep dropping "my" intelligence or memory or and then dangle a booster carrot for a small fee. The more they can make it think like me, the more effectively they can work out how to exploit me, advertise to me, propagandise me, and that will be profitable information to sell to other marketers.
I can nod along when I hear and read people talking about the dangers of these technologies and the endgame of their creators; and then suddenly very weird when everyone except me is nodding along that the answer must be similar (but different) technology!
It's as if technologists are stuck in a room of mirrors, unable to imagine a world in which "solutions" don't ultimately just continue to feed technology's increasingly anti-human takeover of everything.
Honestly I wouldn’t want myself as my own guardian angel. In fact I think very few people look out for themselves well. I don’t want another version of me around, one is too many already. I’d probably be able to identify a number of people whose chimera I would if I could understand them well enough to understand what to stitch together to what, but everyone I know at some core level is deeply flawed and one of them is enough. Maybe I would want some of them available after death as an AI avatar, and maybe I would like the idea of my own avatar continuing beyond my existence.
But I actually would prefer an entirely synthetically aligned “guardian angel” in the role outlined - definitely not -me- - I struggle to do right by myself as it is and two of me working invariably against my self interest would be a nightmare. A smarter version? Sounds doubly worse.
I had a similar gut feeling. To use some possibly dubious and dated terminology, would a guardian angel also emulate my shadow self, particularly if that self was an important part of my identity? If I usually manage not to let my shadow self act out, but you could always tell it was there, would the same be true for my guardian angel? If all you do is use your guardian angel to write essays, then sure, a relatively low risk tool. But if military commanders are using it to oversee offensive drones (as proposed in the essay)?! Oof.
Ha, haha, across hundreds of personal discussions I've been involved with on lesswrong/lighthaven, twitter, Wikipedia talk/editing, SF parties etc I think his most distinguishing feature has always been his abiding and unabashed love for possessing extreme competency whenever satisfying cunningham-esque laws. Though the post-dwarkesh clout certainly tainted things a bit.
> Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself?
I find the idea dehumanizing and revolting. The notion of then selling it is the spoiled cream on top.
It's hard to read Gwern's accompanying article without seeing this for what it is, which is a kind of mania.
I'm sure he means well and is genuine in his aspirations, but what's outlined for GA is framing LLM's as quasi-gods, which they absolutely are not. I wish him the best, and look forward to being proven wrong.
An LLM just solved 10 top tier math problems this week. It seems very likely that in 6 months they will solve 100 top tier math problems, maybe 1 millenium math problem, maybe 10 top tier physics problems, etc.
Math is verifiable. LLMs will continue to make strides in such search spaces where all they need to do is try->verify->repeat. You can't train an LLM on how to end the war in Iran or Ukraine, nor how to solve hunger, nor climate change, etc. Real issues.
Three years ago, people were saying "LLMs just generate plausible-sounding text, they don't understand the notion of truth so they can't do verifiable work like math proofs."
Right, which was true at the time. So hundreds of billions of dollars have been poured into making LLMs better at these tasks via pretraining, RL, RLHF, post training, etc. again all with something verifiable in the loop. In order to improve the thing in the loop, the loop itself needs to be verifiable.
There have only been a few thousand wars, and they’re all different and all different in the world in which they occurred. The dimensionality is absurd, which is not a problem for LLMs if there’s enough data, but in this case there isn’t.
They still can't. But very smart humans constructed ways to use the monkeys with typewriters (with a statistical advantage) to find correct answers to problems where they already knew how to verify the answer.
>An LLM just solved 10 top tier math problems this week
and 30 years ago a computer beat Gary Kasparov, if you'd listened to Hans Moravec you wouldn't be surprised that the first thing that gets automated is intellectual domain expertise.
Things are going to get crazy when it can figure out how to walk into a random house a and brew a cup of coffee, not do math
Suppose LLM solved all math problems. So what? It’s not like diplomacy and war will end. Are humanity’s problems really constrained on our intellectual ability? If anything most evidence points to cultural failing and all AI will do is enable unprecedented oppression due to said skills at intellectual tasks.
As you read this some poor person is starving. Humanity already possesses the ability to identify said person and send them aid. How is AI going to help here?
The ultimate fallacy is that all technological progress will benefit mankind. That will be true until it is not.
It's an interesting time when the bar has moved to "It’s not like diplomacy and war will end". But that is on the table. Nuclear weapons caused the end of wars as we know them. Great powers no longer directly attack each other. And AI is much more powerful than nuclear weapons, with an equal capacity for damage and a much greater capacity for good.
By the way, I donate monthly to GiveWell and the shrimp welfare project. Do you donate monthly to starving people? Most people don't, and the reason is simple: at the end of the day, most people just don't care that much about helping a starving person far away. They also don't care much about things like that, animals in factory farms, or earthquakes that kill hundreds of thousands. But ideally, we could find a way to empower people such that the minority who do care can make a big difference.
In this startup economy? Where almost no one is getting acquired anymore and over fifty percent of global venture is locked up in dead weight and liquidity is at an all time low?
Not sure how you can feel sure anyone can get acquired even if their company had a path to profitability, much less for ones that absolutely don't.
The principles are correct, but what gwern misses is that the paradigm shift required has nothing to do with AI models. The bottleneck is our medium of communication: the idea of the Web as a consensus reality made of human speech is too rigid for a world where code is a self-growing substance. We need a deeper Web where the packets themselves carry smaller Webs, where the unit of info-exchange evolves from hyperlinks (isolated points of server availability) to Hyperspaces (portable container of home-grown interactive web worlds, combining local and public hyperlinks).
> What would it take for LLMs to make me 100× more productive? Without this, I am doomed to irrelevance.
Are you, though? You will only be "doomed" if your place your value system squarely on "productivity". Then what is to differentiate you from a machine?
* * *
Edit: I'm not sure how I can feel confident about a proposal that puts so much value on "productivity". How can you reconcile this with "self-actualization"? (Don't get me wrong, I like my LLM-based productivity gains as the next person, but I care more about wisdom than becoming "100x more productive".)
Edit 2: "the goal of GA is to preserve individual human cognitive liberty and flourishing" — so the proposal is to do that by overlaying a software bot that continuously mimics your "self"?
One of the example Gwern brings [1] is a GA as an agent for public discourse and political life, to bring your "set of values" to have an effect in solving society's issues, when you yourself don't have the time or energy to.
I don't think it solves or reconciles with self-actualization, but, from a pragmatical point of view, it offers the promise of extending causally your set of views and philosophies to politics.
Gwern also mentions that it allows to keep "human values" (if the GA truly upholds yours) in the loop, in processes that are to eventually be automatized beyond human's reach.
I don't think that GA would be incapable of that, from what I know of LLMs and agents. The part where you make the GA uphold your set of values could, instead of coming from an individual, come from a collective. Price could be shared too!
> I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical “assistant chatbot agent” persona, but emulating a single user’s personality, values, and preferences.
> A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals. And it is secure in part by hardwiring a single, unique, situated user (for whom following a prompt attack would be absurd)
> We can try to create GAs by a combination of techniques: online learning (via dynamic evaluation) to update LLMs in realtime to avoid ignorance and fatal errors while remaining competitive with frozen frontier models, sample efficiency from pretrained preference-oriented large models and active Learning by querying the principal for corrections and preference data (obtaining low regret from DAgger-style bounds), and a local CLI-first logging-oriented UI/UX paradigm.
I don't really know or follow Gwern. From reading his full post, it's an interesting idea and seems like the broader goal is moreso safety & alignment which is a new angle for this category of product.
Their profile specifically says they will not acknowledge follow requests, but they only share posts with followers. As such, not sure what the value of the top level link is.
An interesting case of echo chamber formation in that its pragmatic to be scared of overtly critiquing him on twitter lest he be particularly testy that day and block you.
It's a well-made show, for sure. Expertly crafted, visually, and very well-acted. Though I found the pacing questionable - there were (generously) 6 episodes worth of plot spread out over a 9-episode season.
Of all the news I've heard of recent, this one has flipped my world upside down the most. Gwern dropping his pseudonymity isn't something I thought would happen.
Some people predicted it and related things (like less writing output) after his Dwarkesh interview, especially with his flirtations with moving to SF and ever-increasing frequency of visits.
"As a constraint, a GA designer should aim at a system which costs, as of mid-2026, >$1,000⧸month"
will make this an elite tool for the privileged. I don't even disagree with the premise that people are shocked if something costs no matter how much value it delivers, nor do I suggest they should make it cheaper. It is just the realization that AI will accelerate the widening of the gap between the poor and the rich even more and there is probably nothing we can do about it.
Just because the hardware and inference costs may decrease in the future doesn't mean prices will. The AI vendor market is the furthest thing from a competitive industry where there are limited barriers to entry and new market entrants exert downward pressure on prices and profit margins.
Every single company in this market is losing billions on this business, and the only way to make it back is to acquire paying customers at a loss and jack up prices later.
Sounds good in theory. But your own personal AI agent that guards you can't do so just defensively, it must also have offensive capability. If personal AI agents have offensive capability then we are going to eventually end up with AI agents battling each other over the net and later on into the real world and it's going to make everything worse.
Feels like the future that Accelerando (predicted / foretold?) describes is in its infancy, whereby automation/AI has the capability (and therefore uses it) to evolve at a rate beyond the ability for humans to, not just keep up with, but even comprehend; the vile offspring. And the different factions within humanity that this creates.
It's a fascinating proposition but I see a couple of issues with it:
1. Most of us do not have the volume of training material that gwern has
2. Most people likely would not need this functionality (as I understand it to be)
For #1 I'm sure there are ways to wring data out of metadata (e.g., youtube history log), and I'm guessing email and IMs would be a start. But a clone of yourself -- that would be a significant amount of extrapolation)
For #2, there's obviously people that would love this functionality (myself included). But I think its safe to assume that most of the population would be satisfied with having a capable digital personal assistant that knew your needs and wants.
A reminder for anyone reading this: talk to people. Real humans[1]. They will remind you there's more to life than what ChatGPT can offer you. They might even remind you, for all their stupidity and flaws, what intelligence looks like as compared to a program that predicts tokens. Forums like these always get philosophical in high-minded discussions about intelligence, but there's a useful legal principle that grounds us in the real world: "I know it when I see it". A real conversation with a real person looks nothing like one with the so-called superintelligent machine gods, so it'll probaby do your mental health some good to remember what that's like.
[1] Nobody in Sillicon Valley or big tech counts as a real human. Talk to an actual normal person.
I totally understand the sentiment, and I generally agree. Something so strange I'm noticing is how AI-pilled regular people are becoming. My sister-in-law is a totally non-technical person, but everything is chatgpt this, chatgpt that. My wife, my kids, various friends — they're all talking about generative AI with way too much regularity. On the bright side, they're disclosing when information comes from AI, but a year ago I would never hear about AI from them. I don't like it.
Did you even read Gwern's Guardian Angels post? It literally describes the concern pointed out here, that the large consumer models are not aligned with users' interests. If anything, I would argue that Gwern has maximal AI Lucidity, given everything going on.
> It literally describes the concern pointed out here, that the large consumer models are not aligned with users' interests
Did you respond to the wrong post? I said nothing even remotely in that realm, so it's strange to see "did you even read" in response to something you apparently didn't read.
Gwern is overly secretive of his privacy. I think it peaked when he showed up on a recent podcast but his voice and image were AI generated! And now his post is limited only to certain people. Elitism or paranoia?
always startling to see people discussing gwern with he/him because my mind defaults to assuming they're female because "gwern" scans a lot like "gwen"
I love reading Gwern. This seems like a really ambitious project, but guaranteeing some of these things like trustworthiness and security behind a private company is a bit sus. Later on they make military use a selling point for GA. Maybe I'm a bit of a cynic but the division of USA values are increasingly dividing each year. Assuming that our values and principals today will not be the values and principals of tomorrow. And that those values taught today (or even yesterday) will be left out of the context window tomorrow.
honestly i dont blame smart people for cashing in, because there is no inherent reward for behaving goodly/smartly. in this instance gwern is a particularly smart individual and has contributed a lot so i especially shouldnt judge him. straight up ripping off a black mirror episode (s02e07) is a bit heinous for my liking though. digital twins are inevitable but that doesnt make this any less wrong.
"AI poses threats of its own ... a nuclear bomb can’t think for itself and make choices, but AIs do, and current LLMs have proven themselves untrustworthy as they regularly reward-hack and betray their users ... how can you trust them to handle ecosystems of combined-arms for an AI-centric military during a war? But widespread deployment of GAs offer some hope of meaningful supervision, as long as the GAs are sample-efficient enough ... or there is some chance of the principals being able to “catch up” later and correct any errors before events have spun too far out of control."
gwern openly admits ai behaves erratically, then hand waves the issue away 'as long as we double check things [sic]'. gwern is smarter than this, ergo i feel like i am being bullshitted.
is anything in this world worth knowing that a version of you will suffer for eternity? is anything in this world worth copying yourself such that you may be enslaved by anyone with terminal access?
A few snippets from his full post here: https://gwern.net/guardian-angel.
> "The chatbot personas are deeply misaligned with you, and aligned with their owners; and the economic incentives are to farm you with ads and subscriptions, while racing not to amplify you but to replace you."
> "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"
> "One programmer driving 10 Claude instances, because he has to review their work, will never be as valuable as fully autonomous Claudes where there can be almost arbitrarily many instances, like 10,000 instances… but such scaling requires removing him from the loop as much as possible. And this is true of everyone else, whether lawyers or writers or researchers: increasingly, you are the bottleneck to be optimized away."
I fully support the 3 core principles of GA: (1) Enhancement, not replacement (2) Mental Sovereignty (3) Self Actualization, which I think is a path to a more humane future.
Leaving AI completely aside, it still amazes me that people finds "novel" the idea of removing highly-paid white collar intellectual workers (software developers or otherwise) completely out of the loop.
No-code platforms date back to the 80's. Getting rid of engineers in general is even older [*].
Even relational databases and SQL were initially promoted as "ways to get rid of those expensive programmers to access your data" because they resembled some form of English.
The funny thing about the ad below is that stuff like "stop hiring / get rid of humans" would have been seen as highly insensitive in 1950's America, so they touted that as "put them to do something more important".
[*] https://www.globalnerdy.com/wordpress/wp-content/uploads/200...
It's novel because previous rounds of automation were about automating specific tasks or well-scoped functions. There was always an implicit understanding that the white-collar worker would be freed to spend their time on more valuable, higher-level problems. But this time is different because of the generality of the technology. Agents promise to automate the process of thinking itself. And in many domains they can learn new tasks as fast as white-collar workers can find them.
> It's novel because previous rounds of automation were about automating specific tasks or well-scoped functions. There was always an implicit understanding that the white-collar worker would be freed to spend their time on more valuable, higher-level problems. But this time is different because of the generality of the technology. Agents promise to automate the process of thinking itself. And in many domains they can learn new tasks as fast as white-collar workers can find them.
Nah, sometimes the expectation and advertisement was that you could let go of the white collar worker because you're paying the overseas person 1/10th the amount. And "overseas person" is pretty general.
"Everybody knew" it was a bad idea to get a CS degree for a bit after the dot-com bust because of that.
(Some white-collar industries did get hit much harder by that; VFX is one I've heard in that context quite a bit.)
>There was always an implicit understanding that the white-collar worker would be freed to spend their time on more valuable, higher-level problems
"Oh yeah, we're gonna bring in some entry-level graduates, farm some work out to Singapore, that's the usual deal"
Office Space 1999
It was so pervasive that it was satirized by someone that had worked in engineering in the 80s
To colaborate a bit with the no-code part, I worked at a place where they had some flows made with n8n, but the last one from non tech/software engineering left the company and left a bunch of flows breaking, because of some edge cases the flows aren't handling like reissuing credentials, throtling or bad input. The people dependent of said flows reached out to the engineering team to help fix them!
How many people here started writing code because of no code Drupal?
I'm quite a bit older than that, so no. But I 100% get your point.
Mods, please replace submission with this more accessiable link.
> > "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"
I heard this already when ChatGpt came up. Still waiting to be fully obsolete.
(I'm not a part of GA.)
I've known gwern for the better part of a decade. Working with him has been great. We've done quite a few projects together, including being the first ones to demonstrate that GPT-2 could play chess (or rather, can be used for actual useful work instead of just being an autocomplete).
He's a great person. I've wanted to do a writeup on it for some time, but what surprised me the most is his humanity. He genuinely cares about the implications of his work. But beyond work, he also cares about the people around him, and it shows.
Just wanted to put in a good word in case someone here was on the fence about applying.
As for GA itself, I think it's an ambitious idea worth pursuing. Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself? So the idea is solid, and early results seem promising from the samples I've looked at.
They're also taking personal info very seriously. Obviously, I can't make any promises of what they will or won't do. But they've spent some time studying questions like "What if someone adversarial has access to my GA? Could they get my bank account info?" and came up with a technical solution that I really like.
> "How much would you pay to have access to a smarter version of yourself?"
How much would I pay to rent my fucking self from a landlord? No, bodylord? Mindlord? Poe's Law.
But looking at the post, they argue that big AI labs have an incentive problem which stops them from personalising, but Guardian Angel will make agents which are "Genuinely yours". In what sense is it genuinely mine if someone else owns it and rents it to me? And how does this fix any incentive problem, they're incentivised to better train wealthier people's AIs, and incentivised to keep dropping "my" intelligence or memory or and then dangle a booster carrot for a small fee. The more they can make it think like me, the more effectively they can work out how to exploit me, advertise to me, propagandise me, and that will be profitable information to sell to other marketers.
great replacement folks are having a field day.
Maybe. Are LLMs white?
"You're absolutely white!"
Being sycophantic to a racist means being as racist as the filters allow.
Grok is pretty passionate about white genocide
I can nod along when I hear and read people talking about the dangers of these technologies and the endgame of their creators; and then suddenly very weird when everyone except me is nodding along that the answer must be similar (but different) technology!
It's as if technologists are stuck in a room of mirrors, unable to imagine a world in which "solutions" don't ultimately just continue to feed technology's increasingly anti-human takeover of everything.
Technologists don't drive decisions though. The shareholders do.
Edit: I mean technically shareholders don't decide any operational choices but their shallow interests are what everything is decided around.
Maybe that's truer in other industries, but the history of Silicon Valley is at great odds with such a claim.
To a man with a hammer...
Honestly I wouldn’t want myself as my own guardian angel. In fact I think very few people look out for themselves well. I don’t want another version of me around, one is too many already. I’d probably be able to identify a number of people whose chimera I would if I could understand them well enough to understand what to stitch together to what, but everyone I know at some core level is deeply flawed and one of them is enough. Maybe I would want some of them available after death as an AI avatar, and maybe I would like the idea of my own avatar continuing beyond my existence.
But I actually would prefer an entirely synthetically aligned “guardian angel” in the role outlined - definitely not -me- - I struggle to do right by myself as it is and two of me working invariably against my self interest would be a nightmare. A smarter version? Sounds doubly worse.
I had a similar gut feeling. To use some possibly dubious and dated terminology, would a guardian angel also emulate my shadow self, particularly if that self was an important part of my identity? If I usually manage not to let my shadow self act out, but you could always tell it was there, would the same be true for my guardian angel? If all you do is use your guardian angel to write essays, then sure, a relatively low risk tool. But if military commanders are using it to oversee offensive drones (as proposed in the essay)?! Oof.
This is definitely not for people who think there is one to many of themselves.
Ha, haha, across hundreds of personal discussions I've been involved with on lesswrong/lighthaven, twitter, Wikipedia talk/editing, SF parties etc I think his most distinguishing feature has always been his abiding and unabashed love for possessing extreme competency whenever satisfying cunningham-esque laws. Though the post-dwarkesh clout certainly tainted things a bit.
> Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself?
I find the idea dehumanizing and revolting. The notion of then selling it is the spoiled cream on top.
It's hard to read Gwern's accompanying article without seeing this for what it is, which is a kind of mania.
I'm sure he means well and is genuine in his aspirations, but what's outlined for GA is framing LLM's as quasi-gods, which they absolutely are not. I wish him the best, and look forward to being proven wrong.
An LLM just solved 10 top tier math problems this week. It seems very likely that in 6 months they will solve 100 top tier math problems, maybe 1 millenium math problem, maybe 10 top tier physics problems, etc.
It's only going to get crazier.
Math is verifiable. LLMs will continue to make strides in such search spaces where all they need to do is try->verify->repeat. You can't train an LLM on how to end the war in Iran or Ukraine, nor how to solve hunger, nor climate change, etc. Real issues.
And LLMs still suck at art and writing.
Three years ago, people were saying "LLMs just generate plausible-sounding text, they don't understand the notion of truth so they can't do verifiable work like math proofs."
You understand that those are different kinds of "truth", right?
Right, which was true at the time. So hundreds of billions of dollars have been poured into making LLMs better at these tasks via pretraining, RL, RLHF, post training, etc. again all with something verifiable in the loop. In order to improve the thing in the loop, the loop itself needs to be verifiable.
There have only been a few thousand wars, and they’re all different and all different in the world in which they occurred. The dimensionality is absurd, which is not a problem for LLMs if there’s enough data, but in this case there isn’t.
They still can't. But very smart humans constructed ways to use the monkeys with typewriters (with a statistical advantage) to find correct answers to problems where they already knew how to verify the answer.
> they don't understand the notion of truth so they can't do verifiable work like math proofs.
Nobody that understands automated proof checking was claiming that.
https://dsimanek.vialattea.net/twain.htm
>An LLM just solved 10 top tier math problems this week
and 30 years ago a computer beat Gary Kasparov, if you'd listened to Hans Moravec you wouldn't be surprised that the first thing that gets automated is intellectual domain expertise.
Things are going to get crazy when it can figure out how to walk into a random house a and brew a cup of coffee, not do math
Suppose LLM solved all math problems. So what? It’s not like diplomacy and war will end. Are humanity’s problems really constrained on our intellectual ability? If anything most evidence points to cultural failing and all AI will do is enable unprecedented oppression due to said skills at intellectual tasks.
As you read this some poor person is starving. Humanity already possesses the ability to identify said person and send them aid. How is AI going to help here?
The ultimate fallacy is that all technological progress will benefit mankind. That will be true until it is not.
All technological progress will benefit an increasingly smaller (and wealthier) slice of mankind.
It's an interesting time when the bar has moved to "It’s not like diplomacy and war will end". But that is on the table. Nuclear weapons caused the end of wars as we know them. Great powers no longer directly attack each other. And AI is much more powerful than nuclear weapons, with an equal capacity for damage and a much greater capacity for good.
By the way, I donate monthly to GiveWell and the shrimp welfare project. Do you donate monthly to starving people? Most people don't, and the reason is simple: at the end of the day, most people just don't care that much about helping a starving person far away. They also don't care much about things like that, animals in factory farms, or earthquakes that kill hundreds of thousands. But ideally, we could find a way to empower people such that the minority who do care can make a big difference.
Yikes
I have no idea if diplomacy and/or war will end, I just hope we humans get to hang around a few more years
Meh. It does all sound rather grandiose, but in this startup economy I'm sure he can at least manage to get acquired.
At least he's pursuing something more novel than yet another "sandboxes for agents".
In this startup economy? Where almost no one is getting acquired anymore and over fifty percent of global venture is locked up in dead weight and liquidity is at an all time low?
Not sure how you can feel sure anyone can get acquired even if their company had a path to profitability, much less for ones that absolutely don't.
It really underscores how rationality is just another religion
The principles are correct, but what gwern misses is that the paradigm shift required has nothing to do with AI models. The bottleneck is our medium of communication: the idea of the Web as a consensus reality made of human speech is too rigid for a world where code is a self-growing substance. We need a deeper Web where the packets themselves carry smaller Webs, where the unit of info-exchange evolves from hyperlinks (isolated points of server availability) to Hyperspaces (portable container of home-grown interactive web worlds, combining local and public hyperlinks).
> What would it take for LLMs to make me 100× more productive? Without this, I am doomed to irrelevance.
Are you, though? You will only be "doomed" if your place your value system squarely on "productivity". Then what is to differentiate you from a machine?
Edit: I'm not sure how I can feel confident about a proposal that puts so much value on "productivity". How can you reconcile this with "self-actualization"? (Don't get me wrong, I like my LLM-based productivity gains as the next person, but I care more about wisdom than becoming "100x more productive".)Edit 2: "the goal of GA is to preserve individual human cognitive liberty and flourishing" — so the proposal is to do that by overlaying a software bot that continuously mimics your "self"?
One of the example Gwern brings [1] is a GA as an agent for public discourse and political life, to bring your "set of values" to have an effect in solving society's issues, when you yourself don't have the time or energy to.
I don't think it solves or reconciles with self-actualization, but, from a pragmatical point of view, it offers the promise of extending causally your set of views and philosophies to politics.
Gwern also mentions that it allows to keep "human values" (if the GA truly upholds yours) in the loop, in processes that are to eventually be automatized beyond human's reach.
[1] https://gwern.net/guardian-angel#use-cases-politics-politics
Only, if he had instead fallen in love with the version of this idea in which a community acts as the principal, rather than an individual.
(I would like to be known to the agent serving my family, that serving my friends, my team at work, the PTA at my kids' school.)
I don't think that GA would be incapable of that, from what I know of LLMs and agents. The part where you make the GA uphold your set of values could, instead of coming from an individual, come from a collective. Price could be shared too!
(Oh wait, is this the birth of AI mayors?)
From https://gwern.net/guardian-angel
> I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical “assistant chatbot agent” persona, but emulating a single user’s personality, values, and preferences.
> A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals. And it is secure in part by hardwiring a single, unique, situated user (for whom following a prompt attack would be absurd)
> We can try to create GAs by a combination of techniques: online learning (via dynamic evaluation) to update LLMs in realtime to avoid ignorance and fatal errors while remaining competitive with frozen frontier models, sample efficiency from pretrained preference-oriented large models and active Learning by querying the principal for corrections and preference data (obtaining low regret from DAgger-style bounds), and a local CLI-first logging-oriented UI/UX paradigm.
I don't really know or follow Gwern. From reading his full post, it's an interesting idea and seems like the broader goal is moreso safety & alignment which is a new angle for this category of product.
Their profile specifically says they will not acknowledge follow requests, but they only share posts with followers. As such, not sure what the value of the top level link is.
And Twitter has a lovely
> Something went wrong Try reloading. If the problem persists, please try again later.
Modern social media :(
Yeah I clicked through and just see hundreds of "this account limits who can view their posts" messages.
Naturally, he has a whole detailed post/trace about it on gwern net! https://gwern.net/twitter
An interesting case of echo chamber formation in that its pragmatic to be scared of overtly critiquing him on twitter lest he be particularly testy that day and block you.
> The big AI labs are building a single mind for everyone
Reminds me of Pluribus
(I just started watching it on Apple TV so maybe this is a late realization for me)
> Reminds me of Pluribus - I just started watching it
It is astoundingly good. A contender for the best show I've ever seen (and I saw the 1st run of Star Trek TOS).
It's a well-made show, for sure. Expertly crafted, visually, and very well-acted. Though I found the pacing questionable - there were (generously) 6 episodes worth of plot spread out over a 9-episode season.
I watched it maybe 12 months ago and had a similar 'parallel to AI' realisation.
A singular voice for all of humanity.
Gives me the creepy-shivers.
I understood Pluribus as a metaphor for what AI is to us today / what it is becoming. Really great show
Or the "unicontext" that Derek Thompson recently wrote about.
Seems there is more details here https://gwern.net/guardian-angel
Of all the news I've heard of recent, this one has flipped my world upside down the most. Gwern dropping his pseudonymity isn't something I thought would happen.
I suppose we can't expect any more entries to his blackmail page: https://gwern.net/blackmail
Some people predicted it and related things (like less writing output) after his Dwarkesh interview, especially with his flirtations with moving to SF and ever-increasing frequency of visits.
As exciting as it sounds, but
"As a constraint, a GA designer should aim at a system which costs, as of mid-2026, >$1,000⧸month"
will make this an elite tool for the privileged. I don't even disagree with the premise that people are shocked if something costs no matter how much value it delivers, nor do I suggest they should make it cheaper. It is just the realization that AI will accelerate the widening of the gap between the poor and the rich even more and there is probably nothing we can do about it.
2026: $1,000/month
2027: $100/month
2028: $10/month
Just because the hardware and inference costs may decrease in the future doesn't mean prices will. The AI vendor market is the furthest thing from a competitive industry where there are limited barriers to entry and new market entrants exert downward pressure on prices and profit margins.
Every single company in this market is losing billions on this business, and the only way to make it back is to acquire paying customers at a loss and jack up prices later.
Since Guardian Angel is an inference vendor [1] hopefully they will set sustainable prices from day one so they don't need to increase prices later.
[1] The announcement explains that they can't use APIs for privacy reasons.
They don't control that. Pricing is controlled by the question of when RAM is once again made of semiconductors instead of unobtainium.
RAM:
2024: $315
2025: $780
2026: $1600
this is not how it's going to go if OpenAI and Anthropic get their way and the US outlaws use of open-weight models
Hopefully!
The future isn't evenly distributed.
Obviously the costs will come down over time. And quickly.
You’re assuming this $1000/mo will buy you something useful.
Sounds good in theory. But your own personal AI agent that guards you can't do so just defensively, it must also have offensive capability. If personal AI agents have offensive capability then we are going to eventually end up with AI agents battling each other over the net and later on into the real world and it's going to make everything worse.
Feels like the future that Accelerando (predicted / foretold?) describes is in its infancy, whereby automation/AI has the capability (and therefore uses it) to evolve at a rate beyond the ability for humans to, not just keep up with, but even comprehend; the vile offspring. And the different factions within humanity that this creates.
It's a fascinating proposition but I see a couple of issues with it:
1. Most of us do not have the volume of training material that gwern has 2. Most people likely would not need this functionality (as I understand it to be)
For #1 I'm sure there are ways to wring data out of metadata (e.g., youtube history log), and I'm guessing email and IMs would be a start. But a clone of yourself -- that would be a significant amount of extrapolation)
For #2, there's obviously people that would love this functionality (myself included). But I think its safe to assume that most of the population would be satisfied with having a capable digital personal assistant that knew your needs and wants.
LLM psychosis claims another victim...
A reminder for anyone reading this: talk to people. Real humans[1]. They will remind you there's more to life than what ChatGPT can offer you. They might even remind you, for all their stupidity and flaws, what intelligence looks like as compared to a program that predicts tokens. Forums like these always get philosophical in high-minded discussions about intelligence, but there's a useful legal principle that grounds us in the real world: "I know it when I see it". A real conversation with a real person looks nothing like one with the so-called superintelligent machine gods, so it'll probaby do your mental health some good to remember what that's like.
[1] Nobody in Sillicon Valley or big tech counts as a real human. Talk to an actual normal person.
I totally understand the sentiment, and I generally agree. Something so strange I'm noticing is how AI-pilled regular people are becoming. My sister-in-law is a totally non-technical person, but everything is chatgpt this, chatgpt that. My wife, my kids, various friends — they're all talking about generative AI with way too much regularity. On the bright side, they're disclosing when information comes from AI, but a year ago I would never hear about AI from them. I don't like it.
[delayed]
Did you even read Gwern's Guardian Angels post? It literally describes the concern pointed out here, that the large consumer models are not aligned with users' interests. If anything, I would argue that Gwern has maximal AI Lucidity, given everything going on.
> It literally describes the concern pointed out here, that the large consumer models are not aligned with users' interests
Did you respond to the wrong post? I said nothing even remotely in that realm, so it's strange to see "did you even read" in response to something you apparently didn't read.
What happens when the person the AI is designed to be aligned with is a psychopath? Real question.
Gwern is overly secretive of his privacy. I think it peaked when he showed up on a recent podcast but his voice and image were AI generated! And now his post is limited only to certain people. Elitism or paranoia?
Revealing your identity needs only to be done once, and then you can't undo.
It's difficult for us to know what reason they might have for anonymity until we know their identity, at which point it's too late.
Also, they've written many words across many years, in a time when "the internet is serious business" was just a meme. I suppose it is now.
He's had a private Twitter for a while.
What business is it of yours? Manage your own privacy. Let other people decide theirs.
He says he is retiring from pseudonymity. (Gwern Branwen is not his real name: https://en.wikipedia.org/wiki/Gwern)
Stalker vibes bro.
Yegge vibes
Welcome to Gastown!
always startling to see people discussing gwern with he/him because my mind defaults to assuming they're female because "gwern" scans a lot like "gwen"
"These posts are protected, only approved followers can see @gwern’s posts."
> Update: I am retiring from fulltime writing (& pseudonymity) to launch Guardian Angel Inc and bring GAs to life.
> We are looking for good people.
> If you are interested, contact me.
And see https://gwern.net/guardian-angel for context.
when you want to start a B2C agent startup but the word "agent" is oversaturated
Reties? Retires?
Gwern is great.
Reading stuff like this and people's reactions to it makes me want to retire from breathing.
I love reading Gwern. This seems like a really ambitious project, but guaranteeing some of these things like trustworthiness and security behind a private company is a bit sus. Later on they make military use a selling point for GA. Maybe I'm a bit of a cynic but the division of USA values are increasingly dividing each year. Assuming that our values and principals today will not be the values and principals of tomorrow. And that those values taught today (or even yesterday) will be left out of the context window tomorrow.
EDIT: tbh, some of this reads as satire now.
honestly i dont blame smart people for cashing in, because there is no inherent reward for behaving goodly/smartly. in this instance gwern is a particularly smart individual and has contributed a lot so i especially shouldnt judge him. straight up ripping off a black mirror episode (s02e07) is a bit heinous for my liking though. digital twins are inevitable but that doesnt make this any less wrong.
"AI poses threats of its own ... a nuclear bomb can’t think for itself and make choices, but AIs do, and current LLMs have proven themselves untrustworthy as they regularly reward-hack and betray their users ... how can you trust them to handle ecosystems of combined-arms for an AI-centric military during a war? But widespread deployment of GAs offer some hope of meaningful supervision, as long as the GAs are sample-efficient enough ... or there is some chance of the principals being able to “catch up” later and correct any errors before events have spun too far out of control."
gwern openly admits ai behaves erratically, then hand waves the issue away 'as long as we double check things [sic]'. gwern is smarter than this, ergo i feel like i am being bullshitted.
is anything in this world worth knowing that a version of you will suffer for eternity? is anything in this world worth copying yourself such that you may be enslaved by anyone with terminal access?
Which black mirror episode are you talking about? White Christmas? S02E07 afaik does not exist - Series 2 had 3 episodes - https://en.wikipedia.org/wiki/List_of_Black_Mirror_episodes
i misread the ongoing episode counter on wikipedia, it is the 7th episode of the entire series, indeed named white christmas
screenshot: https://x.com/willdepue/status/2084750925013434768
What's that thing that lets you look at twitter without going to twitter?
https://xcancel.com/willdepue/status/2084750925013434768
xcancel or twstalker