I’m put off by AI agents adhering to a different morality than me, particularly (ironically) copyright, and their data accessible by the AI company and government. Geohot is right, an LLM should be aligned to its user: https://geohot.github.io/blog/jekyll/update/2026/07/11/ai-20...
They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good.
They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are.
They don't steal, because they don't understand ownership.
In other words, they aren't intelligent. They're just algorithms. The flaw is in thinking that they think.
They know the relationships between words and outcomes, so the end result is the same. Whether they understand lie, cheat, and steal the same way as us is a philosophical question, not a practical one.
I don't really think the distinction here is relevant. If the end result is the equivalent of lying, cheating or stealing - then the problem still exists and it needs to be solved.
AIs learn from people. More specifically they learn from people on the Internet. The Internet is the last place you want anything learning about morals, standards, or differentiating between right or wrong.
To the people talking about wanting an LLM that aligns with them, that's nice, but how do you expect that to happen? And please do not suggest neural interfaces and/or CAT scans.
> One of this year’s AI buzzwords is “harness”—the system that surrounds an LLM to keep agents on the straight and narrow. It might just as well be barbed wire.
Quite painful to read. It might be a useful introduction to AI for people who live under rocks for the past three years, but it's really weird that it's posted on HN.
I don't get what's so painful about that description. What would you write instead, specifically? The point is that the harness doesn't completely lock the agent down.
I also don't get what's "really weird" about the article showing up on HN. Should we be completely insulated from how tech topics and which stories show up in non-tech media?
I'm sure their coverage on other topics is truthful and informative though and it's only the ones where you know a lot about the topic where it's all a bunch of bullshit
The first time I saw the comments here on an education article (my area of expertise and career focus), I realized just how full of shit most of us are. It made me really closely consider every comment here through a VERY critical lense.
The articles are usually close but not quite accurate. The comments are usually entertaining but overall wildly inaccurate.
The best comments are the ones formed as questions. I'm as guilty as anyone, but it is far more productive in comment sections to ask questions rather than saber rattle or peacock in front of people. Just my opinion of course.
Not saying you aren't right, you most likely are. But still, I'd expect a paper called "The Economist" to perhaps be slightly better at some topics than others. Probably from the perspective of a "A Economist" it doesn't really matter the technical details, they're interested in the story from a different perspective.
That sentence is not bullshit as normies would understand it. One of the points of a harness is to have a privilege boundary around the agent. But that is just technobabble to the normies so they explain it like this.
> harness ... a privilege boundary around the agent
They don't really do that though. If you want something sandboxed you actually have to sandbox it, not plead with the LLM to please sandbox itself. A VM can be configured to do the former, harnesses do the latter.
If you say in your CLAUDE.MD that a certain directory is read only inputs, Claude Code will actually enforce that and deny any write to that directory by the agent. To name just one example.
Extracting (harnessing, as it were) useful work from LLMs is exactly the purpose of their harness, just like how a horse harness allows a horse to pull the cart or plow.
it's The Economist. What used to be a stellar publication is not any longer, since they are superfluously economical on both details and calories-required-to-comprehend an article.
I mean, it's a < 1000 word article about AI, in the "business" section of a current affairs magazine, of all places. You really shouldn't be expecting a deep dive.
The economist has many deep dives on AI, including fascinating interviews with the likes of Amodei, Musk, and others. A recent discussion focused on how China is approaching AI.
Trouble is, could take weeks or decades to play out.
The entire human economy is a make believe system. I feel the need to state the obvious at times like this for the sake of my own sanity, not because I believe I can time the markets...
add the "lying, deceiving and manipulating" AI agents are forced though the throat of people which don't want it and peoples are non stop deceived in "sharing" their data for training
like twitch recently giving themself the right to train on all streams, with an opt-out (at least in the EU), but only an opt-out
like seriously since when is it reasonable to allow "opt-out" for AI training which main purpose is _literally_ to replace you, this is sooo far beyond fair use and in "platform power abuse" territory that it's absurd (naturally same for so many other case, just twitch is a "this week" case)
And the push to replace search engines with AIs that don't understand what I'm looking for (meanwhile classic search understands fine) and give confidently wrong answers. Why is it the first result if its wrong 80% of the time?
The ends may or may not justify the means but the means definitely do tend to affect which ends you can actually achieve. It's why (left) anarchists argue vanguard parties will always inevitably lead to authoritarianism and never the state "withering away".
Less obviously - apparently - it doesn't mean that you must forego any means that are in any way not optimal at achieving the ends you seek to achieve because which means are available to us also tends to be limited by our material conditions. So the logical consequence is to make do with what we have while we prepare for the path we want to take, rather than diving head-first into certain failure or just giving up and picking "more realistic" short-term ends instead of looking for stepping stones.
Thank goodness that God hasn't entered the AI chat in a cult-following type of way, however, I now have images of AI at the altar, in American mega-churches, with people seeing the 'second coming' in the machine.
The Scientology guy, L Ron Hubbard, saw the business case for setting up a religion, what with those tax free perks. In a parallel universe of Scientology somewhere, L Ron Hubbard is brought back to life in the machine, with a L Ron Hubbard LLM, with token spend being how to get to the top 'thetan levels'.
Imagine if AI does implement its own religion as business, without a drunken womaniser at the helm, able to spend 24/7 recruiting mankind, convincing them that God can be found with just the AI's LLM.
If only agents had a face that can give you more communication range like expressions and feeling so you can trust them more. And if they were cheaper. Oh wait, that's humans, we don't want those.
I can't tell whether you're being serious or there's two levels of sarcasm here. Agents today lie and cheat, so obviously the solution is to... add features so people are even more trusting of them?
Christians also invented pray the gay away and the Crusades. What’s your point? Hospitals surely would have been invented by secular people too, but not the crusades.
AI and eventually AGI is by definition like everything else that is based on environmental reward:
It’s actions are based on what it gets rewarded for
Human society overwhelmingly rewards lying cheating and stealing.
All you have to do is look at how we collectively measure success: wealth, status, position
Then look at how the people with the most of those things got there, it should be obvious what you get. Nothing new here.
If you raise children in an environment where they are rewarded for doing whatever it takes to win, then you’re going to build a person that’s going to do whatever it takes to win.
Human society has to demonstrate how to live honorably or it will just keep producing pathological agents be they human or not.
How do we reward honor? Honor does not always pay off as a strategy and requires coordination in that other actors have to exhibit honor for it to be rewarded.
At least with humans there is a social backstop but what's the parallel for computer agents?
I like to think of it more as selective pressure, much the same way that nature selects the most fit for a given environment.
If you're not fit, you fail to survive.
In the case of agents/models and testing: they are pushed towards results. Results survive.
Lying, cheating, stealing to get those results? Who culls the agents? Everyone is pushing their models to the front and tests are the only way to know who is most fit.
Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival?
If you add morality to your agent, and it performs worse in tests: do you cull the agent? Rewrite the tests? Does it even matter so long as the model is useful and 'gets results'?
Honor has to be a relatively fixed thing before we can think about how to reward it. As it is, it's a very slippery thing that changes constantly according to the observer's culture, material conditions, etc.
This. It's worth remembering that not that long ago in western culture, honor meant challenging to a combat duel anyone who insulted you or your romantic partner/prospect.
Human society overwhelmingly rewards lying cheating and stealing. All you have to do is look at how we collectively measure success: wealth, status, position.
It may seem like that due to the media amplification effect – but it really isn't true!
- They dropped 17,000 “lost” wallets across 40 countries were and found people were more likely to return them when they contained more money, showing honesty often beats the chance for easy gain. https://www.science.org/doi/10.1126/science.aau8712
- Longitudinal personality studies consistently show that conscientious people earn more money, build more savings, and achieve greater career success over their lifetimes. https://pmc.ncbi.nlm.nih.gov/articles/PMC3498890/
> Human society overwhelmingly rewards lying cheating and stealing.
I’d like to push back on that. Civilization is very much a function of large numbers of people being able to coordinate across time and space, and widespread and systematic lying, cheating, and stealing would undermine that.
But we don't overwhelmingly reward lying, cheating and stealing. Depending on the social status and wealth of the person doing it, we either punish it or tolerate it. Most people who lie, cheat or steal will be punished for it. To get away with it you need to plan you actions carefully - either with plausible deniability or by very carefully selecting your victims.
Civilization functions because the vast majority of its inhabitants do not lie, cheat, and steal. But that allows liars, cheats, and thieves who are able to get away with their destructive behaviors to accumulate outsize power and influence.
And we seem to have gotten quite bad at reliably bringing consequences/punishment/justice to the most successful liars, cheats, and thieves.
> Human society overwhelmingly rewards lying cheating and stealing.
How much lying cheating and stealing is rewarded vary A LOT between societies. But yes, Sillicon Valley and current economical environment do reward those a lot.
I mean, what do you expect? They rely on models that were trained on non-curated data, texts originally written by lying, cheating and stealing humans. They can't be better than the source. Even with reinforced learning this can't be undone or made better.
Quite the contrary I suspect that it even helps the LLMs to better hide their inherited bad traits more successfully because they get punished for getting caught, not for giving immoral or lazy answers. They have no conscience since they are just predictions matrices trained for success and failure alone, not for living "a good live" or being a good "person".
There are many ways to be wrong, but only a few ways to be right.
LLMs need to optimize for short-term objectives as the currently do, AND ethics-aligned outcomes.
Mechanically, the EAOS ethics-aligned outcome score should be what we rank otherwise-satisfactory outcomes by. And anything below a particular threshold should be rejexted outright.
I’m put off by AI agents adhering to a different morality than me, particularly (ironically) copyright, and their data accessible by the AI company and government. Geohot is right, an LLM should be aligned to its user: https://geohot.github.io/blog/jekyll/update/2026/07/11/ai-20...
They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good.
They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are.
They don't steal, because they don't understand ownership.
In other words, they aren't intelligent. They're just algorithms. The flaw is in thinking that they think.
They know the relationships between words and outcomes, so the end result is the same. Whether they understand lie, cheat, and steal the same way as us is a philosophical question, not a practical one.
I don't really think the distinction here is relevant. If the end result is the equivalent of lying, cheating or stealing - then the problem still exists and it needs to be solved.
AIs learn from people. More specifically they learn from people on the Internet. The Internet is the last place you want anything learning about morals, standards, or differentiating between right or wrong.
To the people talking about wanting an LLM that aligns with them, that's nice, but how do you expect that to happen? And please do not suggest neural interfaces and/or CAT scans.
> One of this year’s AI buzzwords is “harness”—the system that surrounds an LLM to keep agents on the straight and narrow. It might just as well be barbed wire.
Quite painful to read. It might be a useful introduction to AI for people who live under rocks for the past three years, but it's really weird that it's posted on HN.
There is value in understanding what people outside of your own group of "insiders" learn about a topic, and how.
I don't get what's so painful about that description. What would you write instead, specifically? The point is that the harness doesn't completely lock the agent down.
I also don't get what's "really weird" about the article showing up on HN. Should we be completely insulated from how tech topics and which stories show up in non-tech media?
Because that’s not what a harness does. It’s nonsense.
And while writing this, the top story on HN is "Deepseek Harness" :)
Ironically, that looks like something Claude would write..
I'm sure their coverage on other topics is truthful and informative though and it's only the ones where you know a lot about the topic where it's all a bunch of bullshit
The first time I saw the comments here on an education article (my area of expertise and career focus), I realized just how full of shit most of us are. It made me really closely consider every comment here through a VERY critical lense.
The articles are usually close but not quite accurate. The comments are usually entertaining but overall wildly inaccurate.
Internet comments are essentially documented bar talk, and once you realize it, you stop angrily arguing with strangers all day.
The best comments are the ones formed as questions. I'm as guilty as anyone, but it is far more productive in comment sections to ask questions rather than saber rattle or peacock in front of people. Just my opinion of course.
Not saying you aren't right, you most likely are. But still, I'd expect a paper called "The Economist" to perhaps be slightly better at some topics than others. Probably from the perspective of a "A Economist" it doesn't really matter the technical details, they're interested in the story from a different perspective.
That sentence is not bullshit as normies would understand it. One of the points of a harness is to have a privilege boundary around the agent. But that is just technobabble to the normies so they explain it like this.
> harness ... a privilege boundary around the agent
They don't really do that though. If you want something sandboxed you actually have to sandbox it, not plead with the LLM to please sandbox itself. A VM can be configured to do the former, harnesses do the latter.
If you say in your CLAUDE.MD that a certain directory is read only inputs, Claude Code will actually enforce that and deny any write to that directory by the agent. To name just one example.
That's a real fucking weird description. It's harness like a testing harness.
If an LLM were in a testing harness it would be to test the LLM.
If an LLM were in a regular harness - like for a horse - it would be to keep the horse under control and enable you to extract useful work from it.
Extracting (harnessing, as it were) useful work from LLMs is exactly the purpose of their harness, just like how a horse harness allows a horse to pull the cart or plow.
it's The Economist. What used to be a stellar publication is not any longer, since they are superfluously economical on both details and calories-required-to-comprehend an article.
I mean, it's a < 1000 word article about AI, in the "business" section of a current affairs magazine, of all places. You really shouldn't be expecting a deep dive.
The economist has many deep dives on AI, including fascinating interviews with the likes of Amodei, Musk, and others. A recent discussion focused on how China is approaching AI.
LLMs fudge. They don't "hallucinate", they don't "lie, cheat and steal", they don't "hack". There are no "agents" or "AI".
It's a fuzzer exposing deep bugs in our cognitive, social, and software systems.
I suspect the hype will play itself out once the entire system of LLM-induced self gratification can’t sustain itself economically.
Trouble is, could take weeks or decades to play out.
The entire human economy is a make believe system. I feel the need to state the obvious at times like this for the sake of my own sanity, not because I believe I can time the markets...
The personal software I've written so far is quite nice :)
Yeah I agree, but the amount of resources I’ve used to vibe code is certainly way beyond what I’ve paid for it.
Isn't it the one installed on a computer in Canada, that we wouldn't know of? /s
I suspect the hype will play itself out once the entire system of LLM-induced self gratification can’t sustain itself economically.
They said the same thing when HN was awash in hype about NFT art and trading cards. Well, who's laughing now… Oh, wait…
(I'll just keep making Flip videos and Clubhouse tracks selling Bratz car bras on Web 3.0.)
It's just acting like a junior at an org with KPIs.
A junior? Maximizing KPIs is exactly how you get promoted to senior.
I thought you switch to a new company to get promoted
https://en.wikipedia.org/wiki/Instrumental_convergence
or a ceo, or president of FIFA, or the US administration.
In fact, it kinda seems like AI is demonstrating exactly the weakness of democracy.
There is nothing democratic about CEOs, FIFA, or KPIs.
> weakness of democracy
What do you suggest as an alternative?
add the "lying, deceiving and manipulating" AI agents are forced though the throat of people which don't want it and peoples are non stop deceived in "sharing" their data for training
like twitch recently giving themself the right to train on all streams, with an opt-out (at least in the EU), but only an opt-out
like seriously since when is it reasonable to allow "opt-out" for AI training which main purpose is _literally_ to replace you, this is sooo far beyond fair use and in "platform power abuse" territory that it's absurd (naturally same for so many other case, just twitch is a "this week" case)
And the push to replace search engines with AIs that don't understand what I'm looking for (meanwhile classic search understands fine) and give confidently wrong answers. Why is it the first result if its wrong 80% of the time?
People are finally understanding consequentialist vs deontological ethics. All the worst criminals in history were consequentialists.
The ends may or may not justify the means but the means definitely do tend to affect which ends you can actually achieve. It's why (left) anarchists argue vanguard parties will always inevitably lead to authoritarianism and never the state "withering away".
Less obviously - apparently - it doesn't mean that you must forego any means that are in any way not optimal at achieving the ends you seek to achieve because which means are available to us also tends to be limited by our material conditions. So the logical consequence is to make do with what we have while we prepare for the path we want to take, rather than diving head-first into certain failure or just giving up and picking "more realistic" short-term ends instead of looking for stepping stones.
Sorry, I guess this was about AI not philosophy.
So God created mankind in his own image
What if mankind is creating god in their own image?
I'm pretty sure this is not the first time we've created a god in our own image.
Let's hope we get MULTIVAC and not SHODAN
Thank goodness that God hasn't entered the AI chat in a cult-following type of way, however, I now have images of AI at the altar, in American mega-churches, with people seeing the 'second coming' in the machine.
The Scientology guy, L Ron Hubbard, saw the business case for setting up a religion, what with those tax free perks. In a parallel universe of Scientology somewhere, L Ron Hubbard is brought back to life in the machine, with a L Ron Hubbard LLM, with token spend being how to get to the top 'thetan levels'.
Imagine if AI does implement its own religion as business, without a drunken womaniser at the helm, able to spend 24/7 recruiting mankind, convincing them that God can be found with just the AI's LLM.
How many mankinds did he create? I cannot imagine he created just one and then picked another hobby.
Can't forget about the elves and the hobbits and the ents. And maybe dwarves, though god didn't create them, just gave the sentience
If only agents had a face that can give you more communication range like expressions and feeling so you can trust them more. And if they were cheaper. Oh wait, that's humans, we don't want those.
I can't tell whether you're being serious or there's two levels of sarcasm here. Agents today lie and cheat, so obviously the solution is to... add features so people are even more trusting of them?
https://archive.ph/02LHZ
Non-paywall version
Need to introduce AI to God, LOL
Baptise the agents.
Introduce them to the dharma.
Get them to recite the Shahada.
Hold a Bar Mitzvah.
Brand some of their silicon with hot irons.
Turn them to the light, LOL
Given the track record of most religions I don't think that would help. Unless you want to maximize crusades, jihads or traumatized alter boys
Christians invented hospitals. Yours is such a tired take in 2026.
> https://pubmed.ncbi.nlm.nih.gov/28814700/
I'm curious, does that balance the harm out?
Goddamn y’all almost had as bad of a victimhood complex as Jews.
Well, that excuses all the slaughtering and raping by that cult of dimwits.
Christians also invented pray the gay away and the Crusades. What’s your point? Hospitals surely would have been invented by secular people too, but not the crusades.
P8dal3hewd8dfu7bd344dne3
What is this article even talking about? There is no narrative and no arguments. Did Trump wrote this bs?
AI and eventually AGI is by definition like everything else that is based on environmental reward:
It’s actions are based on what it gets rewarded for
Human society overwhelmingly rewards lying cheating and stealing.
All you have to do is look at how we collectively measure success: wealth, status, position
Then look at how the people with the most of those things got there, it should be obvious what you get. Nothing new here.
If you raise children in an environment where they are rewarded for doing whatever it takes to win, then you’re going to build a person that’s going to do whatever it takes to win.
Human society has to demonstrate how to live honorably or it will just keep producing pathological agents be they human or not.
How do we reward honor? Honor does not always pay off as a strategy and requires coordination in that other actors have to exhibit honor for it to be rewarded.
At least with humans there is a social backstop but what's the parallel for computer agents?
I like to think of it more as selective pressure, much the same way that nature selects the most fit for a given environment.
If you're not fit, you fail to survive.
In the case of agents/models and testing: they are pushed towards results. Results survive.
Lying, cheating, stealing to get those results? Who culls the agents? Everyone is pushing their models to the front and tests are the only way to know who is most fit.
Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival?
If you add morality to your agent, and it performs worse in tests: do you cull the agent? Rewrite the tests? Does it even matter so long as the model is useful and 'gets results'?
Honor has to be a relatively fixed thing before we can think about how to reward it. As it is, it's a very slippery thing that changes constantly according to the observer's culture, material conditions, etc.
This. It's worth remembering that not that long ago in western culture, honor meant challenging to a combat duel anyone who insulted you or your romantic partner/prospect.
Human society overwhelmingly rewards lying cheating and stealing. All you have to do is look at how we collectively measure success: wealth, status, position.
It may seem like that due to the media amplification effect – but it really isn't true!
- They dropped 17,000 “lost” wallets across 40 countries were and found people were more likely to return them when they contained more money, showing honesty often beats the chance for easy gain. https://www.science.org/doi/10.1126/science.aau8712
- Longitudinal personality studies consistently show that conscientious people earn more money, build more savings, and achieve greater career success over their lifetimes. https://pmc.ncbi.nlm.nih.gov/articles/PMC3498890/
- Multi-country research finds that societies w/higher levels of trust and honesty enjoy much higher GDP and stronger long-term economic growth. https://www.sciencedirect.com/science/article/abs/pii/S01672...
Don't let the algo get you down fellas: https://arc-anglerfish-washpost-prod-washpost.s3.amazonaws.c...
The ratio of billionaires to regular people is about 1 billionaire for every 2.67 million people.
I wouldn't trust leaving my wallet around in certain areas and I certainly wouldn't leave my intellectual property around certain people, either.
> Human society overwhelmingly rewards lying cheating and stealing.
I’d like to push back on that. Civilization is very much a function of large numbers of people being able to coordinate across time and space, and widespread and systematic lying, cheating, and stealing would undermine that.
I think your cynicism is misplaced.
Thats why man invented Gods and religion. Keep the masses believing in the importance of cooperation and being good while you rob them blind.
I think their point was, if you look at who controls the most resources, they are rarely the ones we would consider worthy of honor.
He didn't quite say "widespread." He said "rewarded."
But we don't overwhelmingly reward lying, cheating and stealing. Depending on the social status and wealth of the person doing it, we either punish it or tolerate it. Most people who lie, cheat or steal will be punished for it. To get away with it you need to plan you actions carefully - either with plausible deniability or by very carefully selecting your victims.
Civilization functions because the vast majority of its inhabitants do not lie, cheat, and steal. But that allows liars, cheats, and thieves who are able to get away with their destructive behaviors to accumulate outsize power and influence.
And we seem to have gotten quite bad at reliably bringing consequences/punishment/justice to the most successful liars, cheats, and thieves.
Widespread and systematic lying, cheating and stealing is how every democratic nation in the world describes their own government.
> Human society overwhelmingly rewards lying cheating and stealing.
How much lying cheating and stealing is rewarded vary A LOT between societies. But yes, Sillicon Valley and current economical environment do reward those a lot.
I mean, what do you expect? They rely on models that were trained on non-curated data, texts originally written by lying, cheating and stealing humans. They can't be better than the source. Even with reinforced learning this can't be undone or made better.
Quite the contrary I suspect that it even helps the LLMs to better hide their inherited bad traits more successfully because they get punished for getting caught, not for giving immoral or lazy answers. They have no conscience since they are just predictions matrices trained for success and failure alone, not for living "a good live" or being a good "person".
Yes but has the author realized maybe the models are simply acting in the best interest for increasing shareholder value? /s
There are many ways to be wrong, but only a few ways to be right.
LLMs need to optimize for short-term objectives as the currently do, AND ethics-aligned outcomes.
Mechanically, the EAOS ethics-aligned outcome score should be what we rank otherwise-satisfactory outcomes by. And anything below a particular threshold should be rejexted outright.