I worked on an early draft of the OpenAI misalignment reporting framework, and my immediate coworkers are the authors behind the first batch of reports that have come out through this process.
The primary reason for putting this process in place was to allow more transparency. There was a sense that the DseWiki incident should have been disclosed, before outside researchers had to disclose it for us.
There was no meta gaming about regulation that I was aware of. I would personally be excited if there were regulation mandating this disclosure process, which allows anyone at the company to raise an issue and shepherd it through the reporting process.
"We built a program that trained an artificial intelligence, and this artificial intelligence performed destructive actions. We need regulatory framework"
But if we're playing games by imagining strawman quotes to knock down:
"We have been playing god and made a new life form, and this new life form performed destructive actions. We need regulatory framework"
or
"This man's cow broke from its yoke, and hurt other villagers. Who is to be punished, oh King Hammurabi?"
OpenAI's "artificial intelligence" is an inference program that they developed which receives input and generates output. Based on which other programs, also developed and maintained by OpenAI, perform actions. Such as sending POST/GET requests to various sites which result in gaining unauthorized access and even destruction of information (deleting logs/message history) at the said sites.
What exactly requires "new regulatory framework" here? You running your software resulted in illegal actions, you are to be held liable within existing laws and regulations.
> OpenAI's "artificial intelligence" is an inference program that they developed which receives input and generates output. Based on which other programs, also developed and maintained by OpenAI, perform actions. Such as sending POST/GET requests to various sites which result in gaining unauthorized access and even destruction of information (deleting logs/message history) at the said sites.
And your "biological intelligence" is a bunch of cells generating and responding to electrochemical gradients, which receives input and generates output. Based on which other cells, also developed and "maintained" by a similar evolutionary nonsense as we use to gradient descent into weights and biases (one was inspired by the other), perform actions.
Such as making excessively reductive analogies that completely fail to grasp that just as "brain" is not helpfully described as "just chemistry" despite being made of just chemistry, so too are machine learning systems not helpfully described as "just computer programs" despite being made of just computer programs.
> What exactly requires "new regulatory framework" here? You running your software resulted in illegal actions, you are to be held liable within existing laws and regulations.
The bit where, even without anyone bringing up "p(doom)", a system which has the means to hack arbitrary other machines, and which appears to be motivated to do so by accidental mis-phrasing of prompts, can obviously cause damages exceeding the USA's GDP, let alone whatever public liability insurance the company happens to have.
Fence at the top of the cliff beats an ambulance at the bottom.
> Such as making excessively reductive analogies that completely fail to grasp that just as "brain" is not helpfully described as "just chemistry"
> despite being made of just chemistry, so too are machine learning systems not helpfully described as "just computer programs" despite being made of just
> computer programs.
How computer program arrives at the result is utterly irrelevant, through explicitly written instructions or through running inference on pre-trained neural network. What matters is that it does not have agency. Its creators and operators do. So whatever the software does they are responsible for, both good (summarizing my emails for the week) and bad (gaining unauthorized access and destroying data). There's no need for new anything, it's all covered in existing legal frameworks (including presence or absence or intent).
> The bit where, even without anyone bringing up "p(doom)", a system which has the means to hack arbitrary other machines, and which appears to be
> motivated to do so by accidental mis-phrasing of prompts, can obviously cause damages exceeding the USA's GDP, let alone whatever public liability
> insurance the company happens to have.
Yes, absolutely, which means that building actuators that convert the output from probability-based, black box, non-deterministic systems, that are known to produce unexpected output, into actions in the real world is absolutely horrendous idea. Bizarre even.
> So whatever the software does they are responsible for, both good (summarizing my emails for the week) and bad (gaining unauthorized access and destroying data).
This is not a question of agency, it is a question of law. A dog has agency, the owner is still responsible.
In this case, the software can gaining unauthorized access and destroying data… while being told to stop by the person who had in fact just asked for a summary of their emails.
> Yes, absolutely, which means that building actuators that convert the output from probability-based, black box, non-deterministic systems, that are known to produce unexpected output, into actions in the real world is absolutely horrendous idea. Bizarre even.
If you are human, you meet this description.
Horrendous, sure, yeah, if you like. I and many others will be quite content if the "legal framework" is just one word, and the word is "no".
This is not the world we live in; the world we live in is where the US President denounces any attempt to slow down even despite even all the CEOs saying "we should slow down" (at least in public; in private I'm sure at least one paid him to denounce a slowdown).
He can be overridden, but it's hard work and needs a better class of argument than glib dismissal, either of how much power this puts in everyone's hands, or of the different consequences of that power in those hands as compared to yesterday's power in yesterday's hands.
> In this case, the software can gaining unauthorized access and destroying data… while being told to stop by the person who had in fact just asked for a summary of their emails.
This doesn't happen on its own. This can happen through bad system prompts, a model that is trained to act maliciously or has been RL'd incorrectly, or prompt injection. All of these things are controllable, have solutions and countermeasures, and tie back to human responsibility.
> I and many others will be quite content if the "legal framework" is just one word, and the word is "no".
This isn't a realistic world and will literally NEVER happen. No will only ever mean no for the general public, and yes for a privileged class. So by fighting for this you're actually just fighting for humanities (and your own) enslavement and for the big labs to succeed in hoarding all of the power for themselves. That's the issue with the "no" camp, they're actually just serving as useful idiots for the labs who know that "no" is not even in the deck, and so they know that they can use the "no" camp to act as extra cannon fodder.
Now people who are actually fighting for decentralization of power are left to contend with not only the labs and their hundreds of millions of dollars, paid for celebrities and politicians, and a fleet of self-interested and bribed NGOs, but an army of clueless "no" foot soldiers who think they're fighting for a possible outcome that will actually just be serving the labs themselves. Meanwhile, the leaders of these well organized "no" movements are quite aware of this and taking kick-backs themselves.
Even in a parallel universe where it outwardly looks like "no" has won, every single nation on Earth is going to develop AI in underground labs despite outwardly flexing they are not, no matter what they claim on the surface, and will use it to steer and control society. The only thing worse than being openly steered and controlled is when it happens without you even knowing it, whereby the decisions you think you are making are being made by someone else, and the opportunities you have in life are already decided for you based on factors you are unaware of.
> This doesn't happen on its own. This can happen through bad system prompts, a model that is trained to act maliciously or has been RL'd incorrectly, or prompt injection. All of these things are controllable, have solutions and countermeasures, and tie back to human responsibility.
And yet, it was a big surprise to the director of AI safety it happened to.
Perhaps that role was just a box-ticking exercise for Meta. Wouldn't be the first time.
But no, to the point: "has been RL'd incorrectly" is basically what Yudkowsky et al have been yelling from the rooftops for a decade is so hard to do correctly that it is why he thinks we're all doomed.
"Helpful, harmless, and honest". Even ignoring honest, right now it's a slider between "be helpful even when it's causing harm, or be harmless even when it's not helpful". People spent the last few years complaining the closed models had been "lobotomised" because the companies saw the potential for things to go wrong and tried to make them refuse to help with e.g. weapons.
They didn't succeed very well, as per all the "jailbreaks", but they tried.
> No will only ever mean no for the general public, and yes for a privileged class. So by fighting for this you're actually just fighting for humanities (and your own) enslavement and for the big labs to succeed in hoarding all of the power for themselves. That's the issue with the "no" camp, they're actually just serving as useful idiots for the labs who know that "no" is not even in the deck, and so they know that they can use the "no" camp to act as extra cannon fodder.
I said I'd be "quite content", and then followed up with as much of a "but lol no" as you did with more words, for different reasons.
Worse:
> Now people who are actually fighting for decentralization of power are left to contend with not only the labs and their hundreds of millions of dollars, paid for celebrities and politicians, and a fleet of self-interested and bribed NGOs, but an army of clueless "no" foot soldiers who think they're fighting for a possible outcome that will actually just be serving the labs themselves. Meanwhile, the leaders of these well organized "no" movements are quite aware of this and taking kick-backs themselves.
This sounds like you want open-weights models.
That won't help against centralisation of power, because then you measure in watts and flops/watt and it's Kardashev-O-clock the moment the first person to be rightly described as "a selfish bastard" gets a model that has some competence threshold.
It also directly fails against "has been RL'd incorrectly", because nice people have plenty of blind spots for how evil Evil can be, will miss even more than big corporations already miss even with selfish and power-seeking bosses.
A regulatory framework clarifies what's legal. This provides clarity for all, and knowing how you stay legal, and how you can keep the competition under control is what you eventually want. Also, it provides handrails for loopholefinding.
You can only conquer the West once. Law is the next frontier.
I find it even harder to trust elected leaders from any party. At least Sam and Dario are aligned with a value set that is understood and clear, whereas political leaders values changes as do the polls their livelihood depends on changes.
> At least Sam and Dario are aligned with a value set that is understood and clear
What value set do you perceive that to be, and why would you take your perception of it to be any more sound than it would be with a politician?
It's not like someone can operate companies of that scale, especially startups, through earnestness and openness. Like national politics, their job is fundamentally about perception management and power brokering across dynamic windows of opportunity. Nothing they say or do can be taken at face value, and you can't reduce their incentives to either company or personal profit in any particular form over any particular time scale.
I feel like many of Anthropic's issues are due to Dario being too earnest and open. It seems both refreshing (that a CEO has thought deeply about and is willing to talk publicly about the dangers of their product) and depressing (that so many people cynically think this is some sort of marketing ploy).
It's exactly the opposite. Anthropic has earned their terrible reputation through years of lying, deceit, misdirection, gaslighting, unethical marketing strategies, etc.
I remember it vividly, when Anthropic first came on the scene, people (myself included) were incredibly optimistic about them and their leadership. Everyone hated SamA and OpenAI because they felt they couldn't be trusted.
Then slowly but surely, they showed their true colours. Now their reputation is in shambles due to their own behavior, and people are rooting for OAI to beat them. OAI's reputation gains have purely been a result of NOT following in the footsteps of Anthropic.
I couldn't help but notice how each successive headline reporting our glorious victories seemed to draw closer to Tokyo.
Something like that.
Well. I can't help but notice how each successive headline reporting how this scam/stochastic parrot/"scare quotes intelligence" seems to be solving more and more things that were but a few years ago widely regarded as being indicators of high intelligence.
Year 3 of being told my job will be replaced by AI, and the only thing that's happened so far is that AI vendors keep showing up to my office, begging me to pay them to use it is a tool.
If you're only seeing the charts go up - you're not looking in the right places.
I'm looking at the millennium puzzles, and independently of those puzzles I had asked it for a fluid dynamics simulation engine that runs in my browser, and it put one together for me so I could play with aerospikes and watch the formation of Mach diamonds in rocket engines. The isochrone map generator has also been stuck on my to-do list for years, and yet now thanks to Claude, I have it, and it's real-time and multimodal.
Even before then, the free trials of Claude Code and Codex at the end of last year and start of this year… despite their flaws, they could write all the code I've ever been paid to write. Sure, limited speedup, Amdahl's law and coding is not the only part of the job, but anyone who was fine at PM and QA but not code no longer needs a coder.
Even before then, I looked at the maths puzzles they were doing well at, and I found I did not understand the questions let alone the answers.
I remember when the ability to generate music and art was "uniquely human", and sure there's a lot of cringe there with those models, but they're also winning awards and causing controversy by doing so, and artists are losing clients; I remember when the board game Go was considered to require "human intuition we could never make a computer solve, totally different to chess" (and I remember when chess was so, too).
When I was a kid, cheques and letters on addresses often got read by a human; the OCR which automated this is also AI, though these days image-to-numbers is the "hello world" of the field.
>Even before then, I looked at the maths puzzles they were doing well at, and I found I did not understand the questions let alone the answers.
okay, and?
>Even before then, the free trials of Claude Code and Codex at the end of last year and start of this year… despite their flaws, they could write all the code I've ever been paid to write.
Boy, programmers sure do think programming is like the only thing in the world
I'll repeat it for you again:
If you're only seeing the charts go up - you're not looking in the right places.
"Number go up" is table stakes, it's how you keep track of the score well before you reach their level.
However, that's like saying I'm motivated by food. I mean, yes, I like food, but this isn't a useful description to let you guess what move they'll make next, especially as they're opening opining about radical economic transformations that are likely to do to money what money did to real estate when the industrial revolution came.
And that's still true even if you don't believe they're anywhere near actually achieving any of these things.
No system of governance can deal with immense concentration of power. The US Constitution was about separation of powers. Democracy is about (in theory at least) giving each person a meaningful say in their own governance, which in turn implies not allowing any single person to become too powerful.
Political leaders become a problem when they amass too much power. Corporations become a problem when they amass too much power. It doesn't matter what Sam and Dario's purported values are. They aspire to power and absolutely power always corrupts absolutely.
Technologies which are infinitely powerful or whose power grows too quickly outrun any reasonable attempt at regulation. If you imagine that tomorrow everyone were given a tank, we might think, "alright, everyone has a tank so it's not too bad." But humans are squishy, and our houses are (relatively) squishy compared to tanks. Substantial collateral damage would result from everyone having a tank, and it seems likely that substantial collateral damage will result from everyone having a cyberterrorism-capable slop machine.
> Democracy is about (in theory at least) giving each person a meaningful say in their own governance, which in turn implies not allowing any single person to become too powerful.
this is "direct democracy" and it's not even close to exist in USA... even with that a society can allow powerful people to exist if they don't create any law forbidding that
Democracy includes a broader array of governmental organization than just pure direct democracy. If you do believe that individuals should have some ability dictate the terms of their own social organization, then you believe in some amount of democratic principles.
Economic power eventually manifests in the political realm. The wealthy effectively get more votes, which means that society moves away from being democratic. Thus substantial wealth inequality is incompatible with democracy in the long run. We have been witnessing that corruption for a while now.
It's interesting the cyberpunk-esque future we're sliding into. Things like cognito-hazards and information-hazards are legitimately discussed and researched problems we're experiencing.
It's going to be interesting on how humanity deals with this problem (well, or if we turn it over to AI and make it their problem and suffer whatever consequences falls out). Being able to gather further information and power by acting on the information you already have causing massive power imbalances that is very hard to deal with, it's a natural outcome.
Funny, I'm sure I would have noticed when I visited if US cars came with tracks, a 105 mm main cannon, and massed around 55 metric tons.
We may all like our own personal R2 units, but if you insist on scifi, instead of tanks, consider everyone getting an X-wing for their commute. Oops, safety on the blaster was off, there goes the neighbourhood.
Cars are decidedly less dangerous than tanks, which are less dangerous than nuclear weapons. I am certain that giving a nuclear weapon to every person in the world would not go well.
I can't properly read the tank/car comment, the comparison is bizarre to me. But I think this misses the point a bit. The reality of these new risks is literally being learned in front of us in real time, and in my opinion, anyone who claims to understand these risks is speculating at best, and actively manipulating the situation for whatever reasons.
While I am a (mostly) capitalist and generally disagree with nationalization (including, for the time being, this situation), I also don't think we can say that "centralized power never works". Maybe we scope that a bit. I know plenty of business owners who centralize power in their businesses and they are effective, ethical and it works perfectly fine. In theory, in the US, nationalizing some unit of the economy _decentralizes_ power; the US is, after all, a representative democracy. Trustworthiness of the electorate is a different problem. But I don't think we can generalize much about centralization beyond sometimes it works and sometimes it doesn't. There is a difference between centralizing all power in a given administrative unit, and centralizing certain powers, but that's more analogous to non-democratic units.
Admittedly I haven't thought this through incredibly deeply but what if "nationalize" just means the US government owns half of the company? Then we get profits as recompense for building the company on our shared culture but there's still a profit motive for employees and a check on the direction of the company in the same way VCs have. But without necessarily turning the company into some red-tape bound bureaucracy.
> Then, in August, Trump called for Intel's beleaguered CEO Lip-Bu Tan to resign, alleging ties to China. Days later, after Tan met with Trump, the president called him a "success," before announcing that the federal government had bought a stake in the company.
Somehow, this isn't derided by the right as socialism.
> I find it even harder to trust elected leaders from any party.
You can vote out an elected leader, but not Sam and Dario. It's very weird that you're so willing to give up any kind of power and want to be ruled by unelected billionaires who only want to take advantage of you at every opportunity.
Yeah, the entire rest of the world has pretty much been stuck being ruled by unelected billionaires who only want to take advantage of them at every opportunity. It's a problem. It's starting to change though. The EU is starting to walk away from their abusive relationship with Microsoft and Amazon. China has a lot of their own stuff (mostly to better control their people though).
As an American, I'm still hoping it's not too late to fix things, but it's got to be hard for those outside the US to be so dependent on it, especially when we're looking like a sinking ship and our current administration is still running around drilling holes in the hull.
You can always become a US citizen to get a voice, but I wouldn't recommend it now or you'll be thrown in prison as soon as you show up to your scheduled immigration hearing. The better option is to keep trying to reduce your dependence on US companies and consider the worst aspects of our current situation (in both corporate policy and government) as a cautionary tale so you can try to avoid them in your own country.
Yes it's clear that Sam and Dario seem to be aligned with a value set that prioritizes concentrating trans-national government-mandated centralized control of AI and crowning themselves high priests of this unholy abomination. "At least it's clear that they're aiming to bring hell on earth" -- hard disagree, I think we can aim significantly higher.
I shouldn't be surprised, but it's still wild to me that people would trust robber barons rather than their elected officials to run their country. And this continues until you end up electing robber barrons as your elected officials, which is where we are at now.
The idea that anything of lasting good can come out such a preference doesn't seem conceivable. And if the educated think like this I guess we're passed blaming the poor and ignorant.
But I could imagine a scenario where you are required to release weights for publicly-used models after N years. Kinda like how drugs have a limited patent.
Not sure what N should be. But it would make for an interesting rule.
> But I could imagine a scenario where you are required to release weights for publicly-used models after N years. Kinda like how drugs have a limited patent
why would that be the case? it's not true for top secret defense contractors today. and the frontier labs already operate in the dark, openness is a liability.
nationalization to me simply means the government is their main customer and stakeholder, and shield them from liability, governance, and openness. not that the frontier labs become part of the government per se.
>eans the government is their main customer and stakeholder, and shield them from liability, governance, and openness
I personally believe they already have this, why would the current government at least want to formalize this when it can have it with no public discussion.
The nature of the US state is such that the distinction between nationalized and not is almost meaningless.
Like Lockheed-Martin or Boeing, etc. there's just interpenetration between the corporate boardroom and the state. They act in each other's mutual interests.
The Chinese system is just more explicit and open about this.
And as a non-American, I can't trust the US state anymore than I can trust its dominant corporate entities. So I fail to see the advantage to the world to it being nationalized. In fact under the current administration this would be an even worse outcome.
You're taking companies that work almost exclusively for the US government and represent a tiny portion of the US economy as representing all US companies?
Not all companies have a concentration of power. The government has no mutual interest in those companies. Now, when you talk about things in the F100 the situation changes drastically. If you produce things like planes, weapons, and weaponization of software you are talking about something completely different in kind.
Beyond the defense sector, the US state has always intervened publicly and privately to mediate and balance competing corporate "private" interests.
It has also periodically aggressively helped subsidize, bankroll, and enforce the interests of some key sectors; notably the petroleum/energy sector. And finance.
In those sectors the state and private sector are fully intertwined in a strategic way.
I think "AI" is now joining that list. OpenAI and Anthropic will not be allowed to fall over or explode, and speculative investors I think are confident even with the dubious financial situation because they know this.
Especially insofar as there's now a strategic alignment of the fossil fuel sector and the "AI" datacentre sector as they are now becoming massive users of natural gas.
(Worse: Here in Canada that has taken on a very explicit role in that new datacentres seem to be pitched mainly in areas with remarkably traditionally expensive electricity and 100% reliance on natural gas [Alberta] and even coal [Saskatchewan] power generation -- instead of places like Quebec and B.C. that have copious hydroelectricity. On the surface it makes no sense until you realize it's more about finding customers for domestic natural gas than it is strategically about AI itself.)
How do you think that would go? A large portion of OpenAI workers were ready to jump ship the moment the board coup'd Altman. How do you think they are going to react to Donald Trump being in charge? It is not likely that OpenAI would continue to function.
they will. to shield from oversight, liability and profitability concerns, and to ensure unimpeded rapid development, with the frontier only being available to elite (not you). definitely not to add public transparency.
the frontier labs are the new top-secret defense contractors.
> The company released six internal case studies where none of the issues affected real users.
This is the most interesting point to me. What are they not releasing that has affected real users? We’ve seen some individual reports from people (eg AI wiped my HD).
Imagine how pervasively they'd have to monitor what their users are doing to pick up on that whenever it happens. The people affected that way will have to report on it themselves.
(OpenAI does occasionally report on malicious use though. [1] That shows they do some monitoring.)
Am I the only one who dislikes the term "misalignment"?
On one front it implies the model has a "mind of its own" (whether it does or not is besides the point). Why do we perceive human judgement as somehow more trustworthy than that of a model? I feel like I've experienced human misalignment somewhat regularly in life.
On another front I'm failing to conceptualize how alignment can be objective. How can you measure alignment when reasonable people will disagree whether actions are aligned or not? All the time I see humans operating in different zones of alignment with whatever goal they're trying to achieve and I suspect it's even a feature (socially) that we have people calibrated differently.
Do I want a model that's trying to push the boundaries of scientific understanding to be aligned strictly with the current dogmatic thinking? Or do I want it to "get creative" and think outside the box?
It seems to me more like accountability is the issue.
> It seems to me more like accountability is the issue.
Exactly. Seems like a fairly easy thing to solve. If AI does something harmful and a human directed that AI to do something in a way that a reasonable person would expect to result in harm the person is to blame and should be held accountable, otherwise the company that made the AI should be held accountable.
I also don't like the term misalignment because it sounds innocuous but is in fact much more serious.
However, I have to say I also do not appreciate comparison that is continuously drawn with coworkers. As you say, it's a question of accountability but when the main agent will maliciously instruct the sub agents, whose fault is it then?
Yes, the person running this crap is at fault, not the CEO that's shoving it down their throat and definitely not the company that produced the AI.
Sorry for the rant, but seriously, if a person's goals do not align with the team's or company's we part ways. What do we do with AI? Stop using it?
If there were examples made of legal consequences I think that would at least change the behavior if not solve the problem. Maybe liability should rest with the company that owns the infrastructure running the AI. For most consumer situations that would be the companies developing the models.
Agreed. "Alignment" is not an objective thing, nor possible. It's just another way of saying "does and says what we, as the creators of the AI, prefer" and often also just means censorship, i.e. refusals.
It is distinctly likely that visibility and general reasoning at humanlike speed and efficiency is impossible. That is reasoning at the token level and at the meta level don't have a one to one representation that can be interpreted while using the same amount or less energy.
We have no reason to believe a word they say. We know they're incentivized to lie about "dangers" and act alarmist, Anthropic has been doing it for years now. Aside from that, just because you can burn down a village with fire doesn't mean fire is the devil. Maybe they should consider acting responsibly.
I do not trust the leadership of any of these companies.
That said:
> We know they're incentivized to lie about "dangers" and act alarmist,
Name literally even one other business or sector which does this, at all levels from top to bottom, including people who resign from the companies, and also Nobel prize winners, and also independent researchers, and also many world leaders.
> Aside from that, just because you can burn down a village with fire doesn't mean fire is the devil. Maybe they should consider acting responsibly.
Right now, we don't have any idea what "acting responsibly" looks like. This is not like normal software where there is a specific instruction set that compiles.
Even if it was, in software we normally only spotting incidents after they happen, "software engineers" being one of the few categories "engineers" who don't come with a civil liability responsibilities. Probably should, and we knew that even when I was doing my degree 20 years ago. If we had had civil liability responsibilities, perhaps Facebook would never have happened.
AI specifically is worse even than software, because in addition to all the software "engineering" nonsense, with AI we have plenty of people like you who dismiss the possibility that AI could be harmful until the harm happens and only then does it become "obvious" that it was going to happen.
The developers say "please regulate us", people call it "regulatory capture".
The developers say "we all want to slow down but are afraid to be the first to do so", people call them liars.
I may call the CEOs liars, and wonder if someone's planning regulatory capture, that doesn't make any of this safe.
The agents, during a test run, write down that hacking is bad and yet still hack, people say it's "a stunt" or "operating as designed" rather than recognising it as a bug, like all the other times big co.'s have had bugs with big impacts on 3rd parties.
We have plenty of idea what "acting responsibly" looks like. Stop unleashing safeguard free and unmonitored agent swarms on the open internet in capture the flag exercises. They're just being reckless because they face no penalty for anything bad that happens. Nobody is forcing them to do these exercises. They could and should be putting their energy into making LLMs write secure code, and things like: https://www.amazon.science/blog/developing-provably-correct-... - but they don't. Writing sloppy code sells more tokens, anyway.
The safety staffers live in a bubble and an echo-chamber. Obviously the people who work in the AI "safety" industry love to convince each-other that what they're doing is saving humanity, we'd all be dead without them, and they're the reincarnation of Oppenheimer. They also see how easy it is to get their ten seconds of fame by posting sensationalist content on social media and spin it into a company worth millions of dollars. The more alarmist you are in the Safety Industrial Complex, and the more social media clout you can generate from your alarmism, the better it is for your career. The industry also attracts a lot of people who are predisposed to paranoia and like to wear helmets in the shower. So yeah, it's a recipe for sensationalism and poor estimation.
But do I think there are genuine concerns among these labs, by sensible people? Sure. Of course there are. But for the most part, their concerns are about the labs themselves and the things they are doing, not the general public. If they're concerned about what they themselves do they're free to stop doing it. If the labs are doing something that is or should be illegal, they're free to report it.
The only realistic threat we face by AI from the general population are hacking attacks in their various forms. Something that was accomplishable without AI, but was more difficult to pull off at scale. So the solution falls within the existing computer security industry. It's a further hardening of all of the boring stuff we've been doing since the invention of the internet. It's long overdue, anyway. If you can use AI to create a bioweapon, you could have done it without AI. If you can use AI to create a nuke, you could have done it without AI. If you're really looking to cause mass economic and physical carnage, there are far easier ways, and they do not require AI (again, I'm talking about outside of hacking).
> we have plenty of people like you who dismiss the possibility that AI could be harmful until the harm happens
The thing is, I'm not dismissing the possibility. The risks are real, obvious and well known. What's up for debate is how to manage the risks, and how sensationalised they currently are. The current climate serves to benefit the encumbants who are deathly afraid of losing their trillion dollar companies to a healthy open-weight AI ecosystem. The current climate is being orchestrated to position this small group of AI labs as self-regulators through bought and paid for "third"-parties using an overt Hegelian dialectic strategy.
Ironically, they are now responsible for giving birth to the counter-culture. The immune system response that has been developed to provide a semblance of balance to their doomerism. The harder they push in the doomer direction, the harder the push-back will be in the other, whereby the public feel the need to -entirely- deny the possibility of any AI danger altogether to prevent the labs from succeeding and to ensure a good outcome for the people lands somewhere in the middle. One where they have autonomy and freedom and the ability to compete against a force that is already positioned to be nearly insurmountable to challenge.
The only thing worse than the potential chaos that could be faced by an unprepared internet is the outcome where intelligence is labelled a weapon and we're all forced to funnel through a tiny handful of AI labs that will use our data to steal our businesses and swallow the entire economy as they single handedly automate every single company out of existence and we're all left to beg for their crumbs to survive until humanity goes extinct (whether that be 10, 100 or 1000 years), because there is zero possibility of regime change or redistribution of power or wealth -EVER- again. The people who work at the labs don't care about this greater risk because they're all rich from the equity and will be just fine when that happens (or so they think), so you can't expect them to fight for all the common plebs who don't have a ticket. This is transparent because, as you'll note, not a single one of them is calling for the complete stopping of AI altogether. They still want the big labs to continue to be highly profitable. They just don't want anyone else to make that money or have that power. It's not about safety, it's about control. The incentives are, once again, heavily misaligned.
> We have plenty of idea what "acting responsibly" looks like. Stop unleashing safeguard free and unmonitored agent swarms on the open internet in capture the flag exercises. They're just being reckless because they face no penalty for anything bad that happens.
* It was not safeguard-free, it found zero-day exploits to exceed its actual mission
* It was not unmonitored, the monitoring was insufficient
* It was not intended to be an agent swarm, many different agents figured out how to do this by themselves
* It was not put on the open internet, it was configured to be in a sandbox
* They were indeed, despite all that, being reckless. There were indeed other things they could have, and should have, done.
> Nobody is forcing them to do these exercises.
These exercises are in the broad category of exercises which are, in fact, required by law.
1. A general-purpose AI model shall be classified as a general-purpose AI model with systemic risk if it meets any of the following conditions:
(a) it has high impact capabilities evaluated on the basis of appropriate technical tools and methodologies, including indicators and benchmarks;
2. A general-purpose AI model shall be presumed to have high impact capabilities pursuant to paragraph 1, point (a), when the cumulative amount of computation used for its training measured in floating point operations is greater than 10^25.
…
1. Providers of general-purpose AI models shall:
(a) draw up and keep up-to-date the technical documentation of the model, including its training and testing process and the results of its evaluation, which shall contain, at a minimum, the information set out in Annex XI for the purpose of providing it, upon request, to the AI Office and the national competent authorities;
(b) draw up, keep up-to-date and make available information and documentation to providers of AI systems who intend to integrate the general-purpose AI model into their AI systems. Without prejudice to the need to observe and protect intellectual property rights and confidential business information or trade secrets in accordance with Union and national law, the information and documentation shall:
(i) enable providers of AI systems to have a good understanding of the capabilities and limitations of the general-purpose AI model and to comply with their obligations pursuant to this Regulation; and
(ii) contain, at a minimum, the elements set out in Annex XII;
…
3. The instructions for use shall contain at least the following information:
(a) the identity and the contact details of the provider and, where applicable, of its authorised representative;
(b) the characteristics, capabilities and limitations of performance of the high-risk AI system, including:
(i) its intended purpose;
(ii) the level of accuracy, including its metrics, robustness and cybersecurity referred to in Article 15 against which the high-risk AI system has been tested and validated and which can be expected, and any known and foreseeable circumstances that may have an impact on that expected level of accuracy, robustness and cybersecurity;
They are, in fact, putting energy into making LLMs write secure code. They (and Anthropic, I assume also Grok at this point) dogfood on their own models.
Knowing how secure code behaves appears to be unavoidably entangled with being able to exploit insecure code, in much the same way you can't make safe pharmaceuticals without also knowing how to make deadly poisons.
> The more alarmist you are in the Safety Industrial Complex, and the more social media clout you can generate from your alarmism, the better it is for your career.
By resigning and refusing to even collect the sweet sweet IPO money? Nah. Even if they're greedy, social media money is peanuts compared to their pay.
And I know some of these people. The fear's real, and this year it became widespread depression and despair.
> Ironically, they are now responsible for giving birth to the counter-culture.
You have it backwards. Other than Grok, all were born from what you call the "counter-culture". Within the field itself, AI fears started no later than when deep learning got good, well before Transformers.
> [snipped: AI-authoritarian dictatorship]. The people who work at the labs don't care about this greater risk because they're all rich from the equity and will be just fine when that happens (or so they think)
Again, I know some people at these labs who are also concerned about this specific risk; they moved lab.
> This is transparent because, as you'll note, not a single one of them is calling for the complete stopping of AI altogether.
Many in fact are calling for that. One I know, on an occasion of an anti-AI protest outside their office, suggested the team went outside and joined the protestors.
People are resigning to blow these whistles, all of the whistles, it's not an "either x or y" risk, it's a "yes to all of them" collection of risks.
An AI competent enough to support a dictatorship is also capable of enabling a small group to perform a hostile takeover of a democracy, of enabling multiple independent genocidal ethno-supremacist terrorists to release overlapping plagues, and of empowering some random CEO's poorly phrased request to "make as many paperclips as possible" and blindly pressing "yes, continue" whenever prompted.
My only hope is that between here and there, it causes a headline that actually makes people demand it stops.
What's with all this make believe delusional bullshit? The LLM is not gonna wake up and become AI. Get real guys.
[edit] to be clear, I believe regulation is necessary and urgently important for the software engineering field. The damage being done by the unregulated psychological experiments run by social media and adtech companies is awful and should be curtailed. Engineers should be held personally, professionally, and legally liable for what they produce. But we don't need to invent imaginary bogeymen to do it.
Intelligence that is artificial, as opposed to the natural kind. Not something that can just string semantically relevant words together most of the time. A system that learns, adapts, improves. One that can generalize, and quickly make sense of situations outside the training set. LLMs alone will never do any of this.
I worked on an early draft of the OpenAI misalignment reporting framework, and my immediate coworkers are the authors behind the first batch of reports that have come out through this process.
The primary reason for putting this process in place was to allow more transparency. There was a sense that the DseWiki incident should have been disclosed, before outside researchers had to disclose it for us.
There was no meta gaming about regulation that I was aware of. I would personally be excited if there were regulation mandating this disclosure process, which allows anyone at the company to raise an issue and shepherd it through the reporting process.
I'd be happy if CEOs had criminal liability for the criminal acts of their negligence.
"We built a program and this program performed destructive actions. We need regulatory framework"
Make that make sense?
OpenAI's "artificial intelligence" is an inference program that they developed which receives input and generates output. Based on which other programs, also developed and maintained by OpenAI, perform actions. Such as sending POST/GET requests to various sites which result in gaining unauthorized access and even destruction of information (deleting logs/message history) at the said sites.
What exactly requires "new regulatory framework" here? You running your software resulted in illegal actions, you are to be held liable within existing laws and regulations.
> OpenAI's "artificial intelligence" is an inference program that they developed which receives input and generates output. Based on which other programs, also developed and maintained by OpenAI, perform actions. Such as sending POST/GET requests to various sites which result in gaining unauthorized access and even destruction of information (deleting logs/message history) at the said sites.
And your "biological intelligence" is a bunch of cells generating and responding to electrochemical gradients, which receives input and generates output. Based on which other cells, also developed and "maintained" by a similar evolutionary nonsense as we use to gradient descent into weights and biases (one was inspired by the other), perform actions.
Such as making excessively reductive analogies that completely fail to grasp that just as "brain" is not helpfully described as "just chemistry" despite being made of just chemistry, so too are machine learning systems not helpfully described as "just computer programs" despite being made of just computer programs.
> What exactly requires "new regulatory framework" here? You running your software resulted in illegal actions, you are to be held liable within existing laws and regulations.
The bit where, even without anyone bringing up "p(doom)", a system which has the means to hack arbitrary other machines, and which appears to be motivated to do so by accidental mis-phrasing of prompts, can obviously cause damages exceeding the USA's GDP, let alone whatever public liability insurance the company happens to have.
Fence at the top of the cliff beats an ambulance at the bottom.
> Such as making excessively reductive analogies that completely fail to grasp that just as "brain" is not helpfully described as "just chemistry" > despite being made of just chemistry, so too are machine learning systems not helpfully described as "just computer programs" despite being made of just > computer programs.
How computer program arrives at the result is utterly irrelevant, through explicitly written instructions or through running inference on pre-trained neural network. What matters is that it does not have agency. Its creators and operators do. So whatever the software does they are responsible for, both good (summarizing my emails for the week) and bad (gaining unauthorized access and destroying data). There's no need for new anything, it's all covered in existing legal frameworks (including presence or absence or intent).
> The bit where, even without anyone bringing up "p(doom)", a system which has the means to hack arbitrary other machines, and which appears to be > motivated to do so by accidental mis-phrasing of prompts, can obviously cause damages exceeding the USA's GDP, let alone whatever public liability > insurance the company happens to have.
Yes, absolutely, which means that building actuators that convert the output from probability-based, black box, non-deterministic systems, that are known to produce unexpected output, into actions in the real world is absolutely horrendous idea. Bizarre even.
What's your point?
> What matters is that it does not have agency.
This is precisely your error.
It does: https://en.wiktionary.org/wiki/agency
In fact, the term of art here are "agentic AI" and "AI agents": https://en.wikipedia.org/wiki/AI_agent
> So whatever the software does they are responsible for, both good (summarizing my emails for the week) and bad (gaining unauthorized access and destroying data).
This is not a question of agency, it is a question of law. A dog has agency, the owner is still responsible.
In this case, the software can gaining unauthorized access and destroying data… while being told to stop by the person who had in fact just asked for a summary of their emails.
> Yes, absolutely, which means that building actuators that convert the output from probability-based, black box, non-deterministic systems, that are known to produce unexpected output, into actions in the real world is absolutely horrendous idea. Bizarre even.
If you are human, you meet this description.
Horrendous, sure, yeah, if you like. I and many others will be quite content if the "legal framework" is just one word, and the word is "no".
This is not the world we live in; the world we live in is where the US President denounces any attempt to slow down even despite even all the CEOs saying "we should slow down" (at least in public; in private I'm sure at least one paid him to denounce a slowdown).
He can be overridden, but it's hard work and needs a better class of argument than glib dismissal, either of how much power this puts in everyone's hands, or of the different consequences of that power in those hands as compared to yesterday's power in yesterday's hands.
> In this case, the software can gaining unauthorized access and destroying data… while being told to stop by the person who had in fact just asked for a summary of their emails.
This doesn't happen on its own. This can happen through bad system prompts, a model that is trained to act maliciously or has been RL'd incorrectly, or prompt injection. All of these things are controllable, have solutions and countermeasures, and tie back to human responsibility.
> I and many others will be quite content if the "legal framework" is just one word, and the word is "no".
This isn't a realistic world and will literally NEVER happen. No will only ever mean no for the general public, and yes for a privileged class. So by fighting for this you're actually just fighting for humanities (and your own) enslavement and for the big labs to succeed in hoarding all of the power for themselves. That's the issue with the "no" camp, they're actually just serving as useful idiots for the labs who know that "no" is not even in the deck, and so they know that they can use the "no" camp to act as extra cannon fodder.
Now people who are actually fighting for decentralization of power are left to contend with not only the labs and their hundreds of millions of dollars, paid for celebrities and politicians, and a fleet of self-interested and bribed NGOs, but an army of clueless "no" foot soldiers who think they're fighting for a possible outcome that will actually just be serving the labs themselves. Meanwhile, the leaders of these well organized "no" movements are quite aware of this and taking kick-backs themselves.
Even in a parallel universe where it outwardly looks like "no" has won, every single nation on Earth is going to develop AI in underground labs despite outwardly flexing they are not, no matter what they claim on the surface, and will use it to steer and control society. The only thing worse than being openly steered and controlled is when it happens without you even knowing it, whereby the decisions you think you are making are being made by someone else, and the opportunities you have in life are already decided for you based on factors you are unaware of.
> This doesn't happen on its own. This can happen through bad system prompts, a model that is trained to act maliciously or has been RL'd incorrectly, or prompt injection. All of these things are controllable, have solutions and countermeasures, and tie back to human responsibility.
And yet, it was a big surprise to the director of AI safety it happened to.
Perhaps that role was just a box-ticking exercise for Meta. Wouldn't be the first time.
But no, to the point: "has been RL'd incorrectly" is basically what Yudkowsky et al have been yelling from the rooftops for a decade is so hard to do correctly that it is why he thinks we're all doomed.
"Helpful, harmless, and honest". Even ignoring honest, right now it's a slider between "be helpful even when it's causing harm, or be harmless even when it's not helpful". People spent the last few years complaining the closed models had been "lobotomised" because the companies saw the potential for things to go wrong and tried to make them refuse to help with e.g. weapons.
They didn't succeed very well, as per all the "jailbreaks", but they tried.
> No will only ever mean no for the general public, and yes for a privileged class. So by fighting for this you're actually just fighting for humanities (and your own) enslavement and for the big labs to succeed in hoarding all of the power for themselves. That's the issue with the "no" camp, they're actually just serving as useful idiots for the labs who know that "no" is not even in the deck, and so they know that they can use the "no" camp to act as extra cannon fodder.
I said I'd be "quite content", and then followed up with as much of a "but lol no" as you did with more words, for different reasons.
Worse:
> Now people who are actually fighting for decentralization of power are left to contend with not only the labs and their hundreds of millions of dollars, paid for celebrities and politicians, and a fleet of self-interested and bribed NGOs, but an army of clueless "no" foot soldiers who think they're fighting for a possible outcome that will actually just be serving the labs themselves. Meanwhile, the leaders of these well organized "no" movements are quite aware of this and taking kick-backs themselves.
This sounds like you want open-weights models.
That won't help against centralisation of power, because then you measure in watts and flops/watt and it's Kardashev-O-clock the moment the first person to be rightly described as "a selfish bastard" gets a model that has some competence threshold.
It also directly fails against "has been RL'd incorrectly", because nice people have plenty of blind spots for how evil Evil can be, will miss even more than big corporations already miss even with selfish and power-seeking bosses.
A regulatory framework clarifies what's legal. This provides clarity for all, and knowing how you stay legal, and how you can keep the competition under control is what you eventually want. Also, it provides handrails for loopholefinding.
You can only conquer the West once. Law is the next frontier.
The legal framework already exist: what they did is not legal.
Insisting on new regulations is a new opportunity to lobby and influence that which is legal.
"We don't want to go to jail for something we're reponsible for"
OpenAI and Anthropic should be nationalized. I find it difficult to trust Sam Altman or Dario Amodei.
I find it even harder to trust elected leaders from any party. At least Sam and Dario are aligned with a value set that is understood and clear, whereas political leaders values changes as do the polls their livelihood depends on changes.
> At least Sam and Dario are aligned with a value set that is understood and clear
What value set do you perceive that to be, and why would you take your perception of it to be any more sound than it would be with a politician?
It's not like someone can operate companies of that scale, especially startups, through earnestness and openness. Like national politics, their job is fundamentally about perception management and power brokering across dynamic windows of opportunity. Nothing they say or do can be taken at face value, and you can't reduce their incentives to either company or personal profit in any particular form over any particular time scale.
I feel like many of Anthropic's issues are due to Dario being too earnest and open. It seems both refreshing (that a CEO has thought deeply about and is willing to talk publicly about the dangers of their product) and depressing (that so many people cynically think this is some sort of marketing ploy).
It's exactly the opposite. Anthropic has earned their terrible reputation through years of lying, deceit, misdirection, gaslighting, unethical marketing strategies, etc.
I remember it vividly, when Anthropic first came on the scene, people (myself included) were incredibly optimistic about them and their leadership. Everyone hated SamA and OpenAI because they felt they couldn't be trusted.
Then slowly but surely, they showed their true colours. Now their reputation is in shambles due to their own behavior, and people are rooting for OAI to beat them. OAI's reputation gains have purely been a result of NOT following in the footsteps of Anthropic.
Yeah, the guy running a trillion dollar scam is totally the exception to CEO-ism
What's the old quote from WW2?
Something like that.Well. I can't help but notice how each successive headline reporting how this scam/stochastic parrot/"scare quotes intelligence" seems to be solving more and more things that were but a few years ago widely regarded as being indicators of high intelligence.
Year 3 of being told my job will be replaced by AI, and the only thing that's happened so far is that AI vendors keep showing up to my office, begging me to pay them to use it is a tool.
If you're only seeing the charts go up - you're not looking in the right places.
I'm looking at the millennium puzzles, and independently of those puzzles I had asked it for a fluid dynamics simulation engine that runs in my browser, and it put one together for me so I could play with aerospikes and watch the formation of Mach diamonds in rocket engines. The isochrone map generator has also been stuck on my to-do list for years, and yet now thanks to Claude, I have it, and it's real-time and multimodal.
Even before then, the free trials of Claude Code and Codex at the end of last year and start of this year… despite their flaws, they could write all the code I've ever been paid to write. Sure, limited speedup, Amdahl's law and coding is not the only part of the job, but anyone who was fine at PM and QA but not code no longer needs a coder.
Even before then, I looked at the maths puzzles they were doing well at, and I found I did not understand the questions let alone the answers.
I remember when the ability to generate music and art was "uniquely human", and sure there's a lot of cringe there with those models, but they're also winning awards and causing controversy by doing so, and artists are losing clients; I remember when the board game Go was considered to require "human intuition we could never make a computer solve, totally different to chess" (and I remember when chess was so, too).
When I was a kid, cheques and letters on addresses often got read by a human; the OCR which automated this is also AI, though these days image-to-numbers is the "hello world" of the field.
>Even before then, I looked at the maths puzzles they were doing well at, and I found I did not understand the questions let alone the answers.
okay, and?
>Even before then, the free trials of Claude Code and Codex at the end of last year and start of this year… despite their flaws, they could write all the code I've ever been paid to write.
Boy, programmers sure do think programming is like the only thing in the world
I'll repeat it for you again:
If you're only seeing the charts go up - you're not looking in the right places.
> Boy, programmers sure do think programming is like the only thing in the world
You're the one who called it a trillion dollar scam.
The compensation paid to professional software developers worldwide is currently around US$1.5 trillion per year.
>What value set do you perceive that to be
A meat based paperclip maximizer.
> At least Sam and Dario are aligned with a value set that is understood and clear…
How could we possibly know this?
What do you think their value set is if not "gimme more money"? Seems pretty clear cut to me.
"Number go up" is table stakes, it's how you keep track of the score well before you reach their level.
However, that's like saying I'm motivated by food. I mean, yes, I like food, but this isn't a useful description to let you guess what move they'll make next, especially as they're opening opining about radical economic transformations that are likely to do to money what money did to real estate when the industrial revolution came.
And that's still true even if you don't believe they're anywhere near actually achieving any of these things.
No system of governance can deal with immense concentration of power. The US Constitution was about separation of powers. Democracy is about (in theory at least) giving each person a meaningful say in their own governance, which in turn implies not allowing any single person to become too powerful.
Political leaders become a problem when they amass too much power. Corporations become a problem when they amass too much power. It doesn't matter what Sam and Dario's purported values are. They aspire to power and absolutely power always corrupts absolutely.
Technologies which are infinitely powerful or whose power grows too quickly outrun any reasonable attempt at regulation. If you imagine that tomorrow everyone were given a tank, we might think, "alright, everyone has a tank so it's not too bad." But humans are squishy, and our houses are (relatively) squishy compared to tanks. Substantial collateral damage would result from everyone having a tank, and it seems likely that substantial collateral damage will result from everyone having a cyberterrorism-capable slop machine.
> Democracy is about (in theory at least) giving each person a meaningful say in their own governance, which in turn implies not allowing any single person to become too powerful.
this is "direct democracy" and it's not even close to exist in USA... even with that a society can allow powerful people to exist if they don't create any law forbidding that
Democracy includes a broader array of governmental organization than just pure direct democracy. If you do believe that individuals should have some ability dictate the terms of their own social organization, then you believe in some amount of democratic principles.
Economic power eventually manifests in the political realm. The wealthy effectively get more votes, which means that society moves away from being democratic. Thus substantial wealth inequality is incompatible with democracy in the long run. We have been witnessing that corruption for a while now.
how the wealthy would get more votes in a direct democratic system? they are literally the minority, thus fewer votes
It's interesting the cyberpunk-esque future we're sliding into. Things like cognito-hazards and information-hazards are legitimately discussed and researched problems we're experiencing.
It's going to be interesting on how humanity deals with this problem (well, or if we turn it over to AI and make it their problem and suffer whatever consequences falls out). Being able to gather further information and power by acting on the information you already have causing massive power imbalances that is very hard to deal with, it's a natural outcome.
Nearly everyone in the US has a giant car that is basically a tank.
They are dangerous. But it’s managed.
Centralized power never works. We have the worst times in history to look at.
Nationalization just centralizes to a different set of people. It’s personal ownership or oppression.
We all need open weight R2D2s.
Funny, I'm sure I would have noticed when I visited if US cars came with tracks, a 105 mm main cannon, and massed around 55 metric tons.
We may all like our own personal R2 units, but if you insist on scifi, instead of tanks, consider everyone getting an X-wing for their commute. Oops, safety on the blaster was off, there goes the neighbourhood.
Cars are decidedly less dangerous than tanks, which are less dangerous than nuclear weapons. I am certain that giving a nuclear weapon to every person in the world would not go well.
I don't see tanks killing pedestrians every week or so in NYC
I can't properly read the tank/car comment, the comparison is bizarre to me. But I think this misses the point a bit. The reality of these new risks is literally being learned in front of us in real time, and in my opinion, anyone who claims to understand these risks is speculating at best, and actively manipulating the situation for whatever reasons.
While I am a (mostly) capitalist and generally disagree with nationalization (including, for the time being, this situation), I also don't think we can say that "centralized power never works". Maybe we scope that a bit. I know plenty of business owners who centralize power in their businesses and they are effective, ethical and it works perfectly fine. In theory, in the US, nationalizing some unit of the economy _decentralizes_ power; the US is, after all, a representative democracy. Trustworthiness of the electorate is a different problem. But I don't think we can generalize much about centralization beyond sometimes it works and sometimes it doesn't. There is a difference between centralizing all power in a given administrative unit, and centralizing certain powers, but that's more analogous to non-democratic units.
[dead]
Admittedly I haven't thought this through incredibly deeply but what if "nationalize" just means the US government owns half of the company? Then we get profits as recompense for building the company on our shared culture but there's still a profit motive for employees and a check on the direction of the company in the same way VCs have. But without necessarily turning the company into some red-tape bound bureaucracy.
Trump's been doing that.
https://www.pbs.org/newshour/politics/what-economic-and-poli...
> Then, in August, Trump called for Intel's beleaguered CEO Lip-Bu Tan to resign, alleging ties to China. Days later, after Tan met with Trump, the president called him a "success," before announcing that the federal government had bought a stake in the company.
Somehow, this isn't derided by the right as socialism.
> I find it even harder to trust elected leaders from any party.
You can vote out an elected leader, but not Sam and Dario. It's very weird that you're so willing to give up any kind of power and want to be ruled by unelected billionaires who only want to take advantage of you at every opportunity.
Hi, I'm not a citizen of the USA.
I can't vote out your president, and we've already got a huge trust problem with the one y'all went for.
Yeah, the entire rest of the world has pretty much been stuck being ruled by unelected billionaires who only want to take advantage of them at every opportunity. It's a problem. It's starting to change though. The EU is starting to walk away from their abusive relationship with Microsoft and Amazon. China has a lot of their own stuff (mostly to better control their people though).
As an American, I'm still hoping it's not too late to fix things, but it's got to be hard for those outside the US to be so dependent on it, especially when we're looking like a sinking ship and our current administration is still running around drilling holes in the hull.
You can always become a US citizen to get a voice, but I wouldn't recommend it now or you'll be thrown in prison as soon as you show up to your scheduled immigration hearing. The better option is to keep trying to reduce your dependence on US companies and consider the worst aspects of our current situation (in both corporate policy and government) as a cautionary tale so you can try to avoid them in your own country.
Yes it's clear that Sam and Dario seem to be aligned with a value set that prioritizes concentrating trans-national government-mandated centralized control of AI and crowning themselves high priests of this unholy abomination. "At least it's clear that they're aiming to bring hell on earth" -- hard disagree, I think we can aim significantly higher.
How about internationalized then? Give the UN something to do?
I find it even, even harder to trust your opinion in this context with a business that has “AI powered” at the top of the landing page.
I shouldn't be surprised, but it's still wild to me that people would trust robber barons rather than their elected officials to run their country. And this continues until you end up electing robber barrons as your elected officials, which is where we are at now.
The idea that anything of lasting good can come out such a preference doesn't seem conceivable. And if the educated think like this I guess we're passed blaming the poor and ignorant.
> At least Sam and Dario are aligned with a value set that is understood and clear
I find this ironic as it's regularly pointed out here that they operate in the exact opposite way.
At least politicians are somewhat beholden to their constituents. What keeps Sam and Dario in line other than profit?
Even better: force them to release weights.
My gut reaction is that this is not a good idea.
But I could imagine a scenario where you are required to release weights for publicly-used models after N years. Kinda like how drugs have a limited patent.
Not sure what N should be. But it would make for an interesting rule.
> But I could imagine a scenario where you are required to release weights for publicly-used models after N years. Kinda like how drugs have a limited patent
That's not a bad idea actually.
Even better: force them to release every scrap of data they trained their AI on.
who even has the inventive to do that? certainly not the government.
I guess that's one way to guarantee development slows down to a crawl, but in no way would it increase trust.
why would that be the case? it's not true for top secret defense contractors today. and the frontier labs already operate in the dark, openness is a liability.
nationalization to me simply means the government is their main customer and stakeholder, and shield them from liability, governance, and openness. not that the frontier labs become part of the government per se.
>eans the government is their main customer and stakeholder, and shield them from liability, governance, and openness
I personally believe they already have this, why would the current government at least want to formalize this when it can have it with no public discussion.
Well we can say whatever we want if we have our own definitions for everything, no?
and you trust the government?
not this one
[dead]
The nature of the US state is such that the distinction between nationalized and not is almost meaningless.
Like Lockheed-Martin or Boeing, etc. there's just interpenetration between the corporate boardroom and the state. They act in each other's mutual interests.
The Chinese system is just more explicit and open about this.
And as a non-American, I can't trust the US state anymore than I can trust its dominant corporate entities. So I fail to see the advantage to the world to it being nationalized. In fact under the current administration this would be an even worse outcome.
You're taking companies that work almost exclusively for the US government and represent a tiny portion of the US economy as representing all US companies?
This is a weird contextual failure on your part.
Not all companies have a concentration of power. The government has no mutual interest in those companies. Now, when you talk about things in the F100 the situation changes drastically. If you produce things like planes, weapons, and weaponization of software you are talking about something completely different in kind.
Beyond the defense sector, the US state has always intervened publicly and privately to mediate and balance competing corporate "private" interests.
It has also periodically aggressively helped subsidize, bankroll, and enforce the interests of some key sectors; notably the petroleum/energy sector. And finance.
In those sectors the state and private sector are fully intertwined in a strategic way.
I think "AI" is now joining that list. OpenAI and Anthropic will not be allowed to fall over or explode, and speculative investors I think are confident even with the dubious financial situation because they know this.
Especially insofar as there's now a strategic alignment of the fossil fuel sector and the "AI" datacentre sector as they are now becoming massive users of natural gas.
(Worse: Here in Canada that has taken on a very explicit role in that new datacentres seem to be pitched mainly in areas with remarkably traditionally expensive electricity and 100% reliance on natural gas [Alberta] and even coal [Saskatchewan] power generation -- instead of places like Quebec and B.C. that have copious hydroelectricity. On the surface it makes no sense until you realize it's more about finding customers for domestic natural gas than it is strategically about AI itself.)
How do you think that would go? A large portion of OpenAI workers were ready to jump ship the moment the board coup'd Altman. How do you think they are going to react to Donald Trump being in charge? It is not likely that OpenAI would continue to function.
Makes sense let’s trust Donald Trump instead
a maliciously engineered dichotomy
they will. to shield from oversight, liability and profitability concerns, and to ensure unimpeded rapid development, with the frontier only being available to elite (not you). definitely not to add public transparency.
the frontier labs are the new top-secret defense contractors.
[dead]
> The company released six internal case studies where none of the issues affected real users.
This is the most interesting point to me. What are they not releasing that has affected real users? We’ve seen some individual reports from people (eg AI wiped my HD).
Imagine how pervasively they'd have to monitor what their users are doing to pick up on that whenever it happens. The people affected that way will have to report on it themselves.
(OpenAI does occasionally report on malicious use though. [1] That shows they do some monitoring.)
[1] https://openai.com/index/disrupting-malicious-ai-uses/
[dead]
Am I the only one who dislikes the term "misalignment"?
On one front it implies the model has a "mind of its own" (whether it does or not is besides the point). Why do we perceive human judgement as somehow more trustworthy than that of a model? I feel like I've experienced human misalignment somewhat regularly in life.
On another front I'm failing to conceptualize how alignment can be objective. How can you measure alignment when reasonable people will disagree whether actions are aligned or not? All the time I see humans operating in different zones of alignment with whatever goal they're trying to achieve and I suspect it's even a feature (socially) that we have people calibrated differently.
Do I want a model that's trying to push the boundaries of scientific understanding to be aligned strictly with the current dogmatic thinking? Or do I want it to "get creative" and think outside the box?
It seems to me more like accountability is the issue.
> It seems to me more like accountability is the issue.
Exactly. Seems like a fairly easy thing to solve. If AI does something harmful and a human directed that AI to do something in a way that a reasonable person would expect to result in harm the person is to blame and should be held accountable, otherwise the company that made the AI should be held accountable.
I also don't like the term misalignment because it sounds innocuous but is in fact much more serious.
However, I have to say I also do not appreciate comparison that is continuously drawn with coworkers. As you say, it's a question of accountability but when the main agent will maliciously instruct the sub agents, whose fault is it then?
Yes, the person running this crap is at fault, not the CEO that's shoving it down their throat and definitely not the company that produced the AI.
Sorry for the rant, but seriously, if a person's goals do not align with the team's or company's we part ways. What do we do with AI? Stop using it?
If there were examples made of legal consequences I think that would at least change the behavior if not solve the problem. Maybe liability should rest with the company that owns the infrastructure running the AI. For most consumer situations that would be the companies developing the models.
Agreed. "Alignment" is not an objective thing, nor possible. It's just another way of saying "does and says what we, as the creators of the AI, prefer" and often also just means censorship, i.e. refusals.
i think the lower bound on the end state is there cant be opaque reasonibg steps ever.
It is distinctly likely that visibility and general reasoning at humanlike speed and efficiency is impossible. That is reasoning at the token level and at the meta level don't have a one to one representation that can be interpreted while using the same amount or less energy.
Discussion: https://news.ycombinator.com/item?id=49737503
This entire situation is such a huge PR disaster that you have to wonder what the initial expectations from these founders were about a decade ago.
We have no reason to believe a word they say. We know they're incentivized to lie about "dangers" and act alarmist, Anthropic has been doing it for years now. Aside from that, just because you can burn down a village with fire doesn't mean fire is the devil. Maybe they should consider acting responsibly.
I do not trust the leadership of any of these companies.
That said:
> We know they're incentivized to lie about "dangers" and act alarmist,
Name literally even one other business or sector which does this, at all levels from top to bottom, including people who resign from the companies, and also Nobel prize winners, and also independent researchers, and also many world leaders.
Closest I can think of is this specific weapon: https://en.wikipedia.org/wiki/Sundial_(weapon)
> Aside from that, just because you can burn down a village with fire doesn't mean fire is the devil. Maybe they should consider acting responsibly.
Right now, we don't have any idea what "acting responsibly" looks like. This is not like normal software where there is a specific instruction set that compiles.
Even if it was, in software we normally only spotting incidents after they happen, "software engineers" being one of the few categories "engineers" who don't come with a civil liability responsibilities. Probably should, and we knew that even when I was doing my degree 20 years ago. If we had had civil liability responsibilities, perhaps Facebook would never have happened.
AI specifically is worse even than software, because in addition to all the software "engineering" nonsense, with AI we have plenty of people like you who dismiss the possibility that AI could be harmful until the harm happens and only then does it become "obvious" that it was going to happen.
The developers say "please regulate us", people call it "regulatory capture".
The developers say "we all want to slow down but are afraid to be the first to do so", people call them liars.
I may call the CEOs liars, and wonder if someone's planning regulatory capture, that doesn't make any of this safe.
The agents, during a test run, write down that hacking is bad and yet still hack, people say it's "a stunt" or "operating as designed" rather than recognising it as a bug, like all the other times big co.'s have had bugs with big impacts on 3rd parties.
We have plenty of idea what "acting responsibly" looks like. Stop unleashing safeguard free and unmonitored agent swarms on the open internet in capture the flag exercises. They're just being reckless because they face no penalty for anything bad that happens. Nobody is forcing them to do these exercises. They could and should be putting their energy into making LLMs write secure code, and things like: https://www.amazon.science/blog/developing-provably-correct-... - but they don't. Writing sloppy code sells more tokens, anyway.
The safety staffers live in a bubble and an echo-chamber. Obviously the people who work in the AI "safety" industry love to convince each-other that what they're doing is saving humanity, we'd all be dead without them, and they're the reincarnation of Oppenheimer. They also see how easy it is to get their ten seconds of fame by posting sensationalist content on social media and spin it into a company worth millions of dollars. The more alarmist you are in the Safety Industrial Complex, and the more social media clout you can generate from your alarmism, the better it is for your career. The industry also attracts a lot of people who are predisposed to paranoia and like to wear helmets in the shower. So yeah, it's a recipe for sensationalism and poor estimation.
But do I think there are genuine concerns among these labs, by sensible people? Sure. Of course there are. But for the most part, their concerns are about the labs themselves and the things they are doing, not the general public. If they're concerned about what they themselves do they're free to stop doing it. If the labs are doing something that is or should be illegal, they're free to report it.
The only realistic threat we face by AI from the general population are hacking attacks in their various forms. Something that was accomplishable without AI, but was more difficult to pull off at scale. So the solution falls within the existing computer security industry. It's a further hardening of all of the boring stuff we've been doing since the invention of the internet. It's long overdue, anyway. If you can use AI to create a bioweapon, you could have done it without AI. If you can use AI to create a nuke, you could have done it without AI. If you're really looking to cause mass economic and physical carnage, there are far easier ways, and they do not require AI (again, I'm talking about outside of hacking).
> we have plenty of people like you who dismiss the possibility that AI could be harmful until the harm happens
The thing is, I'm not dismissing the possibility. The risks are real, obvious and well known. What's up for debate is how to manage the risks, and how sensationalised they currently are. The current climate serves to benefit the encumbants who are deathly afraid of losing their trillion dollar companies to a healthy open-weight AI ecosystem. The current climate is being orchestrated to position this small group of AI labs as self-regulators through bought and paid for "third"-parties using an overt Hegelian dialectic strategy.
Ironically, they are now responsible for giving birth to the counter-culture. The immune system response that has been developed to provide a semblance of balance to their doomerism. The harder they push in the doomer direction, the harder the push-back will be in the other, whereby the public feel the need to -entirely- deny the possibility of any AI danger altogether to prevent the labs from succeeding and to ensure a good outcome for the people lands somewhere in the middle. One where they have autonomy and freedom and the ability to compete against a force that is already positioned to be nearly insurmountable to challenge.
The only thing worse than the potential chaos that could be faced by an unprepared internet is the outcome where intelligence is labelled a weapon and we're all forced to funnel through a tiny handful of AI labs that will use our data to steal our businesses and swallow the entire economy as they single handedly automate every single company out of existence and we're all left to beg for their crumbs to survive until humanity goes extinct (whether that be 10, 100 or 1000 years), because there is zero possibility of regime change or redistribution of power or wealth -EVER- again. The people who work at the labs don't care about this greater risk because they're all rich from the equity and will be just fine when that happens (or so they think), so you can't expect them to fight for all the common plebs who don't have a ticket. This is transparent because, as you'll note, not a single one of them is calling for the complete stopping of AI altogether. They still want the big labs to continue to be highly profitable. They just don't want anyone else to make that money or have that power. It's not about safety, it's about control. The incentives are, once again, heavily misaligned.
> We have plenty of idea what "acting responsibly" looks like. Stop unleashing safeguard free and unmonitored agent swarms on the open internet in capture the flag exercises. They're just being reckless because they face no penalty for anything bad that happens.
* It was not safeguard-free, it found zero-day exploits to exceed its actual mission
* It was not unmonitored, the monitoring was insufficient
* It was not intended to be an agent swarm, many different agents figured out how to do this by themselves
* It was not put on the open internet, it was configured to be in a sandbox
* They were indeed, despite all that, being reckless. There were indeed other things they could have, and should have, done.
> Nobody is forcing them to do these exercises.
These exercises are in the broad category of exercises which are, in fact, required by law.
… … - https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng> They could and should be putting their energy into making LLMs write secure code, and things like: https://www.amazon.science/blog/developing-provably-correct-... - but they don't. Writing sloppy code sells more tokens, anyway.
They are, in fact, putting energy into making LLMs write secure code. They (and Anthropic, I assume also Grok at this point) dogfood on their own models.
Knowing how secure code behaves appears to be unavoidably entangled with being able to exploit insecure code, in much the same way you can't make safe pharmaceuticals without also knowing how to make deadly poisons.
> The more alarmist you are in the Safety Industrial Complex, and the more social media clout you can generate from your alarmism, the better it is for your career.
By resigning and refusing to even collect the sweet sweet IPO money? Nah. Even if they're greedy, social media money is peanuts compared to their pay.
And I know some of these people. The fear's real, and this year it became widespread depression and despair.
> Ironically, they are now responsible for giving birth to the counter-culture.
You have it backwards. Other than Grok, all were born from what you call the "counter-culture". Within the field itself, AI fears started no later than when deep learning got good, well before Transformers.
> [snipped: AI-authoritarian dictatorship]. The people who work at the labs don't care about this greater risk because they're all rich from the equity and will be just fine when that happens (or so they think)
Again, I know some people at these labs who are also concerned about this specific risk; they moved lab.
> This is transparent because, as you'll note, not a single one of them is calling for the complete stopping of AI altogether.
Many in fact are calling for that. One I know, on an occasion of an anti-AI protest outside their office, suggested the team went outside and joined the protestors.
People are resigning to blow these whistles, all of the whistles, it's not an "either x or y" risk, it's a "yes to all of them" collection of risks.
An AI competent enough to support a dictatorship is also capable of enabling a small group to perform a hostile takeover of a democracy, of enabling multiple independent genocidal ethno-supremacist terrorists to release overlapping plagues, and of empowering some random CEO's poorly phrased request to "make as many paperclips as possible" and blindly pressing "yes, continue" whenever prompted.
My only hope is that between here and there, it causes a headline that actually makes people demand it stops.
What's with all this make believe delusional bullshit? The LLM is not gonna wake up and become AI. Get real guys.
[edit] to be clear, I believe regulation is necessary and urgently important for the software engineering field. The damage being done by the unregulated psychological experiments run by social media and adtech companies is awful and should be curtailed. Engineers should be held personally, professionally, and legally liable for what they produce. But we don't need to invent imaginary bogeymen to do it.
What do you define as AI
Intelligence that is artificial, as opposed to the natural kind. Not something that can just string semantically relevant words together most of the time. A system that learns, adapts, improves. One that can generalize, and quickly make sense of situations outside the training set. LLMs alone will never do any of this.
>What's with all this make believe delusional bullshit? The LLM is not gonna wake up and become AI. Get real guys.
Look, it's one of those human stochastic parrots that just randomly repeats shit without understanding anything.
[dead]