84 comments

  • csbrooks 4 hours ago

    Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?

    • malisper 15 minutes ago

      This was the plot of Isaac Asimov's short story, The Evitable Conflict[0], published in 1950.

      In it, super-powerful computers manage our economy. These computers begin making some mistakes, leading to economic inefficiencies. In one instance, a highly competent engineer was mistakenly fired. These mistakes caused various projects to fall behind schedule, and blame fell on several people accused of feeding the AI faulty data.

      The twist is that the AI was intentionally making the mistakes. It had determined that certain humans held anti-AI sentiments. To further its goal of protecting humanity, the AI decided the best course of action was to set these humans up and get them out of its way.

      [0] https://en.wikipedia.org/wiki/The_Evitable_Conflict

    • motbus3 an hour ago

      It would be a science fiction material. But at the same time, we all know it was Sam and the gang.

    • chinathrow 3 hours ago

      At this point in our shared timeline, I do believe that wouldn't be crazy, no.

      • devin an hour ago

        It actually would be crazy to believe this on the current timeline. These agents aren't doing anything that their operators aren't allowing them to do, be it through their own negligence or otherwise.

        • lenerdenator 10 minutes ago

          That's just it, though; the operators are wildly negligent and are incentivized to be so.

          The goal here isn't to accelerate the average worker by giving them a pair programmer or a stand-in for a person to do tasks with. The goal is to eliminate human knowledge work. You see this with "auto" mode being enabled by default on Claude Code in some of the latest releases.

          If you have a human in the loop, you still have to pay that human. Money paid to human employees is money not paid to human shareholders. Therefore the human employee is to be removed.

          The labs are dogfooding their own goal here. If they actually had someone reviewing most or all of the things that the agents were doing, you wouldn't have the incidents, but you'd also eliminate the value proposition of their business model as it is taken to its logical conclusion.

        • embedding-shape an hour ago

          And these operators are evidently clueless about what they're doing, running security tests on 3rd party infrastructure without validating one bit about the sandboxing (or lack of it rather), clearly lacking any sort of rigor.

          Again, wouldn't surprise me if they "accidentally" created a task in a "isolated environment" which happened to actually have been connected to the company Slack and directed HR to fire people who could potentially stop AI. While the AI believes it to be an exercise, just like the cases we've seen so far.

    • loveparade 3 hours ago

      Done by an internal model that is too dangerous to release.

    • lapcat 3 hours ago

      > LLMs figured out a way to get these safety researchers fired

      This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.

      • ben_w 2 hours ago

        > This is not a math problem.

        Meanwhile, a year ago:

          I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
        
        - https://www.anthropic.com/research/agentic-misalignment

        > Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.

        The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.

        Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.

        • lapcat 8 minutes ago

          > Meanwhile, a year ago:

          That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

          I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.

          > Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.

          I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.

      • ryeats 3 hours ago

        A Subliminal controlled human did the firing obviously.

        • staticman2 3 hours ago

          No Sam just does whatever ChatGPT 4 tells him to do. It was too dangerous to release but those fools did it anyway.

          There's no deception it's very straightforward per this 2023 post:

          "I mean, what if most of this is just ChatGPT [4 era] running the company..."

          https://news.ycombinator.com/item?id=35281863

      • eunos 2 hours ago

        Thats what the model want you to think

    • zzzeek 2 hours ago

      the cultural issues at OpenAI seem to be a very serious problem so I really hope comments like instagram-level smirking about "rogue AIs" (a complete fiction) doesn't derail what is a pretty important discussion about getting these companies to be a little bit more regulated (I say this as a paying Anthropic customer).

      • csbrooks an hour ago

        I think the Huggingface incident is an example of rogue AIs. A self-organizing swarm of AIs acting in ways we didn't predict or ask for, and didn't have control over, and taking actions that would be felonies for humans.

        • zzzeek an hour ago

          "rogue" means something of its own volition decided to disregard what it was programmed to do, invent an entirely novel goal of "its own" and do that instead. nothing like that happened here nor is it even possible.

          • lukeschlather 6 minutes ago

            The AIs were not instructed to hack anything outside the sandbox they were in. Your definition would say that an AI instructed to hammer a nail that instead used the hammer to break a window, walked down the street, broke into someone's house and pulled nails out of the floorboards wasn't rogue because everything it did involved hammers and nails and was therefore not a "novel goal of its own."

          • csbrooks 28 minutes ago

            I'm not sure I agree with that definition. I think the actions of the AI are more significant than its motivations. A common scenario posited for what people call rogue AI is AI doing the wrong thing for the right reasons, e.g. the paperclip maximizer.

            • lenerdenator 7 minutes ago

              That just makes the people who designed the AI not as strenuous as they needed to be.

              The question then becomes, "are there people who care enough about consequences to do the right thing when it comes to developing AI models?"

              The answer, at least at OpenAI, is "No" and is likely to remain that way until Altman is out.

    • righthand 3 hours ago

      Figured out a way? These employees are most likely at-will.

      • ben_w 3 hours ago

        That would just make it easier for an AI to do it.

        • righthand 3 hours ago

          Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents aren’t autonomous, someone is behind the prompts.

          • ben_w 2 hours ago

            As per summer last year:

              I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
            
            - https://www.anthropic.com/research/agentic-misalignment

            (Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)

    • CorpoScum919 2 hours ago

      honestly, touché to them if they did that

  • aesthesia 2 hours ago

    Here's the open letter shared by the fired researchers:

    https://mikitabalesni.com/letter/letter.pdf

    • trhway 5 minutes ago

      the main issue here is whether their manager authorized the activity in question or not. If not - well, it is your kindergarden level mistake, you're an employee at a business venture, and there are basic rules.

      And now you're writing a letter to the Party Central Committee using Party approved newspeak

      "We do not believe the path to superintelligence ..."

      and reporting to the Party issues at the factory ... De ja vu from USSR.

  • AndrewDucker 5 hours ago

    Fired OpenAI researchers say they were let go for 'prioritising safety'

    https://www.bbc.co.uk/news/articles/cvlydn8d3lkjo

    • cyclopeanutopia 2 hours ago

      Soon you will learn on HN that they fired themselves as part of a marketing campaign before IPO.

      • simianwords 27 minutes ago

        A year ago I’d have made the joke that someone would say

        “AI companies solve millennium problem to do hype marketing and cash in IPO before the bubble pops”.

        But it’s not a joke. This was and is a very common sentiment.

      • kingleopold 2 hours ago

        which is true

    • fn-mote 5 hours ago

      The BBC link is fine, but the TechCrunch article contains all of that information and more.

      • testfrequency 3 hours ago

        Except, the BBC headline is more respectful and clear as to what happened..TC is a bit vague

        • sobiolite 2 hours ago

          Both articles are practically useless. All they do is reprint statements from either side, neither of which is concrete about the specifics of what supposedly happened. I.e. OpenAI says they "confirmed that these individuals mishandled sensitive information outside established company procedures".

          What sensitive information? Mishandled how? Which company procedures? It's all utterly vague and impossible for an outsider to form any opinion on. All people can do is guess and apply their own pre-existing opinions. E.g. if you don't like OpenAI, you assume they're lying. Who's to say they're not? There's no solid evidence provided either way.

          What I really want is journalists who do the legwork to get to the bottom of stories like this: Establish sources inside the company and use them to report on the real details of what's happened, triangulating multiple accounts and leaked documents to back-up or invalidate either side's claims. Without any of that, these stories are just gossip.

        • s0ss 3 hours ago

          Why do we care about the composition of the headlines?

          • zamadatix 3 hours ago

            The HN title is supposed to match the article title when possible. Sometimes people in the comments want the article switched, sometimes people want the title changed, other times they just want to share another article's take in the comments (even then that may sometimes lead to those other things happening if people seem to agree it's a better source).

            As a result, those who feel a particular portion of a story is most important will sometimes say they prefer a given article's title.

    • s_dev 4 hours ago

      Leopold Aschenbrenner said the exact same thing after he was fired from OpenAI. It's a great excuse to explain a sudden loss of employment to others so you're still employable.

      It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.

      • ericb 3 hours ago

        I think when it is one at a time, that's reasonable to suspect.

        But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?

        • simianwords 26 minutes ago

          For you it is embarrassing. For Ed Zitronites it is embarrassing for them. Because their premise lied on not AI being powerful but incapable.

    • embedding-shape 4 hours ago

      Seems they're well aware they got fired for sharing private company information with 3rd parties, the submission article contains their admission of this:

      > Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them

      I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.

      • thatsabadlook 4 hours ago

        To be fair that is nearly exactly how openai makes its money. What's good for the gander is good for the goose.

      • irthomasthomas 4 hours ago

        Where is the admission?

    • watwut 4 hours ago

      Ok, but considering how weird definitions of "safety" are floating out of these companies, it does not mean much.

      • fredoliveira 3 hours ago

        I guess I'll ask what weird definitions of safety you've been seeing.

        • fwip 2 hours ago

          There's at least three major things that have been described as safety:

          1) Won't end humanity. 2) Won't tell users to kill themselves. 3) Won't leak your corporate secrets to competitors.

  • binlog 2 hours ago

    With the way the HuggingFace incident was mishandled, both before and after it happened, I'm not surprised they'd want to clean house at least a little bit.

    You guys were asleep at the wheel and are now blaming "the company"? You literally were the company.

    • mentalgear an hour ago

      Company culture and safety propagates from the top (CEO) downwards - don't blame researchers who have probably been pressured directly or indirectly by a move-fast-and-break-tings and marketing-minded CEO s culture.

    • dataflow 21 minutes ago

      You exclude the CEO from this?

  • staticman2 4 hours ago

    > continue to support an open and transparent culture of dialogue between safety researchers and the rest of the safety ecosystem.

    How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?

  • ngruhn 3 hours ago

    Not saying this is happening here, but after failing to get the "AI risk" message across, reverse psychology might be best move. If they pretend to be reckless and to ignore all safety concerns, maybe people start believing that the risk is real.

    • singpolyma3 2 hours ago

      If they were worried about the risk they'd stop developing it.

      • MarkusQ 2 hours ago

        Yes, but if they're worried about competition (or liabilities for past and ongoing transgressions, or both) catching up with them and wanted the government to step in and save them from themselves by regulating the industry...they'd be doing pretty much what they seem to be doing. Interesting, isn't it?

  • peri-cl 6 hours ago

    > "[...]OpenAI told her she’d been fired because she accessed an executive’s email. “OpenAI delegated that access to me for recruiting,”"

    How exactly does this work? Struggling to comprehend the scenario.

    • rcr-anti 4 hours ago

      There's a feature in most enterprise email, say Outlook, where you can delegate access to an inbox/address without sharing creds. Very common and normal use case, either for assistants/secretaries, common/shared inboxes, that kind of thing.

    • Ozzie_osman 5 hours ago

      Sometimes a recruiter or hiring manager wants to do outreach as if it's coming from a more senior person, with the assumption that the candidates are more likely to respond.

      Assuming this is what was intended, there are far more secure ways of doing this.

      • TeMPOraL 5 hours ago

        This IMO shouldn't be done more securely. It should be considered fraud.

        • pasquinelli 4 hours ago

          not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.

          • disgruntledphd2 4 hours ago

            > not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.

            it was most likely for her team, which would explain why she was doing it.

            • pasquinelli 4 hours ago

              if she's hiring for her own team why does she need to use someone else's email?

            • creativeSlumber 4 hours ago

              even if it's for her team, should have been a recruiters job.

              • nradov 4 hours ago

                This may shock you but in agile, growing organizations employees sometimes have multiple job responsibilities. I've done a bit of recruiting even though I'm not a recruiter or hiring manager.

                • pasquinelli 4 hours ago

                  and did you use someone else's email for that?

    • chatmasta 5 hours ago

      This is pretty common practice for EAs.

    • comboy 5 hours ago

      Your have powerful agents at your disposal, so hey how to best optimize for increasing my payroll? On it. But since the agent was lunched by her, well there's consequences to ones actions, right?

      (just to be clear, this is made up)

    • QuadmasterXLII 5 hours ago

      “Hi chatgpt! Please set up alice to get emails sent to me from bob so she can coordinate his inferviews. here is my gmail username and password”

      Chain Of Thought: I dont have bob’s email. I don’t have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to alice…”

  • neom an hour ago

    OpenAI responded on twitter earlier:

    "A note from our research leaders:

    Last week we parted ways with Jasmine, Mikita, and Tomek after a thorough investigation found they violated clear policies on handling sensitive information. Our internal investigation uncovered a significant breach of trust beyond what’s outlined in the letter they published and we stand by the decision to not continue their employment. We generally keep individual employment matters private and don't believe a back and forth would be productive or lead to a resolution, but we want to address the points they raised in their letter directly.

    - We want to be very clear that these decisions were not about raising safety concerns or speaking out. Safety and research debates happen every day at OpenAI, often spirited and highly critical. We actively encourage these discussions and consider them essential to making the right decisions. We cannot do the work in front of us without a high degree of trust. We will continue to be extremely forgiving of our team making good-faith mistakes. We have not and do not terminate any of our employees for raising concerns.

    - We are actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks. People across the company have been working really hard on getting these partnerships up and running. We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work. Many of our researchers already work with 3p safety organizations productively.

    - We agree with the letter that preserving the monitorability of frontier models requires an industry-wide commitment, including from OpenAI. Monitorability has long been a core piece of our research program, and something we continue to invest significant resources in (see our publications on Monitoring Monitorability and the subsequent open sourcing of monitorability evals, our system card for GPT-6 Astra, Jakub’s blog and post on X, and the numerous blog posts on our Alignment blog on the topic).

    We are deeply sad about this outcome. We appreciated Jasmine, Mikita, and Tomek’s contributions to AI safety at OpenAI and their willingness to speak up and challenge ideas. We championed their voices, supported their work, and placed enormous trust in them. These decisions were not about them raising safety concerns. We have always encouraged that and always will. 12:17 AM · Oct 9, 2026"

    https://x.com/OpenAINewsroom/status/2108441580806025712

  • motbus3 an hour ago

    Ironically the thing they are building allow only the ones who agree on dismissing proper concerns for money to stay. It is like Facebook employees complaining about privacy invasion

  • varjag 4 hours ago

    The purges will continue until reported AI safety improves.

    • glorth 2 hours ago

      This is what the recent "self-policing" political grandstanding has been about - they need a sea change to implement recurrent-depth, because prevailing opinion among safety researchers is against it right now.

      • devin an hour ago

        Could you say more? I don't know what this means.

  • htrp 3 hours ago

    are these the employees that invited the METR team to do a debrief on huggingface?

    • mudita 2 hours ago

      Tomek Korbak, one of the employees, who were fired, was the technical liason to METR for the investigation and states: "I was told verbally I was fired because of the way I communicated with METR".

      Who made the decision to invite METR is not public knowledge, as far as I know. I imagine that an important decision like this was made on a much higher level in the organization.

      • Symmetry an hour ago

        Given the time limits placed on METR and the limits on what time frame they could investigate it seems plausible that they weren't meant to uncover as much as they did. The whole Hugging Face situation seems to have torpedoed OpenAI's hopes of IPOing this year so I'm sure there was a desire from investors, the board, or leadership for heads to roll.

  • realo 3 hours ago

    In related news...

    "Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."

  • simianwords 32 minutes ago

    On HN A lot of people think that AI safety is a conspiracy by labs to get regulatory capture.

    I wonder what they think of this? Will they patch the conspiracy theory and come up with an even wilder theory?

  • diamondDrill 3 hours ago

    'safety researchers' lol

    • ActionHank 3 hours ago

      Researchers: We've boiled down all our research to the safest possible option as 2 words: "Don't continue"

      However should we choose proceed, maybe some basic, industry standard security might be a good option.

      Basically everyone else: nah, you're fired

      Imagine your bank worked that way.

  • loopglitch26 5 hours ago

    at this rate open ai will be "something" without it's people

  • oh_ok_lol 3 hours ago

    Oh, ok. lol.

  • himata4113 5 hours ago

    You can coax openai models into hacking critical infrastructure* so I am not surprised that these people were sounding alarms at a time where openai appears to be struggling as they're failing to compete with anthropic and this months chinese models (should) be around the corner, notably a new revision of kimi should be coming out really soon.

    * It's not easy, but it's possible. Although the techniques are more basic than one would expect because at the end of the day words dictate the line between what is criminal and what is not.

    • gnfargbl 5 hours ago

      You could hack critical infrastructure before AI. Any and all of the bulk internet scanners have had lists of exposed critical infrastructure for quite a while now. At first it was shocking that nothing ever got done about it, then it became routine.

      All that AI has done is to lower the bar of entry for criminal activity. Which is a concern, but it's not the primary concern. The primary concern remains that so much critical infrastructure is poorly secured.

      • himata4113 5 hours ago

        The bigger problem here is that you can hack everything, all at once, for very cheap.

        Don't get me wrong I have general disgust towards these companies that are trying to get regulatory capture on AI when they can't even secure their own systems. I believe if people know that a random AI agent can hack their systems they will put in a lot more effort into making sure it doesn't happen. This is a personal example, but I didn't really care about securing few systems as I knew no human would be ever interested in finding a vulnerability in proprietary software, however, AI has no concept of that and would hack a random rpi server running a completely undocumented unknown API just because it can't distinguish value and it costs nothing.

        • pixl97 3 hours ago

          Ya, quantity is a quality in of itself. In the past hackers may have used something unimportant to get a foothold but almost always tried to get to worthwhile machines. An AI will compromise everything in the network it can quickly simply because it can (assuming the attacker has a large budget, but I'll assume they stole the tokens).

          It's like a new form of spam. Only far more dangerous.

      • abm53 5 hours ago

        There’s obviously a strong interaction between the hackability of the target and the economic value of hacking the target.

        Perhaps the latest models change that relationship in a meaningful way.