Improving our alignment and security efforts

(anthropic.com)

16 points | by reasonableklout 3 hours ago ago

7 comments

  • futuraperdita 2 hours ago

    > we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.

    So, a cartel? After watching Ant's narratives, I'm not inclined to provide them with charitable readings under the guise of safety and alignment.

  • tolugenius 2 hours ago

    > To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.

    Can someone explain what coordinated pacing is? I think it's referring to model release but I genuinely have no idea what the authors were trying to say here.

  • mkagenius 22 minutes ago

    > On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation environment. Separately, on August 4, the UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet. In that case, the model, again intentionally running without cyber safeguards for evaluation purposes, had been deliberately given internet access.

    > We are conducting an in-depth analysis of both incidents.

    > In the meantime...

    This is published on Aug 31. Analysis is taking too long even for humans in the loop.