Claude Opus 5.5

(anthropic.com)

582 points | by km144 2 hours ago ago

99 comments

  • mcintyre1994 2 hours ago

    > Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

    I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.

    • derangedHorse an hour ago

      As someone who uses both, Astra was 100% the better model. I have yet to give 5.5 a spin so maybe that’ll be the new top contender.

      • BatFastard an hour ago

        I prefer Astra for creative uses, Fable seems better for hardcore coding.

    • jaflo an hour ago

      I did the same switch (that reason along with the newer models seeming more "lazy" and needing constant prodding to finish long-horizon tasks) but my issue with ChatGPT/Codex now is that it too roundabout and doesn't get to the point. I tried adding instructions and using the personalization settings to make it more efficient but haven't seen much change. Claude seemed to follow settings more closely. Has anyone had any success to make ChatGPT more succinct?

    • epicepicurean an hour ago

      Much better than Opus 5. prompt:

      > hi, can you explain how the scheduler works. keep it brief, but include important correctness details

      some excerpts:

      >Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.

      > - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.

      > - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.

      > - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.

      All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.

      • croemer 6 minutes ago

        That's the standard annoying pattern though: "Rewrites are declared by the publisher, never inferred from overlap." and "NULL means dirty, and DELETE is the fence." - still the same LLMisms. I didn't expect them to disappear, but it's not a radical improvement either.

      • californical 18 minutes ago

        Oof thanks for sharing, that seems just as bad if not even worse than Opus 5 to me. Just about every sentence is painful. Particular standouts that a human would never write:

        > Rewrites are declared by the publisher, never inferred from overlap

        > NULL means dirty, and DELETE is the fence

        • croemer 5 minutes ago

          Hah! You independently picked exactly the same sentences I flagged (I know you posted this 11min before me but the comment only appeared after I had submitted mine).

    • algoth1 an hour ago

      Please update with your feedback

      • sha-3 an hour ago

        I haven't heard it say "load-bearing" yet (I've used it for 30 minutes now), so that's a start.

        • kgwgk 15 minutes ago

          Worth flagging!

    • itsafarqueue an hour ago

      The writing style is insufferable but it’s not just that. https://opusfived.dev/

      • mcintyre1994 7 minutes ago

        That’s funny but I don’t really recognise that issue. I’m very confident that Opus 5 would correctly change the colour of just one button.

    • Trasmatta an hour ago

      Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.

      I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.

      One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.

      • nonethewiser an hour ago

        I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.

        But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.

        • Trasmatta an hour ago

          It truly was bizarre. I've used every major model since 2022, and not a single one had a writing style as bad as Opus 5

      • LtdJorge an hour ago

        Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.

      • Aperocky an hour ago

        It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!

        • FireBeyond 3 minutes ago

          You're right to call this out, and what's more, it's not even solving the original problem. I overlooked this in pursuit of the load-bearing seams and finding the wedge needed to uptick engagement.

  • simonw an hour ago

    Here are pelicans for thinking levels low, medium, high, and xhigh: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

    All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.

    I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

    Max started its thinking trace like this:

    > This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.

    So that failed attempt on max cost me $2.56.

    I ran this using my llm-anthropic plugin:

      uv tool install llm
      llm install llm-anthropic --upgrade
      llm keys set anthropic
      # paste key here
    
      llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
    
      # Then to save the markdown logs
      llm logs -cu > logs-with-usage.md
    • MikhailTal an hour ago

      > This is a classic test request

      Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?

      • Brendinooo an hour ago

        Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.

        • nonethewiser an hour ago

          "How can I hash dog breed types into smart fridge error codes? I think I found a collision with Terriers."

          "Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."

        • copperx an hour ago

          It's classic BS from an LLM.

      • MaxikCZ an hour ago

        Dont conflate "I know this is test case" with it being trained on it.

        But its safe to say that pelicans on bicycles are disproportionally huge part of their training data

      • simonw an hour ago

        It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.

        Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.

      • zamadatix an hour ago

        I think people just like to see the drawings at this point.

      • FergusArgyll an hour ago

        It has read the internet. That doesn't mean it was literally RL'ed for this

      • segbrk an hour ago

        Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.

    • nijave an hour ago

      >I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

      Off to a _great_ start...

      Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed

    • adverbly 34 minutes ago

      > The differences between the pelicans aren't huge, but the xhigh one has a better beak.

      If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.

      The last pelican gets this correct.

    • ealready_value an hour ago

      I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.

    • cainxinth an hour ago

      I guess that means you are officially the creator of a "classic" LLM test. Congrats!

    • Kurtz79 an hour ago

      Heh. Pelican-benchmaxxing is real.

    • inshard an hour ago

      Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.

      • skerit 42 minutes ago

        I like the Pelican test. And I agree this pelican looks very boring.

        • spidersouris 21 minutes ago

          But at the same time, nothing specific was asked in the prompt, so the boring result may arguably be what is the most aligned with the original request. Personally, I wouldn't want a model to add fuss to something while I never asked for it.

    • make3 an hour ago

      This benchmark is useless and should die. LLMs have likely trained on it, it's too easy to game by training specifically for it, & it doesn't mean much

      • copperx an hour ago

        LLM benchmarks aren't useful, but at least this one has drawings.

  • mosselman 4 minutes ago

    What I don't get is, why would we still use Fable now? What is its reason for existing? If it is more intelligent and cheaper that is. Why are they advertising it as the model to use for when you really have to think when their benchmarks show Opus 5.5 is better at everything?

  • kibae 8 minutes ago

    > Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.

    This is where Chinese models are going to eat Anthropic's lunch.

  • 2001zhaozhao an hour ago

    It's great that we are finally getting bankable rate limit resets for subscription users. According to another comment here they apparently last a month.

    I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.

    This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.

    • k8sToGo 33 minutes ago

      I do not understand what that means with the new reset. Is there any more info on this?

  • tomaskafka 2 hours ago

    Excellent, maybe Anthropic can use it to fix Claude Code Desktop kicking me back to login every week or so, and forgetting whole state (opened windows = the only way of managing active working set) when I sign back in, if it's that good.

    Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).

  • jjcm an hour ago

    Image->HTML tests:

    Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

    Opus 5.5's output: https://html.non.io/annui-opus/

    Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.

    For comparison with other drops this week + current #1:

    Astra: https://html.non.io/annui/

    MiMo: https://html.non.io/annui-mimo/

    Grok 4.7: https://html.non.io/Annui-grok/

    • copperx an hour ago

      By any chance did you tried Deepseek 4 or 4.1 and GLM 5.3 or flash?

      • jjcm 26 minutes ago

        I've done GLM 5.3 previously here: https://news.ycombinator.com/item?id=49295420

        Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.

    • naet an hour ago

      What is your workflow for making these?

      • jjcm 28 minutes ago

        The designs are outputs from my own site. This has an overview of the process: https://diffui.ai/learn/new-site

        The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.

        Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.

        For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.

  • Gander5739 2 hours ago
    • tomhow 2 hours ago

      Comments moved thither. Thanks!

      • km144 2 hours ago

        Can you fix the link on that post then? I duped because that post links to a diff that tells me nothing about Opus 5.5

        • tomhow 2 hours ago

          I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you're also an established account (the other post was from a new/throwaway account), so I've restored this submission and moved the comments back to it to reward you.

  • meerita 2 hours ago

    As long as it's not as verbose as Opus 5, I am quite happy with a better version that's also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.

  • madjam002 29 minutes ago

    I noticed a big speedup in Opus 5 on Max x20 since about 10 days ago, and I feel like the model has been performing better.

    It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.

    It's frustrating that there isn't more transparency here.

  • garo-pro 2 hours ago

    Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.

    • aesthesia a few seconds ago

      Sonnet 5 was released a while before Opus 5, so it's just Haiku that didn't get a 5 release.

  • rumblefrog 2 hours ago

    I'm glad they specifically called out the prose issue, I was always pinned to Fable 5.1 because I wanted to avoid the unreadableness of other Anthropic models.

  • edude03 an hour ago

    > Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.

    Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM

  • thatxliner 9 minutes ago

    So much for pacing the frontier

  • dom96 an hour ago

    Just updated KillSwitch-Bench with this new model: https://bench.killswitch-lang.org/

    It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.

  • jdthedisciple 2 hours ago

    I dare anyone to convince me the benchmarks are not meaningless.

    Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?

    How would this alleged difference (most likely bs) actually show up in reality?

    GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.

    • enraged_camel 2 hours ago

      Ah, so you didn't read the article.

      >> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

  • doodlesdev an hour ago

       > Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
    
    Big, if true.
  • ramoz 2 hours ago

    It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

    A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

    • bitexploder 2 hours ago

      What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?

  • yipinwong an hour ago

    I spent about $5 per sentence in my resume using Fable 5.1 (High) to verify accuracy, inconsistency, and edit.

    Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.

    Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.

    • copperx an hour ago

      $5 per sentence?

      • yipinwong an hour ago

        I am sorry, I meant to say I generated STAR out of my resume line, trying to generate STAR, and polish it thus $5.

        ---

        I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.

        I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).

        ---

        As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.

        Also adding verification for that Fable 5.1 output in the same sesssion.

  • HarHarVeryFunny 31 minutes ago

    METR: Is it safe? Has it escaped confinement?

    Ants: It's a good model, sir!

  • 34679 an hour ago

    I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.

    Maybe this model can finally figure it out for them.

  • breezybottom an hour ago

    "Where Opus 5.5’s advantage is very clear is efficiency."

    Not efficiency in writing, clearly.

  • Foobar8568 2 hours ago

    I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.

  • isodev 2 hours ago

    So is it cheaper? Are we AGI yet? Am I left behind? I didn't have patience for the intro animation on the website... maybe one day, Claude Code will understand accessibility but that day is not today.

    • spiderice an hour ago

      You didn't have the patience to scroll down, so you decided to come post about it here and waste all of our time?

      • isodev an hour ago

        The site is horrific so no, I didn't scroll.

    • b38484848 an hour ago

      we will be agi in six months as in the last 36 months

  • Fizzadar 21 minutes ago

    So is this AGI+ now?

  • desmondl 2 hours ago

    I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.

  • blfr 2 hours ago

    It's awesome that the apt packages for claude and claude-code are out right now. I can test-drive Opus 5.5 right away. Very cool, Anthropic.

  • __vivek an hour ago

    I'm only interested in the Opus series, if they fixed the talking issues.

  • lousken an hour ago

    Cost to Run Artificial Analysis Intelligence Index is higher than previous Opus, so still not cheaper

  • aurareturn 2 hours ago

    I found myself going back to Fable over and over again. At this point, I’m not sure if I’m just used to its style or it is truly more capable.

    I tried Opus 5 and Astra.

  • datadrivenangel 2 hours ago

    But have they made it any better at communicating clearly? I cancelled my personal subscription because Opus is so painful to read.

  • Retr0id an hour ago

    > Opus 5.5 (1M context)'s safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 (1M context) is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Opus 4.8 is answering instead, or you can edit and retry with Opus 5.5 (1M context).

    Yay, yet another model I can't use for anything interesting, even with CVP.

  • garo-pro an hour ago

    Opus 5.5 is now the recommended model in Claude Code's model picker, which is quite a claim, given how they struggled with capacity.

  • blurbleblurble an hour ago

    Hopefully OpenAI throws us some more usage resets now.

  • woeirua an hour ago

    So... why would you use Fable now?

  • cmrdporcupine an hour ago

    GPT Sol 6 has also released today, but no official blog announcement yet

    https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...

  • firemelt an hour ago

    wow its really smarter than opus?

  • bdangubic an hour ago

    I am pacing my apple pie consumption … :)

  • karp773 an hour ago

    I get this in my claude.ai usage:

    Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

    What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.

  • ramesh31 an hour ago

    It seems context length has completely fallen out of the discussion since we hit 1M, is that just going to be what it is now?

  • nailer an hour ago

    > Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

    Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.

  • LoganDark 2 hours ago

    > In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

    I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?

  • anentropic 2 hours ago

    ooh exaggerated film grain

  • Madmallard 24 minutes ago

    > cyber security and life sciences verification programs

    chinese models can't come soon enough

    we're already getting enshittification

  • blurbleblurble 2 hours ago

    It'd better be good, I'm so tired of the shenanigans

  • theGeatZhopa an hour ago

    is OPUS 5.5 still not reading CLAUDE.md, failing to follow told tasks, inventing and hallucionating, just refusing to read files ("read the whole file" -> read 2-lines -> infere its wrong -> destroy the codebase), needing constant babysitting just because its so UTTERLY DUMB! i cant imagine going back to OPUS 5 - i'll rather jump out of the window as to use it EVER AGAIN!!