Skip to content
HN On Hacker News ↗

Claude Opus 5.5 Should Raise Your Ambitions

▲ 10 points • 5 comments • by 7777777phil • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

2 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,676
PEAK AI % 2% · §1
Analyzed
Sep 26
backend: pangram/v3.3
Segments scanned
1 windows
avg 1676 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,676 words · 1 segments analyzed

Human AI-generated
§1 Human · 2%

When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too. Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again.Feedback is almost universally positive. Claude was never gone, but also is so back. The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations. If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That’s even better.By Claude Opus 5.5, for this postThe pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch.We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.They could have reasonably pitched this as above-Fable-5.1-level performance. Better pitch, but Anthropic tends to keep its pitches conservative. They highlight agentic coding, security and improved communications. The early tester blurbs flag as AI generated and they have a set for each feature area. They praise agentic coding skills, efficiency, readability and communication, and ability to effectively run on its own for extended periods and increased reliability, pitching it as a major upgrade.Benchmarks look fantastic, with occasional spots where Astra or Fable is still ahead.Opus 5.5 is available with zero data retention (ZDR). Given it is at least as capable as Fable 5.1, and clearly more capable on cyber tasks (they say ‘extremely strong’ cyber capabilities), this seems unprincipled. They have a technical explanation but I do not buy it. From where I sit, either Fable 5.1 can have ZDR, or Opus 5.5 can’t. Guardrails are similar to Fable 5.1. Sholto Douglas highlights improved ability to understand and model in 3D, and also a key double-edged sword.Sholto Douglas (Anthropic): also important news we fixed the writingTom Brown: Plus we fixed the accentClaude: Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5. It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.Ado was one of many praising how easy Opus 5.5 is to work with and talk to.Ado (Anthropic): If you loved how Opus 4.6 felt to work with, Opus 5.5 feels like coming home. Same easy back-and-forth with a lot more horsepower underneath. It's the most fun I've had in Claude Code in a while.It's also much cheaper (~40% less than Opus 5 on typical workloads).Why release Opus 5.5 when you have to Pace the Frontier? That’s a whole 0.5 of Opus.Sam Bowman: We think Opus 5.5 is sufficiently safer than it's predecessors that releasing it, more likely than not, reduces risks related to misalignment.I agree that not releasing would not help matters, given the model exists. The real frontier is training new internal models. We don’t get to see it in real time.Pricing is $4/$20, 20% lower than Opus 5, and $0.20 for the cache which is 60% lower. They estimate overall costs will typically drop 40%, while speed is up 30%. Limits on subscriptions have been increased. Fast mode is available at $8/$40, with up to 2.5x the speed. I’d be tempted.This is all a very good price for Fable or Astra level performance. OpenAI focused on lower costs, with GPT-6 Sol prices cut 50% to $2/$10 and Luna cut to only $0.10/$0.50. OpenAI wants you to mix and match models depending on task level. GPT-6 Sol and Luna are pitched as big quality improvements over their old versions, but Sol is not pitched as matching Astra. Sam Altman (CEO OpenAI): GPT-6 Sol and Luna are big improvements on intelligence, alignment, work output, coding, computer use, and more over their 5.6-family predecessors. They are also half the price per token, and even less per task!Whereas Claude now essentially says you should use Opus 5.5 for most tasks. As I said, that’s a strong pitch across the board. They are very good benchmarks. Opus 5.5 is almost universally better on benchmarks than every model except Astra, and is usually above Astra.They add various additional scores, including via graphs, and they mostly look like this, although many are missing Astra from the chart:It goes on like this, and I don’t think it is worth anyone’s time to go through it. Artificial Analysis puts Opus 5.5 into the clear lead overall in intelligence at 58.As tested on max, Opus 5.5 uses so many tokens it is slightly more expensive than Opus 5, and only slightly cheaper than Fable 5.1. You can also run it cheaper, and often should. Opus 5.5 on medium, high and xhigh settings are all on the dotted line that marks the intelligence-cost frontier, as are GPT-6 Luna, GPT-6 Sol and MiMo v2.6 Pro.As its five point lead indicates, Opus 5.5 dominates Astra across most of AA’s tracked benchmarks, including big leads in AA-Briefcase, GDPval-AA SciCode, AA-Omniscience Index and HLE. Astra’s biggest lead is in GDP.pdf.Artificial Analysis adds the new Terminal-Bench-Science 0.1, with GPT-6 Astra at 63% and Opus 5.5 on xhigh at 62%. No other model breaks 50%. Opus 5.5 wins on Omniscience (46 versus 43 for both fable 5.1 and Astra) despite slightly fewer right answers, because it guesses (or hallucinates) less, and is willing to more often admit it doesn’t know. WeirdML v3 still has Astra out in front at 42.2%, with Opus 5.5 in clear second at 31.2% with Fable 5.1 in third at 26%. The best non-OpenAI, non-Anthropic model is Kimi K3 at 7.3%. Sometimes the benchmarks are secret.James Moughan: In my testing it's stronger overall than every other model and cracks a couple of questions nothing else has. And then just occasionally it randomly emits a sentence that is totally gibberish and illogical. It's fun to talk to. Still a huge math nerd. Odd but very good.If you like benchmarks, here is the Chart of Utter Abomination, updated.The places Opus 5.5 falls short of Fable 5.1 fall into a pattern. Fable 5.1 takes all seven Vals professionals rows, as well as MedCode, SAGE and both MLCRs. My Opus 5.5 speculates that this cluster is ‘domain answer graded purely for correctness.’ When presentation and writing do not matter, Fable 5.1’s advantages dominate. When other output details matter, both AIs and humans prefer Opus 5.5. One particularly impressive jump was ProgramBench (fully resolved), where Astra scores 5.5%, Fable 7% and Opus 5.5 jumps to 18.5%, but there are many such cases. I have been pleasantly surprised so far how hard it is to hit the classifiers. Editing my post on the system card still dropped me to Opus 5, which is annoying, but I get it. Billy Gigurtsis: The "Fable 5.1" classifiers they're using for 5.5 have improved considerably when it comes to blocking non-cybersecurity related tasks, at least in my cybersecurity adjacent workloads.gavin leech (Non-Reasoning): Hasn't refused once in a hundred sessions, which is quite surprising. I was sceptical of the claims about improved writing quality but it's certainly much less bad than Fable. As always, for interesting perceptual reasons the cracks will only show up after two weeks.Jai: I ran into an unexpected refusal wall while (ironically) putting together a presentation on effectively working with frontier AI models. I'm not sure what set it off and I haven't heard of anyone else having similar issues yet.The classifiers do still bite in the right locations. Vals in particular tracks this. The classifiers overall fire at similar rates to before, but seem to do so more sensibly, and Opus 5.5 does much better at recovering when it does temporarily hit a classifier. There are still some cases of hitting classifiers where you shouldn’t, but in practice I expect this to be only a minor annoyance unless you are working on bio or cyber.Pliny has you covered, as per usual.As usual, I have included every reaction until I felt like things were repeating themselves, after which I included everything that felt like a fresh take. This was the most consistently positive set of reactions I have ever seen. Astra also had extremely positive reactions. We have two highly excellent models. From what I am seeing here, most prefer Opus 5.5 to Astra if you have to choose one, especially factoring in cost, but they are great models, sir. Astra impressed us by being able to create lots of things in 3D.There are claims that Opus 5.5 can do similar things. The benchmarks certainly say that it can: BenchCAD Vision2Code 73%, with tools 96.2% vs. Astra 95.9%. Its scores on Furniture Assembly (83 vs. Astra’s 80) and Chartography (64.4 vs. Fable’s 44.8) also reflect this. The reactions related to visuals all indicate clear improvement as well. Pulling forward those that mention vision:AllTime: Great conversationalist, I find it about as good as Fable in this regard (much better than Opus 5). Vision is much better, it's not awfully blind! Seeing many my vision tests finally pass with a Claude is great. Seems like a good editor so far, but I haven't used it enough yet.kyle: would it be too much to suggest it’s a bigger jump than opus 4.5 was? incredibly creative, vision is a clear step up even over fable, speaks like a normal human, and i’m struggling to burn through usage.ant cooked hard here. x.5 releases continue to be peakCormundus: It’s wicked fast, which is a nice boon. Overall vision improvements are super for any kind of creative work or tasks where interpreting visual information is key. Computer use is improved, and another fun thing I found having Opus 5.5 play DOOM in my harness: Better spatial awareness and reasoning.4MinuteWarning: First model that is really good at riddles - creating, and solving. Less whimsical