Sonnet 5 is LIVE And It Competes With Opus

Study Guide

Overview

In this video, Chase AI runs through everything you need to know about Anthropic's newly released Sonnet 5. Sonnet is positioned as the less powerful, less expensive tier below Opus and Fable, and this is the first upgrade to the line since 4.6. The core question Chase keeps returning to is not whether Sonnet 5 is good, but how much falloff there is versus the far pricier Opus 4.8, and the surprising answer is: not much. He argues the real story is nuance, where the right model is now a case-by-case decision driven by task difficulty and effort level rather than a blanket "newer is always better."

Key Takeaways

  • Sonnet 5 is a straight upgrade over 4.6. Across agentic coding, multidisciplinary reasoning, computer use, and knowledge work, it is a large leap forward from the previous Sonnet.
  • It closes the gap to Opus 4.8 dramatically. Versus Opus, the question is only how much falloff there is, and it is small: computer use and agentic coding are only about 2% behind, and knowledge work is actually better.
  • The pricing is the headline. Sonnet 5 is roughly 40% of Opus's input cost and less than half its output cost, yet the benchmark numbers are close. That makes it "huge news" for routine tasks that still want power.
  • Effort level, not just the model, drives results. Low effort on Sonnet 5 is not better than low on 4.6; low mostly means "super cheap." Matching past high-level performance often means running Sonnet 5 on high.
  • Model choice is now case-by-case. For simpler tasks, Sonnet 5 is the clear pick. For complex problems, Opus 4.8 can end up cheaper because it is so token efficient, and it reaches levels Sonnet cannot.

The Benchmarks

Sonnet 5 vs Sonnet 4.6

Chase starts with the benchmarks and reminds viewers that Sonnet is the less powerful, less expensive tier compared with Opus and especially Fable, and that there has not been an upgrade since 4.6. Comparing 5 to 4.6, he calls it a huge leap forward across agentic coding, multidisciplinary reasoning, computer use, and knowledge work. In his words, it is a straight upgrade across the board.

Sonnet 5 vs Opus 4.8

Against Opus 4.8, the framing shifts: "the question isn't, is it close? It's how much of a falloff do we have." And the falloff is small. Sonnet 5 actually posts better numbers on knowledge work, sits only about 2% behind on computer use and roughly 2% behind on agentic coding, and trails by only a handful of percentage points on multidisciplinary reasoning. The biggest gap he calls out is SWE-bench Pro at 63% versus 69%, while Terminal Bench 2.1 is nearly even at 80 versus 82.

Pricing in Context

Chase insists the performance discussion only makes sense alongside pricing. He contrasts the top tiers (Fable 5 / Mythos 5) at around $10 per million input tokens, double Opus, with the mid and value tiers:

  • Opus: $5 per million input tokens, $25 per million output tokens.
  • Sonnet 5: $2 per million input tokens (about 40% of Opus's input cost) and $10 per million output tokens (less than half Opus's output cost).

The upshot: significantly cheaper, yet the numbers land "pretty dang close." For anyone doing more routine rather than bleeding-edge AI work who still wants power, Sonnet 5 is now closing that gap.

Beyond the Benchmarks

Agentic search by effort level

Looking at agentic search performance across effort levels (Opus 4.8 vs Sonnet 5 vs Sonnet 4.6), Chase highlights a large spread in pass rates as effort rises for Sonnet 5. On low, Sonnet 5 hits about 55%, but low on Sonnet 4.6 actually performs better, though at higher cost. On medium, Sonnet 5 roughly matches 4.6's performance at a significant discount. Only at high effort does Sonnet 5 surpass 4.6, and at that point it delivers better performance than 4.6 for what would have been 4.6's low-effort cost.

His caution: it is not that "Sonnet 5 is better across the board." Low on 5 is not automatically better than low on 4.6; low on 5 mostly means "super cheap." To reach past high-level performance you likely need the high range. And on high, Sonnet 5 ends up costing about what Opus does, since Opus on medium and high costs the same as Sonnet on high, but Opus returns a better pass rate.

Agentic computer use

For agentic computer use the picture is cleaner: across the board Sonnet 5 mostly beats Sonnet 4.6, and at a low cost, so here you would essentially always pick Sonnet 5 over 4.6. But he points to Opus again: Opus high performs better than max Sonnet 5, and is cheaper, underscoring how token-efficient Opus is and how it hits higher ceilings Sonnet cannot reach.

Which Model Should You Use?

Chase frames this as the interesting place the field is now in: a case-by-case decision. For a more complex problem, Opus 4.8 may end up cheaper because it is so token efficient, and Sonnet simply may not be the right tool. For tasks that are not as challenging for these models, where using Opus 4.8 may not make sense at all, that is where Sonnet 5 comes in. Even though Opus has a higher cost per million tokens, there remain many cases where Opus is still the efficient choice.

Notable Quotes & Data Points

  • "The question isn't, is it close? It's how much of a falloff do we have. And honestly, it's not that much of a falloff."
  • Knowledge work: Sonnet 5 boasts better numbers than Opus 4.8.
  • Computer use and agentic coding: Sonnet 5 only about 2% behind Opus.
  • SWE-bench Pro: 63% (Sonnet 5) vs 69% (Opus). Terminal Bench 2.1: 80 vs 82.
  • Pricing: Sonnet 5 at $2 input / $10 output vs Opus at $5 input / $25 output; Fable 5 / Mythos 5 around $10 input.
  • Agentic search: Sonnet 5 low around 55%; medium roughly matches 4.6 at a discount; only high surpasses 4.6.

Conclusion

Sonnet 5 meaningfully narrows the distance to Opus while costing far less, making it a strong default for routine work that still needs capability. But the real lesson is nuance: effort level shapes both cost and quality, and for genuinely hard problems Opus 4.8's token efficiency and higher ceiling can make it the better and even cheaper choice. The right pick is now a case-by-case call.

YouTube