In this video, Chase AI runs through everything you need to know about Anthropic's newly released Sonnet 5. Sonnet is positioned as the less powerful, less expensive tier below Opus and Fable, and this is the first upgrade to the line since 4.6. The core question Chase keeps returning to is not whether Sonnet 5 is good, but how much falloff there is versus the far pricier Opus 4.8, and the surprising answer is: not much. He argues the real story is nuance, where the right model is now a case-by-case decision driven by task difficulty and effort level rather than a blanket "newer is always better."
Chase starts with the benchmarks and reminds viewers that Sonnet is the less powerful, less expensive tier compared with Opus and especially Fable, and that there has not been an upgrade since 4.6. Comparing 5 to 4.6, he calls it a huge leap forward across agentic coding, multidisciplinary reasoning, computer use, and knowledge work. In his words, it is a straight upgrade across the board.
Against Opus 4.8, the framing shifts: "the question isn't, is it close? It's how much of a falloff do we have." And the falloff is small. Sonnet 5 actually posts better numbers on knowledge work, sits only about 2% behind on computer use and roughly 2% behind on agentic coding, and trails by only a handful of percentage points on multidisciplinary reasoning. The biggest gap he calls out is SWE-bench Pro at 63% versus 69%, while Terminal Bench 2.1 is nearly even at 80 versus 82.
Chase insists the performance discussion only makes sense alongside pricing. He contrasts the top tiers (Fable 5 / Mythos 5) at around $10 per million input tokens, double Opus, with the mid and value tiers:
The upshot: significantly cheaper, yet the numbers land "pretty dang close." For anyone doing more routine rather than bleeding-edge AI work who still wants power, Sonnet 5 is now closing that gap.
Looking at agentic search performance across effort levels (Opus 4.8 vs Sonnet 5 vs Sonnet 4.6), Chase highlights a large spread in pass rates as effort rises for Sonnet 5. On low, Sonnet 5 hits about 55%, but low on Sonnet 4.6 actually performs better, though at higher cost. On medium, Sonnet 5 roughly matches 4.6's performance at a significant discount. Only at high effort does Sonnet 5 surpass 4.6, and at that point it delivers better performance than 4.6 for what would have been 4.6's low-effort cost.
His caution: it is not that "Sonnet 5 is better across the board." Low on 5 is not automatically better than low on 4.6; low on 5 mostly means "super cheap." To reach past high-level performance you likely need the high range. And on high, Sonnet 5 ends up costing about what Opus does, since Opus on medium and high costs the same as Sonnet on high, but Opus returns a better pass rate.
For agentic computer use the picture is cleaner: across the board Sonnet 5 mostly beats Sonnet 4.6, and at a low cost, so here you would essentially always pick Sonnet 5 over 4.6. But he points to Opus again: Opus high performs better than max Sonnet 5, and is cheaper, underscoring how token-efficient Opus is and how it hits higher ceilings Sonnet cannot reach.
Chase frames this as the interesting place the field is now in: a case-by-case decision. For a more complex problem, Opus 4.8 may end up cheaper because it is so token efficient, and Sonnet simply may not be the right tool. For tasks that are not as challenging for these models, where using Opus 4.8 may not make sense at all, that is where Sonnet 5 comes in. Even though Opus has a higher cost per million tokens, there remain many cases where Opus is still the efficient choice.
Sonnet 5 meaningfully narrows the distance to Opus while costing far less, making it a strong default for routine work that still needs capability. But the real lesson is nuance: effort level shapes both cost and quality, and for genuinely hard problems Opus 4.8's token efficiency and higher ceiling can make it the better and even cheaper choice. The right pick is now a case-by-case call.