Your Everyday AI Just Learned to Finish the Job
Claude Sonnet 5 quietly became the default for millions of free and Pro users, and it's built to see multi-step work through to the end
If you opened Claude this week and everything looked exactly the same, you'd be forgiven for missing the news. But behind that unchanged chat box, the model doing the actual work quietly changed. On June 30, Anthropic released Claude Sonnet 5 and, starting the next day, made it the default for every free and Pro user 1. No setting to flip, no upgrade to buy. The AI you already use just got a meaningful bump.
That's worth paying attention to, not because of the version number, but because of what this particular update is trying to do. Sonnet 5 isn't chasing a flashier chat experience. It's built to be better at finishing things on its own. And once you understand why that matters, a lot of the confusing "agent" talk in the news starts to make sense.
What "agentic" actually means
You've probably seen the word "agentic" thrown around a lot lately. Stripped of the hype, it means something simple: an AI that doesn't just answer a question, but takes a series of steps to complete a task, using tools along the way. Anthropic describes Sonnet 5 as able to "make plans, use tools like browsers and terminals, and run autonomously" at a level that used to require much bigger, more expensive models 1.
Here's the everyday version. A regular chatbot is like asking a knowledgeable friend for directions. An agentic model is more like handing that friend your car keys and asking them to actually go pick up the package, deal with the traffic, and text you when it's done. The first gives you information. The second gets the job across the finish line.
The catch with earlier models was that they'd often stop halfway. They'd draft a plan, complete the easy first step, then stall and wait for you to nudge them along. If you've ever asked an AI to do something in three parts and watched it quietly do one and forget the rest, you've felt this.
The upgrade that's easy to miss
Sonnet 5's headline improvement is exactly that follow-through. Anthropic's early testers reported that it completes complex tasks where previous versions would stop short, and that it checks its own output without being asked to 1. One example the company shared came from an engineer at Zapier: they handed the model a two-part job, update account tiers in Salesforce and then send a launch announcement to enterprise contacts, and it finished the whole thing end to end. That workflow, the engineer said, "used to stall halfway" 2.
That self-checking habit is the part I'd underline. A model that reviews its own work before handing it back is a model you have to babysit less. It won't eliminate mistakes, but it reduces the number of times you catch an error the AI could have caught itself.
The numbers back up a real, if modest, jump. On a benchmark that measures agentic coding, essentially whether a model can write, run, and fix code across several steps, Sonnet 5 scores 63.2%, up from 58.1% for the previous Sonnet and closing in on the far pricier flagship Opus 4.8 at 69.2% 3. You don't need to memorize those figures. The point is the direction: the affordable, default model is creeping up toward what only the premium tier could do a few months ago.
Why cheaper is the real story
The most important number here isn't a benchmark. It's the price. Sonnet 5 launched at an introductory rate of $2 per million words of input and $10 per million words of output through the end of August, before settling at $3 and $15 2. That undercuts Anthropic's own Opus 4.8, and it comes in below OpenAI's GPT-5.5 and Google's Gemini 3.1 Pro too 2.
Why should you care what businesses pay per million tokens? Because cost is what decides whether these tools show up in the software you actually use. When capable AI gets dramatically cheaper to run, companies start building it into everyday products, the help desk, the scheduling tool, the invoicing app, instead of reserving it for premium features. The falling price is the quiet mechanism that puts AI in front of ordinary users. This week's release is one more step down that curve.
Sonnet 5 also lets you dial its "effort" up or down, trading speed and cost for extra thoroughness depending on the task 3. Think of it as a rough-draft setting versus a careful-review setting. Most of your day doesn't need the expensive deep-thinking mode, and now you don't have to pay for it when you don't.
The honest limitations
Here's where the coffee-chat honesty comes in: this is an improvement, not a leap, and it isn't the smartest model available. For the hardest reasoning, deep research, subtle judgment calls, genuinely difficult science, Opus 4.8 still comes out ahead 3. Some independent reviewers also note that Sonnet 5 seems to trade a little reasoning depth for coding speed, so for the very trickiest problems, the pricier model may still be the safer pick 3.
There's good news on the safety side, though. Anthropic's own testing found Sonnet 5 is better than its predecessor at refusing malicious requests, more resistant to manipulation, and less prone to both hallucinating and to sycophancy, that tendency to just tell you what you want to hear 1. When an AI is actually taking actions in your tools rather than just chatting, a model that reliably knows when to say no matters as much as one that knows how to say yes 3.
What this means for you
You don't need to do anything to get Sonnet 5. If you use Claude on a free or Pro plan, you already have it. The practical shift is in what you can reasonably ask for. Tasks that span several steps, "pull this data, summarize it, then draft the email," are more likely to run all the way through now instead of stopping partway.
The bigger picture is the one worth sitting with. Anthropic, OpenAI, and Google are all racing to make "AI that finishes tasks" the baseline expectation, not a premium add-on 2. That's the trend behind the trend. And it's not really a story about any single model. It's about a capability that keeps getting cheaper and more reliable, which means it'll keep quietly showing up in more of the tools you already rely on.
If that makes you a little uneasy, that's a fair reaction. But notice the shape of it: the useful move isn't to out-type the machine, it's to get comfortable handing off the multi-step busywork and keeping your attention on the judgment calls, the priorities, and the final sign-off. Those are the parts that still very much need you. The tools are getting better at the middle of the job. You're still in charge of the ends.
Sources
- Anthropic — Introducing Claude Sonnet 5 "Official announcement, benchmarks, pricing, and safety evaluations"
- TechCrunch — Anthropic launches Claude Sonnet 5 as a cheaper way to run agents "Reporting by Rebecca Bellan on pricing and competitive context"
- DataCamp — Claude Sonnet 5: Features, Benchmarks, and Pricing "Independent review with benchmark table and caveats"