This website uses cookies

Read our Privacy policy and Terms of use for more information.

On 22 September Anthropic released Claude Opus 5.5, and OpenAI followed the same day with GPT-6 Sol and GPT-6 Luna. Both announcements led on price. Anthropic says Opus 5.5 will cost 40% less than Opus 5 on typical workloads at default settings. OpenAI says it has cut API prices for Sol and Luna by 50% compared with their GPT-5.6 promotional pricing.

This is Anthropic's second price cut in three weeks. On 7 September we argued that the Fable 5.1 price cut was real but the savings were conditional, because most of it came from cheaper cache reads. The same issue now applies to both vendors. For firms paying for seats as well as API credits, the final bill depends on how many tokens each task uses and which setting it runs on.

What each vendor actually cut

Opus 5.5 lists at $4 per million input tokens and $20 per million output tokens, which Anthropic says is 20% below Opus 5. Cache reads, the charge for re-reading context the model has already seen, fall 60% to $0.20 per million. Anthropic says the 40% saving comes from those lower prices combined with the model using fewer tokens per task. Subscribers see the difference too. Anthropic says the lower price means Pro, Max and Team plans can go about 25% further than they could on Opus 5.

OpenAI's cut sits on a different rung of its range. Sol falls from $4 to $2 per million input tokens and from $20 to $10 for output. Luna falls from $0.20 to $0.10 and from $1.20 to $0.50. The baseline is GPT-5.6 promotional pricing, and OpenAI still describes GPT-6 Astra as its best model across the board. The two 'price cuts' measure different things against different starting points.

Cost per finished task is the number to watch

Anthropic's own guide to Opus 5.5 costs makes the point plainly. Two models at the same token price can cost very different amounts on the same job, because every extra turn resends the whole conversation. Anthropic says a session made mostly of cache reads can save up to 60% on input, while a short question with a long answer saves up to 20%, since output dominates it. Output tokens, which include the model's thinking, cost five times the input price.

Independent testing adds a warning. Artificial Analysis scores Opus 5.5 at 58 on its Intelligence Index, first among the 211 models it compares it with, but at maximum effort it generated 260M output tokens across the index against a median of 88M. Artificial Analysis finds it very verbose, which can quickly eat into the savings from lower token costs.

One early tester reports a different picture at lower settings.

❝

On US consulting analysis, low thinking effort matched its higher thinking settings on half the output and passed our quality checks.

Carl Bennett, CIO of Deloitte Consulting LLP, told Anthropic

For consultancies and cost managers, that suggests the cheapest setting may be enough for routine analysis, and it is worth proving your own work before paying for maximum effort.

Effort and routing are now the controls

Both vendors are steering buyers towards the same dials.

  • Anthropic advises raising effort before switching to a bigger model, and trying medium effort for well-scoped day-to-day work.

  • Anthropic suggests a small model for lookups, Opus 5.5 for supervised work and Fable 5.1 for the hardest tasks, with Fable 5.1 listing at two and a half times the Opus 5.5 price.

  • OpenAI says developers can now raise or lower reasoning effort mid-conversation without losing cached context, and that cached input reads are discounted by 90%.

  • OpenAI's published cost comparisons are against Opus 5 and Fable 5.1, so they say nothing about how Sol performs against Opus 5.5.

What the outside checks say

METR, which evaluated Opus 5.5 before release, concluded that it likely represents a modest improvement upon Fable 5.1 in AI research and development capability and does not represent a huge leap. METR notes that Anthropic had the chance to review and edit its summary. For buyers, the release is mainly a price and efficiency change, and the bill may move more than the capability.

Box reported a similar gain in token use. Yashodha Bhavnani, VP of AI Products at Box, described the effect on document-heavy work in Anthropic's announcement.

❝

In our evaluations, Claude Opus 5.5 used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy.

Yashodha Bhavnani, VP of AI Products, Box

Firms running agents across contracts, specifications and reports should expect similar gains only if they test for them, because the Artificial Analysis figures show how much the effort setting changes the result.

Takeaway

Price moves are now arriving weeks apart, and each vendor quotes its cut against its own baseline. A firm that builds its AI budget around one rate card will keep planning from stale numbers. As we cover elsewhere this week, Microsoft's new Copilot, announced on 25 September, lets users choose OpenAI's GPT, Anthropic's Opus or an Auto mode, which puts that routing choice in front of ordinary staff. The sensible response is to own the measurement yourself.

  • Keep a small set of real test tasks, such as a cost report narrative, a tender query log or a programme review, and re-run them whenever a vendor changes its prices.

  • Record cost per finished task, including retries, and stop comparing vendors on cost per million tokens.

  • Check which effort level your tools default to, and test whether medium or low holds quality on routine drafting and review.

  • Keep the top model for work where a wrong answer is expensive, and route lookups and summaries to cheaper models.

Two price cuts in three weeks mean the assumptions in most AI budgets are already out of date. Project Flux tracks what these models cost to run on built-environment work every week, so join the Project Flux newsletter and catch the next price move before it reaches your invoice.

All content reflects our personal views and is not intended as professional advice or to represent any organisation.