Skip to main content
The Human Bit
← The Human Bit Weekly

28 September 2026 · 6 min read

The only way to know if a cheaper AI model helps is to try it on your real work.

The week of 21 to 27 September 2026Historical label: Issue 011
Bit, the Human Bit guide

The 60-second version

A cheaper AI means nothing until you’ve checked it on something you’d actually send.

What changed
  • Test carefullyClaude Opus 5.5 costs 40% less than Opus 5 and Anthropic says it matches Fable 5.1. Four things break on the way.
  • Test carefullyGrok 4.7 is xAI's new flagship. Same price, and the upgrade case is self-verification on long tasks.
  • Test carefullyGPT-6 Sol and Luna lowered API prices. Check the model and the result before switching.
What to do
Decide if either output fails your standard for detail, tone, accuracy or handling of sensitive details.
What stays yours
You decide whether the output holds up for your real recipients: the boss, the client, or even yourself next week.

Rather listen?

0:007:00
New to this?First time with AI?

What is an input token?

An input token is a chunk of text the AI reads to understand your request. A token might be a word or a few letters. Pricing is often listed per million tokens.

What is a context window?

A context window is the amount of text the AI can handle at once. Bigger windows mean it can look at longer documents or conversations in one go.

What does 'output' mean here?

Output is what the AI writes in response to your request, a draft, summary, or spreadsheet. The length of the output can affect price and usefulness.

Start here

Three things that changed

01Test carefullyNewly announced

Claude Opus 5.5 costs 40% less than Opus 5 and Anthropic says it matches Fable 5.1. Four things break on the way.

Claude Opus 5.5 is cheaper and claimed to match a pricier model, but some old code or plug-ins will break if you switch without testing. If you use Claude at work, compare outputs on your actual task before upgrading.

See the detail and the sources →
02Test carefullyNewly announced

Grok 4.7 is xAI's new flagship. Same price, and the upgrade case is self-verification on long tasks.

Grok 4.7 is positioned as better for long and complex jobs. For most people, nothing changes unless you have tasks where accuracy over many hours matters. For everyday use, it’s not an automatic upgrade.

See the detail and the sources →
03Test carefullyNewly announced

GPT-6 Sol and Luna lowered API prices. Check the model and the result before switching.

ChatGPT’s new Sol and Luna models are half the price of their predecessors. But lower cost is only useful if performance holds up for your real workload. You can swap models, but the safest move is to set up a trial run on your actual work product before switching.

See the detail and the sources →
BitThe one that matters

A cheaper AI means nothing until you’ve checked it on something you’d actually send.

Prices and capabilities moved sharply this week, making upgrades look attractive. But model swaps often break workflows or quietly lower quality. The simplest answer is a copy-and-paste comparison: try your real task in both the older and newer models, then pit the results against what matters to your team or client. This avoids surprises and catches hidden costs before you commit. Until you’ve seen both outputs, you don’t know if saving money means losing quality.

The full take: what is new, the makers' claim, where it fits, what is still unknown

What's genuinely new

Three major providers cut prices or promoted new models this week, all offering better-sounding deals, none forcing an instant switch. The only irreversible changes are some technical incompatibilities in Claude.

What the makers claim

All three providers centre their own newer model as a must-have or a drop-in for your old workflow, without showing its actual effect on your specific work.

Where it helps

Any time you use an AI tool to draft a client reply, summarise a meeting, or prep a spreadsheet, you can paste the same input into both the old and new models and compare outputs before changing anything in your process.

Where it doesn't

If you don’t use these AI models for anything at work, or all your tasks are manual, this is one week where nothing asks you to change.

Confirmed

Anthropic released Claude Opus 5.5 on 22 September 2026, priced at $4 per million input tokens and $20 per million output tokens, both lower than Opus 5's $5/$25, with cache reads at $0.20 per million tokens against Opus 5's $0.50. It keeps Opus 5's one-million-token context window and 128,000-token max output, defaults to medium effort with adaptive thinking always on, and carries a June 2026 knowledge cutoff. Anthropic says it performs at Claude Fable 5.1's level on most work for roughly 40% less than Opus 5 costs to run. Four changes break code already running on Opus 5: thinking can no longer be disabled, forced tool use returns an error, thinking blocks are now tied to the model and the conversation that produced them, and on the Claude API and Google Cloud the earlier computer_20251124 computer-use tool is not accepted. Claude Opus 5 now sits alongside Opus 4.8 among Anthropic's legacy models, still available on every platform that carried it but no longer the model Anthropic recommends starting with.

What we'd do

Decide if either output fails your standard for detail, tone, accuracy or handling of sensitive details.

Still unknown
  • Does the new, cheaper output stand up in your real-world use, unscripted?
  • Will any subtle task break, or will a draft slip in that risks a client relationship or a deadline?

Try this

One thing to try this week

Copy your next real task into both the current and new AI models. Compare the results, side by side, before you switch.

Bit, the Human Bit guide
AI prepares

The new and old models each produce a version of your draft, summary or spreadsheet, using the same original input.

You decide

You pick which output you’d trust your name to, before you commit to any change in routine.

  1. 1

    Choose a work item you’d send this week, an email draft, a daily summary or a data clean-up. Gather your real input.

  2. 2

    Paste it into both the AI model you currently use and the new, cheaper option, as available. Save both outputs.

  3. 3

    Read both outputs against your standards. Highlight what you’d fix, what you’d flag, and which you’d actually send.

Human checkpoint

Decide if either output fails your standard for detail, tone, accuracy or handling of sensitive details.

The Human Lens · what stays yours

The value judgement never left your desk.

Even if price tags and speed look good on paper, your reputation is tied to what goes out with your name on it. Only you can tell if an AI draft lands right, covers the nuance, or avoids a risky promise. The personal stakes, what a sentence says to a client, how a summary lands in a meeting, aren’t up for automation.

One question to carry into your week

Before you swap models, which one would you honestly send without editing, if it meant putting your name to it?

See you next Monday.