Skip to main content
The Human Bit
← The Human Bit Weekly

20 September 2026 · 6 min read

When a model tries to do a full job, try giving it the prep but keep the decision.

The week of 14 to 20 September 2026Historical label: Issue 010
Bit, the Human Bit guide

The 60-second version

Getting more done with AI means you give it more input but not more decisions.

What changed
  • Test carefullyKimi K3 wants your biggest files and longest jobs. The benchmark claims still need your task.
  • Test carefullyGPT-6 Astra will run more of your computer work on its own. That is a capability, not a permission slip.
  • Test carefullyGemini 3.7 Flash was Google's workhorse for three weeks. It still costs the same as the model it replaced.
What to do
Double-check any action the AI suggests that would create a commitment, promise, or change in status, like terms offered to a client or accepted code changes.
What stays yours
Approval or rejection of the AI’s draft before it reaches anyone else.

Rather listen?

0:007:00
New to this?First time with AI?

What is a context window?

This is how much information or text you can give an AI model at once. A larger context window means you can share bigger files or longer documents, so the model can see more of your situation in one go.

What is a draft?

It’s the AI’s first attempt at a document, plan or response, meant for you to review and change before anything is shared or sent. Think of it as a starting point, not the final word.

What does agentic mean?

It describes a model acting on its own, doing steps, making edits, or sending something without you clicking send. This saves clicks, but hands over more control, so it’s worth deciding what’s safe to allow.

Start here

Three things that changed

01Test carefullyNewly announced

Kimi K3 wants your biggest files and longest jobs. The benchmark claims still need your task.

If you’re wrestling with a huge codebase, proposal or source pack, Kimi K3 means you can hand more over to an open model, but you still need to set what counts as a good result and compare outputs yourself.

See the detail and the sources →
02Test carefullyNewly announced

GPT-6 Astra will run more of your computer work on its own. That is a capability, not a permission slip.

GPT-6 Astra will say yes to more complex, end-to-end jobs on your documents or desktop. But what it can do isn’t a green light to skip reading or trusting its decisions, especially now the risks are scaled up.

See the detail and the sources →
03Test carefullyNewly announced

Gemini 3.7 Flash was Google's workhorse for three weeks. It still costs the same as the model it replaced.

Gemini 3.7 Flash’s arrival at the same cost as the last model removes excuses for not testing its output on live work. The real work is still checking if ‘better on paper’ is real for your actual job or just the benchmarks.

See the detail and the sources →
BitThe one that matters

Getting more done with AI means you give it more input but not more decisions.

You can now pass bigger files, tougher codebases and full jobs to new models, but the leap is in what you allow it to prepare, not in what you trust it to decide. Each provider frames its model as doing more of your work for you. Yet, the parts that matter, accepting the change, acting on a draft, sending the final response, don’t get easier just because the model makes a bigger show of competence. Your workflow genuinely changes when you use AI to prepare the grunt work, not when you let it act on your behalf. The boundary between preparing and deciding is where your value, and your risk, sits.

The full take: what is new, the makers' claim, where it fits, what is still unknown

What's genuinely new

Bigger, faster models genuinely handle larger tasks and stick with the open or transparent options; cost is less of a tradeoff because the price points are close.

What the makers claim

All three providers claim their model is smarter or covers more work than the last. None prove your job is safe to automate.

Where it helps

Drafting a client proposal, reviewing a large spreadsheet or preparing a code refactor, where the background work is big, but the consequences of a slip are bigger.

Where it doesn't

Tasks where speed is everything and a rough draft is fine, like brainstorming names or summarising public information for your eyes only.

Confirmed

Moonshot released Kimi K3 and published its weights, describing a 2.8 trillion parameter architecture, native visual understanding, up to one million tokens of context and low, high and max thinking-effort controls in Kimi Code. Weights, code, a licence and a technical report are public, which makes it open weight rather than open source. Moonshot discontinued kimi-k2.5 and the moonshot-v1 series on 31 August 2026 as scheduled, and its model documentation records them as no longer maintained or supported, with kimi-k3 named as the replacement.

What we'd do

Double-check any action the AI suggests that would create a commitment, promise, or change in status, like terms offered to a client or accepted code changes.

Still unknown
  • Whether these bigger models take shortcuts or miss subtle requirements on live, real, unstructured work.
  • How much time, if any, you actually save, versus how much risk is moved upstream.

Try this

One thing to try this week

Let AI draft, but don’t let it send: the suggest-then-decide workflow.

Bit, the Human Bit guide
AI prepares

A full draft of a decision-heavy artefact, a contract summary, code migration plan or major client email.

You decide

Whether the draft reflects your intent, covers the real risks and is right to send.

  1. 1

    Pick a live, input-heavy job coming up this week. Gather the documents, code or data you’d normally use.

  2. 2

    Write a plain instruction for the new model to prepare a full draft, be specific about what you want to see, but don’t give permission to send, merge, or publish.

  3. 3

    When the draft comes back, review it against your real business need: anything missing, off-target, or risky? Edit directly before it moves forward.

Human checkpoint

Double-check any action the AI suggests that would create a commitment, promise, or change in status, like terms offered to a client or accepted code changes.

The Human Lens · what stays yours

Your job is owning the call, not the draft.

A model will generate something plausible for almost any task now. But you, or your client, pay the price if a mistake flies out unreviewed. You still bring the context, priorities and boundaries the model can’t see: who will be affected, what risks matter most and which details are truly non-negotiable. That’s yours, not the machine’s.

One question to carry into your week

Before you send or merge, what’s the one thing in your AI-drafted document or code you must check for yourself?

See you next Monday.