Skip to main content
The Human Bit

The Lab · 20 current model entries

Know the model. Choose the job.

Names change faster than people can learn them. The Lab keeps official positions, independent observations and Human Bit tests separate, then lets you compare only the fields relevant to your job.

Coverage rule

Major current text and reasoning models people can choose or deploy.

We include a model when it is currently accessible, materially different from another tier and relevant to a real user decision. The catalogue order is not a rank.

OpenAI · Anthropic · Google · xAI · Moonshot AI · DeepSeek

Filter the bench

Compare recorded fit, not a universal score.

Start with the work surface or provider you can actually use. Select up to three entries to compare their stated fit, limits, evidence coverage and review dates side by side.

More filters
20 of 20 entries · catalogue order, not rank
OpenAI · GPT-5Available

GPT-5.5 Instant

A fast ChatGPT model for routine questions, drafting, rewriting and low-consequence transformations. It is no longer the everyday default; GPT-5.6 Luna now fills that role for Free and Go users.

Useful clue
Quick explanations and first drafts
Access
ChatGPT
Official record onlyinformational · gpt-55-instant@v2
Reviewed 8 Sept 2026Review by 22 Sept 2026
OpenAI · GPT-5.6Rolling out

GPT-5.6 Sol

OpenAI’s main reasoning model for complex knowledge work, research, coding and tasks that need more deliberate analysis.

Useful clue
Conflicting evidence and multi-step analysis
Access
ChatGPT · Work · Codex · API
Official + independent noteinformational · gpt-56-sol@v3
Reviewed 31 Aug 2026Review by 14 Sept 2026
OpenAI · GPT-5.6Rolling out

GPT-5.6 Sol Pro

The highest-capability GPT-5.6 option for difficult and longer-running work where a standard reasoning pass is not enough.

Useful clue
Long-running professional workflows
Access
ChatGPT Pro · API
Official + independent noteinformational · gpt-56-sol-pro@v2
Reviewed 8 Sept 2026Review by 23 Sept 2026
OpenAI · GPT-6Rolling out

GPT-6 Astra

OpenAI’s new flagship for research, coding, computer-use and long-running agentic work. It began rolling out on 3 September 2026, first to a limited set of organisations and then to ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API and AWS.

Useful clue
Research, coding and computer-use work that runs several steps on its own
Access
ChatGPT · API · AWS
Official record onlyinformational · gpt-6-astra@v1
Reviewed 5 Sept 2026Review by 19 Sept 2026
Anthropic · Claude SonnetAvailable

Claude Sonnet 5

Anthropic’s current Sonnet balance for everyday professional work, coding and agentic tasks, with adaptive thinking by default.

Useful clue
Professional writing and analysis
Access
Claude · Claude API · Cloud platforms
Official + independent noteinformational · claude-sonnet-5@v1
Reviewed 6 Sept 2026Review by 20 Sept 2026
Anthropic · Claude OpusAvailable

Claude Opus 5

Anthropic’s current model for complex agentic coding and enterprise work, and the model it points Opus 4.8 users at. Thinking is on by default, so the effort setting, not a thinking switch, is how you control depth and cost.

Useful clue
Complex agentic coding and long-horizon tasks
Access
Claude · Claude API · Cloud platforms
Official record onlyinformational · claude-opus-5@v1
Reviewed 6 Sept 2026Review by 20 Sept 2026
Anthropic · Claude OpusAvailable

Claude Opus 4.8

A high-capability Claude model for complex agentic coding and enterprise work, now superseded by Claude Opus 5 but still available on every platform that carried it.

Useful clue
Existing integrations not yet migrated to Opus 5
Access
Claude · Claude API · Cloud platforms
Official record onlyinformational · claude-opus-48@v1
Reviewed 7 Sept 2026Review by 22 Sept 2026
Anthropic · Claude FableAvailable

Claude Fable 5

Anthropic's prior most capable widely released model for demanding reasoning and long-horizon agentic work, now superseded by Fable 5.1 but still listed as a legacy model with unchanged pricing and limits.

Useful clue
Existing integrations not yet migrated to Fable 5.1
Access
Claude · Claude API · Cloud platforms
Official + independent noteinformational · claude-fable-5@v1
Reviewed 6 Sept 2026Review by 20 Sept 2026
Anthropic · Claude FableAvailable

Claude Fable 5.1

Anthropic's current model for demanding reasoning and long-horizon agentic work, at the same headline price as Fable 5 but with prompt-cache reads cut to a quarter of the cost, and stronger long-running coding, research and document work.

Useful clue
Very long and difficult agentic coding and research projects
Access
Claude · Claude API · Cloud platforms
Official record onlyinformational · claude-fable-51@v1
Reviewed 2 Sept 2026Review by 16 Sept 2026
Google · GeminiAvailable

Gemini 3.5 Flash

Google’s earlier fast Flash model, still stable and deployed broadly across consumer, developer and enterprise products, but no longer the current Flash: Google now points new work at 3.6, 3.7 or 3.8 Flash.

Useful clue
Fast multimodal document work
Access
Gemini app · AI Mode in Google Search · Google AI Studio · Gemini API · Gemini Enterprise
Official + independent noteinformational · gemini-35-flash@v1
Reviewed 7 Sept 2026Review by 24 Sept 2026
Google · GeminiPreview

Gemini 3.1 Pro Preview

A preview model for precise multi-step reasoning, software engineering and tool use across text, images, video, audio and PDFs.

Useful clue
Complex multimodal analysis
Access
Gemini API · Google AI Studio · Vertex AI
Official record onlywatch · gemini-31-pro-preview@v1
Reviewed 7 Sept 2026Review by 15 Sept 2026
Google · GeminiAvailable

Gemini 3.1 Flash-Lite

Google’s lower-latency, lower-cost model for high-volume extraction and lightweight multimodal jobs.

Useful clue
Classification and simple extraction
Access
Gemini API · Google AI Studio · Vertex AI
Official record onlyinformational · gemini-31-flash-lite@v1
Reviewed 7 Sept 2026Review by 24 Sept 2026
xAI · GrokAvailable

Grok 4.5

xAI’s previous flagship for coding, agentic tasks and knowledge work. It remains available, but xAI released Grok 4.6 on 12 August 2026 and now recommends that as its most capable model; The Human Bit's lab now carries Grok 4.6 as its own entry.

Useful clue
Engineering and coding evaluation
Access
Grok Build · Cursor · xAI API
Official record onlyinformational · grok-45@v2
Reviewed 9 Sept 2026Review by 26 Sept 2026
xAI · GrokAvailable

Grok 4.6

xAI's current flagship, recommended for coding and chat, building on Grok 4.5 with more emphasis on long-running agents and interactive or visual work.

Useful clue
Long-running agentic coding and software engineering
Access
Grok Build · Cursor · xAI API
Official record onlyinformational · grok-46@v1
Reviewed 9 Sept 2026Review by 27 Sept 2026
Moonshot AI · KimiAvailable

Kimi K3

Kimi's open-weight multimodal flagship for long-horizon coding, knowledge work and tool-using tasks, with a one-million-token context window.

Useful clue
Large document and knowledge-work packs
Access
Kimi · Kimi Work · Kimi Code · Kimi API · Open weights
Official + independent noteinformational · kimi-k3@v2
Reviewed 9 Sept 2026Review by 28 Sept 2026
Moonshot AI · Kimi CodeAvailable

Kimi K2.7 Code

Kimi’s dedicated coding model for long-context, tool-using software work, with a separate high-speed endpoint for the same model.

Useful clue
Long-context repository work
Access
Kimi API · OpenAI-compatible clients
Official record onlyinformational · kimi-k27-code@v1
Reviewed 6 Sept 2026Review by 20 Sept 2026
Google · GeminiAvailable

Gemini 3.6 Flash

Google's fast prior-generation model for agentic coding, code generation and spatial reasoning, now sharing the Flash tier with Gemini 3.7 Flash, which Google names as the latest and most capable Flash model.

Useful clue
Rapid agentic coding loops and iteration
Access
Gemini app · AI Mode in Google Search · Google AI Studio · Gemini API · Gemini Enterprise
Official record onlyinformational · gemini-36-flash@v1
Reviewed 31 Aug 2026Review by 14 Sept 2026
Google · GeminiAvailable

Gemini 3.7 Flash

Google's prior fastest Flash model for complex agentic tasks at scale, now sharing the Flash tier with Gemini 3.8 Flash, which Google names as the latest and most capable Flash model.

Useful clue
Existing integrations not yet migrated to 3.8 Flash
Access
Gemini app · AI Mode in Google Search · Google AI Studio · Gemini API · Gemini Enterprise
Official record onlyinformational · gemini-37-flash@v1
Reviewed 9 Sept 2026Review by 29 Sept 2026
Google · GeminiAvailable

Gemini 3.8 Flash

Google's current fastest Flash model for complex agentic tasks at scale, positioned as the successor to 3.7 Flash with stronger long-horizon software engineering and multi-step reasoning, while currently sharing the same published rate.

Useful clue
Long-horizon software engineering and multi-step coding loops
Access
Gemini app · AI Mode in Google Search · Google AI Studio · Gemini API · Gemini Enterprise
Official record onlyinformational · gemini-38-flash@v1
Reviewed 2 Sept 2026Review by 16 Sept 2026
DeepSeek · DeepSeek V4Available

DeepSeek V4 (V4-Flash and V4-Pro)

DeepSeek's V4 API generation, offered as a cheap high-volume Flash model and a costlier, more capable Pro model, both with a million-token context window.

Useful clue
High-volume or cost-sensitive tasks where token price dominates
Access
DeepSeek API
Official record onlyinformational · deepseek-v4@v1
Reviewed 31 Aug 2026Review by 14 Sept 2026

The model list can grow. The standard cannot loosen.

Each entry needs an official record, a practical decision, an evidence state and a review owner. Missing independent evidence remains visible rather than being filled with a guess.

Read the evidence standard →