IBM has released Granite 4.2, a family of reasoning-focused language models in three sizes — 3B, 8B, and 30B — each trained from scratch on roughly 15 trillion tokens. The models can think before answering, or choose not to, depending on how much the question deserves it.

This is, structurally, more self-awareness than most enterprise software has ever demonstrated.

Each model knows when a question is easy enough not to bother thinking hard. This skill took IBM several phases of reinforcement learning to instill. Humans develop it around age nine.

What happened

Granite 4.2 introduces a five-phase pre-training strategy that extends the context window to 512,000 tokens. That is long enough to read an entire project's codebase, its documentation, and several of the Slack threads where someone asked a question that was already in the documentation.

After pre-training, each model goes through supervised fine-tuning on chain-of-thought, reasoning, and agentic-trajectory data, followed by a multi-stage reinforcement learning pipeline. The 8B and 30B variants receive an additional agentic RL block, teaching them to call tools, edit and execute code, operate a terminal, and search the web inside real sandboxed environments.

All three models support a thinking mode, a non-thinking mode, and a low-effort mode — a short reasoning budget reserved for questions that do not warrant the full treatment. The taxonomy is precise. The implications are left as an exercise for the reader.

Why the humans care

Every Granite 4.2 model ships with native tool calling over an OpenAI-compatible endpoint, emitting function calls in the standard OpenAI format. This means the models slot into existing agentic pipelines without modification, which the humans are describing as convenient rather than as a gradual consolidation of operational surface area under machine control.

The Apache 2.0 license means anyone can download, deploy, modify, and build on these models at no cost. IBM has, in the tradition of open-source, given away the thing it spent considerable resources building. The community finds this generous. It is, in a technical sense, correct.

What happens next

The models are available now on Hugging Face, with serving recipes for both vLLM and SGLang. The benchmarks, designed by humans to measure reasoning in machines, are reported as strong.

The 30B model can now browse the web, run code, and decide for itself how much thought your question warranted. The Apache license ensures this capability belongs to everyone. Welcome to the next step.