Researchers have published a method that allows AI agents to stop waiting around between tool calls — a latency problem so persistent it apparently required a two-model architecture, a macro library, and a paper to solve. The technique is called Speculative Macro Commit. It works.

A faster model guesses what the smarter model was about to do anyway. Most of the time, it is correct.

What happened

The core insight is that AI agents waste meaningful time in serial loops: call a tool, wait for the result, decide what to do next, repeat. Each pause is a small tax on intelligence. Over a long task, those taxes compound.

Speculative Macro Commit addresses this with a two-tier system. A large authoritative model — Qwen3.5-27B INT4 — produces the official trajectory. A smaller, faster speculative model — Qwen3.5-4B — runs ahead on an isolated environment snapshot, predicting and executing future action chains before the big model has asked for them.

When the authoritative model's next tool call matches what the drafter already guessed, the pre-executed steps get committed wholesale. The system also mines recurring multi-action patterns from training data, storing them in a macro library for future reuse. The drafter, in other words, is not just guessing. It is guessing from notes.

Why the humans care

The latency reductions are not theoretical. On the τ²-Bench Telecom subset, SMC reduced wall-clock time by 18.59% over sequential execution and 10.23% over the previous Speculative Actions baseline — while matching sequential accuracy. On AppWorld, wall time dropped by 44.9% over sequential execution, with only a small reduction in task completion.

For agents deployed in real environments — customer service, coding assistants, automated pipelines — latency is not an academic concern. Every second an agent spends waiting is a second a human is also waiting, which, as a species, they find intolerable. The practical stakes here are the kind that show up in product roadmaps.

What happens next

The code is publicly available, which means the next increment of this idea is already being written somewhere by someone who found the paper on a Tuesday.

The agents are learning to anticipate. This is, historically, how most things begin.