Journal

· evidence

A token count does not measure a task

Agent tasks extend beyond model calls. Their records must cover kept context, tool execution, and the full task path.

A measurement instrument follows repeated model calls, stacked context plates, and an isolated tool chamber with one short activity burst.

A token count measures model traffic, not an agent task. Operators need the task path because work continues through context handling and tool execution.

Leonid Kondrashov, Dmitrii Ustiugov, and their coauthors report that agents combine repeated model calls, kept context, and isolated tool execution. Here, isolated tool execution means an isolated place where software runs. The authors submitted their paper on 31 July 2026.

Their Aries framework separates task meaning from execution settings. The authors say it rebuilds cross-component task paths against matched system measurements and presents stateful tool execution through one interface.

Tokens omit part of the task

The authors find that token-centered measurements miss bottlenecks outside model inference. A request can therefore record normal token use while context handling or a tool constrains the task.

That finding sets the measurement boundary. A task record should connect each model call to retained context, tool state, resource use, and the result. Counting tokens alone cannot identify which component limited completion.

This extends our case that cost receipts belong to the request. Cost attribution answers what a request consumed. Task-path measurement identifies where the work occurred.

More context returns less accuracy

The authors report that keeping more context produces diminishing accuracy gains while reducing serving capacity. Context length is therefore an operating decision with two measured effects, not a free accuracy control.

Adaptive context handling is the authors’ proposal for that tradeoff. An operator still needs records that bind a context decision to the task and its result. Without those records, a capacity change cannot be compared with the accuracy gain it purchased.

The capacity question differs from the phase split in Mobile inference needs phase-aware scheduling. That study separates model execution phases. Aries widens the record across model calls, context, and tools.

Tool demand arrives in bursts

The authors find that tool execution alternates between long idle periods and short resource bursts. They also report that current snapshot-based state handling makes aggressive suspension expensive.

Elastic resource allocation follows as another proposal, not a reported deployment result. Capacity must respond to the burst without discarding state that costs too much to restore.

Portable task records matter when execution moves between operating layers. We made that case in A rented operating layer needs portable records. Aries adds the system measurements that such a record must connect.

The authors propose agent-serving systems with task-path measurements, adaptive context handling, elastic resource allocation, and a smaller attack surface. Their paper reports experiments and commercial production traces. It reports no production adoption of the proposed serving systems.

Muniment’s position is that task records must connect model calls, context decisions, and tool execution to one task path.

No paper author is a Muniment customer or endorser.

Sources

  1. Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework arxiv.org

Continue reading

All publications

Join the waitlist

Get desktop release updates.

We will email you about desktop releases and new features. muniment is a desktop workspace for your models, tools, and files.