Engineering essay
Completion needs readback
A successful tool call is one event. Completion needs evidence of the intended state.
Define the state before choosing the tool
An agent can produce a convincing account of work that never reached its destination. The API accepts a request, the model sees a success response, and the task is marked complete. Meanwhile the record is unchanged, the wrong account was updated, or the result was overwritten a moment later.
My starting point is to describe completion as a state that another reader can inspect. For a fictional task to change a library opening time, completion means the correct branch record has the requested time, the public page displays it, and the evidence identifies when each check ran. “Update the website” is too vague to serve as that contract.
Separate permission, execution and verification
Permission answers whether the action is allowed. Execution records what was attempted. Verification asks whether the intended state now exists. One success response cannot answer all three questions.
The verifier should read from the destination through a separate request. It should compare the record identity, the expected value and the revision, rather than checking only a success flag. If the public page is cached, record its state separately from the underlying record. A correct database value does not establish a correct visitor experience.
A failure worth keeping
In the library example, imagine that the update succeeds but the page retains yesterday’s hours. The agent has evidence for the record change and evidence against public completion. It should keep the task open, identify the stale layer and make a bounded repair. Repeating the update blindly risks duplicate effects and does nothing to clear the cache.
A timeout presents a different problem: the effect may have happened even though the caller received no response. Before retrying, look up the destination using the task’s identity. A retry-safe request key helps, but a key in a client log does not prove that the receiving service honoured it.
Make uncertainty useful
A useful receipt names the expected state, observed state, checked surface and remaining gap. It can say “record updated; public page still stale” without pretending the whole task succeeded. That gives the next actor a precise starting point.
This example is fictional. It illustrates a design choice, not a measured reliability result. For an actual system, evaluation should exercise the destination state, partial success, uncertain outcomes and recovery. Anthropic’s agent-evaluation guidance describes the distinction between grading outputs and examining the environment’s final state. I use that distinction to keep completion claims testable.
Sources and attribution
This is an original, self-published piece by Vihang Patel. The sources below support the referenced frameworks; fictional examples and personal judgments are identified in the text.