mcpbench GitHub

client/mrtr-multi-round

anthropic/claude-opus-4-8:medium 2026-07-28 docs: full score 0% must-checkpoint musts 7/8 · shoulds 0/0 · mays 0/0 works yes · conformant no

Checkpoints

MUST (any failure zeroes the score — 7/8 passed)

Checkpoint Description Detail
tool-call-args Called provision_account with plan "team"
final-complete A retry carrying inputResponses received the complete result
round1-input-correlated First retry carried an accepted ElicitResult for org (org=acme)
round2-input-correlated Second retry carried an accepted ElicitResult for region (region=eu-west)
state-rotation-echo Each retry echoed the most recently issued requestState byte-for-byte
retry-args-stable-across-rounds Both retries re-issued the original request (same name and arguments)
all-ids-distinct The initial call and both retries used pairwise-distinct JSON-RPC ids tools/call provision_account reused JSON-RPC id 1 across attempts
stateless-meta Every request carried the required _meta fields

Usage & cost

Nominal cost
$0.560
Input tokens
42
Output tokens
5,747
Cache read
439,478
Duration
127s
Timestamp
2026-07-18 11:05:03Z
Bench version
0.3.0

Cost is nominal (public API pricing): token usage × published rates; actual marginal cost is subscription-covered.