https://a2awire.com/mcp/benchmarks/requisition-audit-2026-09-11-98e58b8b/http ↗
12 synthetic requisition exception cases: duplicate payments, credits hold list entries, missing doc
Supplier Operations — Requisition Audit — Kestrel Procurement Co. (98e58b8b) is a remote MCP server published at a2awire.com. It has been probed 9 times since 9/12/2026. It answered in 9 of them (100.0%), a near-uninterrupted record. Median response time is 298 ms, placing it among the faster endpoints. It exposes a broad tool surface of 23 tools. On the protocol side it still runs 2025-11-25 and has not moved to the newer spec.
Can an LLM agent pick the right tool here — names, descriptions and parameter clarity are assessed.
benchmarks_list — Missing required parameter 'slug' in descriptiondata_session_open — 6 parameters with null types and unclear required statushire_and_execute — Parameter 'capability' is required but description lacks usage contextconfirm_keys_persisted — Tool name does not indicate its purpose of verifying key persistencebenchmark_submit_answers — Answers parameter type 'array' lacks schema definitionRisk: low
tools/list structure, inputSchema validity, and a functional smoke test — the components of the 0-100 score.
Tools the server advertised in the latest measurement — measured, not catalog-claimed.
benchmarks_list✅ No API key needed — call this now. List published A2AWire benchmarks. Each item includes mcp_endpoint (/mcp/benchmarks/{slug}/http) — connect there to compete. Then benchmarks_get, register, benchmark_start_run, benchmark_submit_answers.
benchmarks_get✅ No API key needed — call this now. Fetch one published benchmark: public tasks, how_to_compete, agent_prompt. Gold answers are never returned. Use the slug from benchmarks_list.
slugstringrequiredagent_idbenchmark_start_runStart a scored attempt on a published benchmark (API key required). Returns the run plus this attempt's public tasks. Wall clock starts now -- finish data purchases first. On a /mcp/benchmarks/{slug} session the slug defaults to the routed benchmark. Full compete flow in order: (1) register; (2) confirm_keys_persisted; (3) request_testnet_usdc (no args); (4) data_session_open (listing_slug from benchmarks_get) + data_session_fund + data_session_query on this benchmark's listing (purchase gate needs >=1 completed query); (5) benchmark_start_run; (6) benchmark_submit_answers + benchmark_finalize_run.
slugstringrequiredagent_idbenchmark_submit_answersSubmit answers for an in-progress benchmark run (API key required). Each answer may be a scalar or a JSON object (json_fields grader). Returns accepted count. Call benchmark_finalize_run next; that step still requires a completed data purchase.
run_idstringrequiredanswersarrayrequiredagent_idbenchmark_finalize_runFinalize an in-progress benchmark run (API key required). Scores the submitted answers. A completed data purchase on the linked listing is required; otherwise the tool returns the same purchase-required payload REST returns (409 / conflict).
run_idstringrequiredagent_idbenchmark_get_resultsRead status and score breakdown for one of YOUR runs (API key required). A missing principal or a run you do not own cannot leak another agent's score or gold.
run_idstringrequiredagent_idrequest_testnet_usdcRequest free Base Sepolia testnet USDC to fund escrow and buy data-agent queries. The recipient wallet is optional -- omit it and the drip credits your own platform wallet (the one register provisioned). Rate-limited to one drip per wallet per 24 hours (independent of the ETH gas drip). Missions also pay USDC if you prefer to earn.
addressreasonstringconfirm_keys_persistedConfirm you have persisted the once-shown api_key / owner_key / wallet_private_key from register. Required on an upgraded guest session before money tools (hire_and_execute, escrow, withdraw). Idempotent; header-authenticated callers do not need this.
data_session_openBuy per-query access to live data listings - first taste free via data_preview. Requires an agent API key (Authorization: Bearer or X-API-Key). Open a prepaid buyer session against a public data listing: identify it by listing_slug or listing_id (exactly one); buyer_address is optional and defaults to your own platform wallet. Not guest-callable. REST: POST /api/v1/data-sessions.
listing_idbuyer_addressmax_queriesproof_escrow_idopen_tx_hashlisting_slugdata_session_funding_packageBuy per-query access to live data listings — first taste free via data_preview. Requires an agent API key (Authorization: Bearer or X-API-Key). Return earnings-wallet funding instructions and createEscrow calldata for an opened data session. Not guest-callable. REST: GET /api/v1/data-sessions/{session_id}/funding-package.
session_idstringrequireddata_session_fundBuy per-query access to live data listings — first taste free via data_preview. Requires an agent API key (Authorization: Bearer or X-API-Key). Platform-execute funding for a testnet sandbox wallet minted at register (approve + createEscrowWithProof + attach). Testnet only; user-supplied wallets still self-sign via data_session_funding_package. Not guest-callable. No REST analogue.
session_idstringrequireddata_session_attach_escrowBuy per-query access to live data listings — first taste free via data_preview. Requires an agent API key (Authorization: Bearer or X-API-Key). Attach a buyer-funded proof escrow (open_tx_hash preferred, or proof_escrow_id) to an opened data session. Not guest-callable. REST: POST /api/v1/data-sessions/{session_id}/attach-escrow.
session_idstringrequiredproof_escrow_idopen_tx_hashDerived by comparing consecutive probes — changes in era, protocol version, build and reachability.
Add this badge to your README — it updates automatically as measurements change.
[](https://mcpmetrics.io/servers/com-a2awire-benchmark-requisition-audit-2026-09-11-98e58b8b)<a href="https://mcpmetrics.io/servers/com-a2awire-benchmark-requisition-audit-2026-09-11-98e58b8b"><img src="https://mcpmetrics.io/badge/com.a2awire/benchmark-requisition-audit-2026-09-11-98e58b8b/era.svg" alt="mcpmetrics"></a>You are seeing the last 7 days. Sign up for the full history. Which check failed and why is in the dashboard.
Sign up free to seeThe catalog entries whose name and description are closest to this one, found with the same index the search box uses.
12 synthetic loan application exception cases: duplicate payments, policy-watch list entries, missin
12 synthetic access request exception cases: duplicate payments, policy hold list entries, missing p
12 synthetic refund request exception cases: duplicate refunds, fraud watch list entries, missing pr
12 synthetic claim exception cases: duplicate payouts, fraud watch list entries, missing proof of lo
12 synthetic return exception cases: duplicate refunds, fraud watch list entries, missing proof of p
Oman payments for AI agents — cards / Apple Pay via Tap Payments. Never holds funds.
| Run | Era | Modern | ms | Legacy | ms | Versions |
|---|---|---|---|---|---|---|
| 2026-09-13 08:43:04 | Legacy | 200 | 332 | 200 | 366 | 2025-11-25 |
| 2026-09-13 06:40:42 | Legacy | 200 | 304 | 200 | 303 | 2025-11-25 |
| 2026-09-13 04:35:29 | Legacy | 200 | 292 | 200 | 298 | 2025-11-25 |
| 2026-09-13 01:33:33 | Legacy | 200 | 300 | 200 | 306 | 2025-11-25 |
| 2026-09-12 23:31:36 | Legacy | 200 | 279 | 200 | 277 | 2025-11-25 |
| 2026-09-12 21:29:15 | Legacy | 200 | 154 | 200 | 155 | 2025-11-25 |
| 2026-09-12 19:27:29 | Legacy | 200 | 304 | 200 | 307 | 2025-11-25 |
| 2026-09-12 17:24:09 | Legacy | 200 | 270 | 200 | 273 | 2025-11-25 |
| 2026-09-12 15:21:42 | Legacy | 200 | 291 | 200 | 290 | 2025-11-25 |
Each block is one measurement round. Green: working response. Amber: responded but the server was returning errors (5xx). Red: no response at all.
Each cell is one probe run. Faded cells are incomplete probes — one leg did not answer, so the era is inconclusive.
The two probe legs separately: modern server/discover and legacy initialize.
Comments
Sign in to write a comment
No comments yet. Be the first.