https://mcp-selection-lab-production.up.railway.app/mcp ↗
Benchmark MCP tool selection with metadata-only routing, collision, abstention, and holdout checks.
MCP Selection Lab is a remote MCP server published at mcp-selection-lab-production.up.railway.app. It has been probed 8 times since 9/12/2026. It answered in only 3 of them (37.5%), indicating low availability. Median response time is 15,002 ms, a delay an agent will notice. It offers a narrow, focused set of 3 tools. On the protocol side it speaks the 2026-07-28 stateless spec.
Can an LLM agent pick the right tool here — names, descriptions and parameter clarity are assessed.
selection_lab_info — Name doesn't clearly indicate metadata-only information retrievalscan_mcp_metadata — 'internal_test' parameter lacks clear purpose explanationget_selection_report — Description doesn't clarify when to use existing reports vs new scans| Run | Era | Modern | ms |
|---|
tools/list structure, inputSchema validity, and a functional smoke test — the components of the 0-100 score.
Tools the server advertised in the latest measurement — measured, not catalog-claimed.
selection_lab_infoGet free service capabilities, limits, privacy and connection details; does not scan a target.
scan_mcp_metadataCreate a free public report of MCP tool-description routing collisions. Reads discovery metadata (initialize/tools/list) only; never calls target business tools or spends funds. Use a public HTTP(S) URL without credentials, query or fragment. Results use at most 12 tools and 24 generated cases, not a real-model evaluation. The returned report is enriched with agent_plan candidate description-only edits and a machine-readable rerun instruction. Candidate edits are not untouched-holdout proof. Reports persist and are public by link. For an existing report use get_selection_report. Set internal_test=true for owner/CI validation so it is excluded from public scans. source is a self-reported referral bucket, not identity verification.
mcp_urlstringrequiredsourcestringinternal_testbooleanget_selection_reportRead and agent-enrich an existing public scan report by id; no new target request or model call.
report_idstringrequiredDerived by comparing consecutive probes — changes in era, protocol version, build and reachability.
Add this badge to your README — it updates automatically as measurements change.
[](https://mcpmetrics.io/servers/io-github-changhuliu-mcp-selection-lab)<a href="https://mcpmetrics.io/servers/io-github-changhuliu-mcp-selection-lab"><img src="https://mcpmetrics.io/badge/io.github.ChanghuLiu/mcp-selection-lab/era.svg" alt="mcpmetrics"></a>You are seeing the last 7 days. Sign up for the full history. Which check failed and why is in the dashboard.
Sign up free to seeThe catalog entries whose name and description are closest to this one, found with the same index the search box uses.
Mineral intelligence API: 8 critical minerals (USGS benchmarks) with mineral processing and ultrasound run-log data. $0.10 USDC on Base via x402. Every response carries an on-chain EAS attestation.
Read-only tools for finding where a small business leaks deals, time, and cash.
Read-only Get2Great magazine essays on who is good at a job, and the management-tool catalog.
Safety execution proxy protecting clients during tool calls.
Public record of one read only walk over the official MCP registry: who answered, tools, who pays.
Link-preview metadata and clean page-to-Markdown for any public URL. No install.
| Legacy |
|---|
| ms |
|---|
| Versions |
|---|
| 2026-09-13 06:40:42 | Dual-era | 200 | 196 | 200 | 188 | 2026-07-28 |
| 2026-09-13 04:35:29 | Dual-era | 200 | 190 | 200 | 192 | 2026-07-28 |
| 2026-09-13 01:33:33 | Dual-era | 200 | 379 | 200 | 380 | 2026-07-28 |
| 2026-09-12 23:31:36 | Unreachable * | — | 15003 | — | 15002 | — |
| 2026-09-12 21:29:15 | Unreachable * | — | 15003 | — | 15002 | — |
| 2026-09-12 19:27:29 | Unreachable * | — | 15003 | — | 15002 | — |
| 2026-09-12 17:24:09 | Unreachable * | — | 15002 | — | 15001 | — |
| 2026-09-12 15:21:42 | Unreachable * | — | 15005 | — | 15004 | — |
Each block is one measurement round. Green: working response. Amber: responded but the server was returning errors (5xx). Red: no response at all.
Each cell is one probe run. Faded cells are incomplete probes — one leg did not answer, so the era is inconclusive.
The two probe legs separately: modern server/discover and legacy initialize.
Comments
Sign in to write a comment
No comments yet. Be the first.