Lists built from measurement. Each one is computed from the last 30 days of observations; a server needs at least 20 measurement rounds to appear.
Ranked by median response time. Median, not mean: a single slow round distorts an average.
Share of rounds the server answered in the last 30 days. Ties break toward more measurement.
How clearly a model can read the server’s tool names, descriptions and schemas.
Share of rounds whose tool schema validated. Servers whose schema could never be read are not listed.
The highest tool count seen across measured rounds. More tools is coverage, not quality.
Servers speaking the modern transport, ranked by availability.