Profiling a production server (Windows, bun): every web UI page load calls GET /provider, which walks the whole models.dev catalog through deep clones + per-model schema validation, JSON.stringifies ~5.5MB, then gzips it — ~700ms p50, 2s+ p95 under load (~33% of server CPU was schema validation, 13% JSON.stringify). Under concurrent load this starves unrelated endpoints (measured p95 up to 11.7s on other routes while /provider churned).
The four inputs (config, models.dev catalog, connected providers, credentials) are reference-stable between reloads, so the encoded payload can be memoized on input identity and invalidated instantly when a reload swaps any input. Same for the gzip/deflate compression of identical response bodies.
Happy to send a patch — measured p50 726ms → 22ms, throughput ~5 → ~205 req/s.
Profiling a production server (Windows, bun): every web UI page load calls GET /provider, which walks the whole models.dev catalog through deep clones + per-model schema validation, JSON.stringifies ~5.5MB, then gzips it — ~700ms p50, 2s+ p95 under load (~33% of server CPU was schema validation, 13% JSON.stringify). Under concurrent load this starves unrelated endpoints (measured p95 up to 11.7s on other routes while /provider churned).
The four inputs (config, models.dev catalog, connected providers, credentials) are reference-stable between reloads, so the encoded payload can be memoized on input identity and invalidated instantly when a reload swaps any input. Same for the gzip/deflate compression of identical response bodies.
Happy to send a patch — measured p50 726ms → 22ms, throughput ~5 → ~205 req/s.