feat(inference): log gateway quota headers when LLM requests hit a 429 - #6736
Conversation
The agent gateway stamps X-LiveKit-Inference-{RPM,TPM,Credits}-{Limit,Used}
on completions responses, including rate-limit rejections. Surface that
snapshot in a warning log when inference.LLM receives a 429 so customers
can see which quota they hit and by how much, instead of a bare status error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Question for reviewers: should we extend this to warn before the 429? The gateway stamps these quota headers on every successful completions response too, so we could log a warning when usage crosses e.g. 50% and 80% of the RPM/TPM limit — giving customers a heads-up to self-pace before requests start failing. Worth doing in this PR, or as a follow-up? |
I'm not against it, this is a great idea too, but if we have multiple concurrent voice sessions, it could be quite spammy? vs an email we're already sending customers? |
|
Yup, there are supposed to be emails. I was thinking it would be nice to be in the logs and only do it 1 time per session. That is, 1 time at > 50% < 80% and then 1 time > 80%. So, at most a session would only see 2 logs. |
Summary
When an
inference.LLMrequest is rejected with a 429, customers currently get a bareAPIStatusErrorwith no indication of which limit they hit or by how much. The agent gateway now stamps quota telemetry headers on LLM completions responses (livekit/agent-gateway,X-LiveKit-Inference-*), including on rate-limit rejections — this PR surfaces that snapshot in the agent logs.On a 429,
inference.LLMnow emits a structured warning before raising:X-LiveKit-Inference-RPM-{Limit,Used}— requests/min for the project × model bucketX-LiveKit-Inference-TPM-{Limit,Used}— tokens/minX-LiveKit-Inference-Credits-{Limit,Used}— cumulative token-credit balanceOnly dimensions the gateway actually stamped are logged (a missing header means "not enforced / no data", never zero), so responses from older gateways just log model + request_id. The existing
APIStatusErrorraise path is unchanged.Changes
inference/_utils.py:extract_quota_usage()maps the gateway's quota headers to log-friendly fields (rpm_limit,tpm_used, …)inference/llm.py:LLMStream._log_rate_limited()logs the quota snapshot on 429 before re-raisingtests/test_inference_utils.py: unit coverage for header extraction (all dimensions, partial, empty, case-insensitivehttpx.Headers) and the 429 log path (with and without quota headers)Testing
ruff format/ruff checkclean;make type-checkshows only the pre-existingboto3stub error in the AWS plugin.🤖 Generated with Claude Code