Skip to content

feat(inference): log gateway quota headers when LLM requests hit a 429 - #6736

Merged
adrian-cowham merged 1 commit into
mainfrom
ac/quota-logging-rate-limits-85d5e5
Aug 7, 2026
Merged

adrian-cowham merged 1 commit into
mainfrom
ac/quota-logging-rate-limits-85d5e5

Conversation

@adrian-cowham

Copy link
Copy Markdown
Contributor

Summary

When an inference.LLM request is rejected with a 429, customers currently get a bare APIStatusError with no indication of which limit they hit or by how much. The agent gateway now stamps quota telemetry headers on LLM completions responses (livekit/agent-gateway, X-LiveKit-Inference-*), including on rate-limit rejections — this PR surfaces that snapshot in the agent logs.

On a 429, inference.LLM now emits a structured warning before raising:

LLM request rate limited by inference gateway
  model=openai/gpt-4o request_id=req_123
  rpm_limit=100 rpm_used=101 tpm_limit=50000 tpm_used=48000
  credits_limit=1000000 credits_used=999999
  • X-LiveKit-Inference-RPM-{Limit,Used} — requests/min for the project × model bucket
  • X-LiveKit-Inference-TPM-{Limit,Used} — tokens/min
  • X-LiveKit-Inference-Credits-{Limit,Used} — cumulative token-credit balance

Only dimensions the gateway actually stamped are logged (a missing header means "not enforced / no data", never zero), so responses from older gateways just log model + request_id. The existing APIStatusError raise path is unchanged.

Changes

  • inference/_utils.py: extract_quota_usage() maps the gateway's quota headers to log-friendly fields (rpm_limit, tpm_used, …)
  • inference/llm.py: LLMStream._log_rate_limited() logs the quota snapshot on 429 before re-raising
  • tests/test_inference_utils.py: unit coverage for header extraction (all dimensions, partial, empty, case-insensitive httpx.Headers) and the 429 log path (with and without quota headers)

Testing

uv run pytest tests/test_inference_utils.py --unit   # 11 passed

ruff format / ruff check clean; make type-check shows only the pre-existing boto3 stub error in the AWS plugin.

🤖 Generated with Claude Code

The agent gateway stamps X-LiveKit-Inference-{RPM,TPM,Credits}-{Limit,Used}
on completions responses, including rate-limit rejections. Surface that
snapshot in a warning log when inference.LLM receives a 429 so customers
can see which quota they hit and by how much, instead of a bare status error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@adrian-cowham
adrian-cowham requested a review from a team as a code owner August 6, 2026 22:09

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no bugs or issues to report.

Open in Devin Review

@adrian-cowham

Copy link
Copy Markdown
Contributor Author

Question for reviewers: should we extend this to warn before the 429? The gateway stamps these quota headers on every successful completions response too, so we could log a warning when usage crosses e.g. 50% and 80% of the RPM/TPM limit — giving customers a heads-up to self-pace before requests start failing. Worth doing in this PR, or as a follow-up?

@theomonnom theomonnom left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's great!

@theomonnom

Copy link
Copy Markdown
Member

Question for reviewers: should we extend this to warn before the 429? The gateway stamps these quota headers on every successful completions response too, so we could log a warning when usage crosses e.g. 50% and 80% of the RPM/TPM limit — giving customers a heads-up to self-pace before requests start failing. Worth doing in this PR, or as a follow-up?

I'm not against it, this is a great idea too, but if we have multiple concurrent voice sessions, it could be quite spammy? vs an email we're already sending customers?

@adrian-cowham

Copy link
Copy Markdown
Contributor Author

Yup, there are supposed to be emails. I was thinking it would be nice to be in the logs and only do it 1 time per session. That is, 1 time at > 50% < 80% and then 1 time > 80%. So, at most a session would only see 2 logs.

@adrian-cowham
adrian-cowham merged commit 78a8080 into main Aug 7, 2026
6 checks passed
@adrian-cowham
adrian-cowham deleted the ac/quota-logging-rate-limits-85d5e5 branch August 7, 2026 17:07
AALG123 pushed a commit to AALG123/agents that referenced this pull request Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants