feat(schema): support time-of-day cost tiers - #4892
guillaumegay13 wants to merge 3 commits into
Conversation
Action items
|
|
Good catch from the review bot — this was a real hole, not a false positive. [[cost.tiers]]
input = 0.44
output = 1.32Since ~13 syncers carry Fixed in the latest commit: time tiers serialize as |
|
No actionable findings. |
|
Follow-up after an adversarial second-pass review of the branch: Verified clean: full catalog generation at merge-base vs this branch over all 364 tiered provider TOMLs produces byte-identical JSON — the output shape of existing data is provably unchanged, not just test-covered. Addressed in the latest commit: the schema didn't define what happens when a context tier and a time tier both match a request (>200k tokens during a peak window). Now documented in both the schema comment and the SDK types: the context tier wins; tiers replace the base cost and never compose. A provider charging a combined long-context peak rate can't be expressed with one tier of each kind — if that ever materializes it needs its own schema discussion. Two caveats worth recording for reviewers, no code change:
|
|
No actionable findings. |
|
ill have to merge something like this soon, openrouter has been thinking about how to model these things so ill look at them |
|
@rekram1-node happy to work on it if you have any feedback! |
They've already shipped it. It's live in I censused the full model list this morning (2026-08-22 14:42 UTC, 421 models): curl -s https://openrouter.ai/api/v1/models \
| jq '[.data[] | select(.pricing.overrides)] | length' # 60
curl -s https://openrouter.ai/api/v1/models \
| jq -r '.data[] | select(any(.pricing.overrides[]?; .utc_start)) | .id' # 160 of 421 models carry "pricing": {
"prompt": "0.00000022", "completion": "0.00000066", "input_cache_read": "0.000000007",
"overrides": [
{ "utc_start": 1000, "utc_end": 100, "prompt": "0.00000022", "completion": "0.00000066", "input_cache_read": "0.000000007" },
{ "utc_start": 100, "utc_end": 400, "prompt": "0.00000044", "completion": "0.00000132", "input_cache_read": "0.000000014" },
{ "utc_start": 400, "utc_end": 600, "prompt": "0.00000022", "completion": "0.00000066", "input_cache_read": "0.000000007" },
{ "utc_start": 600, "utc_end": 1000, "prompt": "0.00000044", "completion": "0.00000132", "input_cache_read": "0.000000014" }
]
}( Five things in there that bear on this PR. 1. The wrap is real, not hypothetical. 2. 3. Precedence is undefined upstream too. Zero models carry both kinds today ( 4. This turns caveat #2 into a fix rather than a caveat. You noted that 5. The gap neither schema covers: day of week. From midnight Beijing time on 2026-08-23 — which is 16:00 UTC today, a little over an hour from now — DeepSeek bills off-peak all day on Saturdays and Sundays. Worth noting the timezone edge if a I'd suggest not blocking this PR on that. A 2× overstatement on 6 of 168 hours a week is a much smaller error than the flat number #4891 had to ship, and it's one-directional, so it can go in |
|
@guillaumegay13 you asked for feedback — I read the branch against the live OpenRouter response and DeepSeek's own pricing page. Three notes, and the first two are validation rather than complaints. 1. The midnight wrap is required, not defensive. The only production example of time-of-day pricing on OpenRouter today is 2. "Context tier wins, tiers never compose" is safe against today's data. I censused all 421 models this afternoon — 68 override entries across 60 models: No provider currently ships a combined long-context peak rate, so the restriction in your comment costs nothing today. It's worth asserting 3. The one gap I'd flag before merge: day of week. DeepSeek's page says that from midnight Beijing time on 2026-08-23 (= 2026-08-22 16:00 UTC) both V4 models bill off-peak all day on Saturdays and Sundays. Two things make it nastier than it looks: the rule is stated in Beijing time while the windows are UTC, so "is it the weekend" is not a property of the UTC date — 2026-08-23 00:30 UTC is Beijing Sunday (off-peak) while 2026-08-24 00:30 UTC is Beijing Monday (peak). If Importer detail, if the openrouter syncer starts deriving these: OpenRouter's @rekram1-node on "openrouter has been thinking about how to model these things" — the shape they shipped is the table above: both tier kinds in one |
|
Closing this pull request as stale because it has not been updated in 30 days. Feel free to reopen it or submit a new pull request if the work is resumed. |
Adds a
timevariant tocost.tiers[], so providers that bill by time of day can be expressed as data. Follow-up to #4891, which had to flatten DeepSeek V4's peak/off-peak rates into a single number because the schema has no way to say "this rate applies between these hours".Windows are UTC
HH:MM-HH:MM, start inclusive and end exclusive; an end that precedes its start wraps past midnight (22:00-02:00). The base[cost]applies outside every window — same override relationship context tiers already have.DeepSeek is the immediate case (V4 moved to peak/off-peak on 2026-08-16), but the pattern isn't new — Chinese labs have run off-peak discount windows for a while, and this stops each one from silently becoming a wrong flat number in the catalog.
Notes on the implementation
z.union, notz.discriminatedUnion. The obvious move is a discriminated union ontier.type, but it breaks every tier already in the repo: authored context tiers omittypeand rely onz.literal("context").default("context"), and a discriminated union can't match a missing discriminator — it fails with "No matching discriminator" before the default is ever applied. A plain union keeps existing data valid, at the cost of slightly noisier error messages. There's a test pinning that behaviour.Validation. Windows must be well-formed, must not start and end at the same minute, must be non-empty, and must not overlap each other — including across separate time tiers, and including wrap-around windows. Adjacent windows (
01:00-04:00and04:00-06:00) are allowed since the end is exclusive.The duplicate-size check. It did
tiers.map((tier) => tier.tier.size)unconditionally, which yieldsundefinedfor a time tier and would fire spuriously on two of them. Now it only looks at context tiers.Legacy
context_over_200k. Bothgenerate.tsandcompare-model-migrations.tsbailed out unless there was exactly one tier, so adding a time tier to a model would have silently dropped the legacy field. They now count context tiers only — behaviour is unchanged for every model without a time tier.SDK types.
CostTier.tierbecomesContextTier | TimeTier. The drift-protection assertions inpackages/sdk/test/types.tspin these to the inferred Zod types, and they pass.Verification
bun test— 206 pass, 12 of them new. The 4 failures (repository open-weight model metadata includes weights links,DeepInfra preserves live modalities for new base models, and twosnapshotmodule-resolution failures) reproduce identically on an unmodifieddev.bun run validate— passes.tsc --noEmitinpackages/sdk— clean.No data uses the new variant yet. If this lands I'll follow up by converting the DeepSeek entries to carry both rates.