Skip to content

feat(schema): support time-of-day cost tiers - #4892

Closed
guillaumegay13 wants to merge 3 commits into
anomalyco:devfrom
guillaumegay13:feat/time-cost-tiers
Closed

guillaumegay13 wants to merge 3 commits into
anomalyco:devfrom
guillaumegay13:feat/time-cost-tiers

Conversation

@guillaumegay13

Copy link
Copy Markdown

Adds a time variant to cost.tiers[], so providers that bill by time of day can be expressed as data. Follow-up to #4891, which had to flatten DeepSeek V4's peak/off-peak rates into a single number because the schema has no way to say "this rate applies between these hours".

[cost]
input = 0.22
output = 0.66
cache_read = 0.007

[[cost.tiers]]
tier = { type = "time", windows = ["01:00-04:00", "06:00-10:00"] }
input = 0.44
output = 1.32
cache_read = 0.014

Windows are UTC HH:MM-HH:MM, start inclusive and end exclusive; an end that precedes its start wraps past midnight (22:00-02:00). The base [cost] applies outside every window — same override relationship context tiers already have.

DeepSeek is the immediate case (V4 moved to peak/off-peak on 2026-08-16), but the pattern isn't new — Chinese labs have run off-peak discount windows for a while, and this stops each one from silently becoming a wrong flat number in the catalog.

Notes on the implementation

z.union, not z.discriminatedUnion. The obvious move is a discriminated union on tier.type, but it breaks every tier already in the repo: authored context tiers omit type and rely on z.literal("context").default("context"), and a discriminated union can't match a missing discriminator — it fails with "No matching discriminator" before the default is ever applied. A plain union keeps existing data valid, at the cost of slightly noisier error messages. There's a test pinning that behaviour.

Validation. Windows must be well-formed, must not start and end at the same minute, must be non-empty, and must not overlap each other — including across separate time tiers, and including wrap-around windows. Adjacent windows (01:00-04:00 and 04:00-06:00) are allowed since the end is exclusive.

The duplicate-size check. It did tiers.map((tier) => tier.tier.size) unconditionally, which yields undefined for a time tier and would fire spuriously on two of them. Now it only looks at context tiers.

Legacy context_over_200k. Both generate.ts and compare-model-migrations.ts bailed out unless there was exactly one tier, so adding a time tier to a model would have silently dropped the legacy field. They now count context tiers only — behaviour is unchanged for every model without a time tier.

SDK types. CostTier.tier becomes ContextTier | TimeTier. The drift-protection assertions in packages/sdk/test/types.ts pin these to the inferred Zod types, and they pass.

Verification

  • bun test — 206 pass, 12 of them new. The 4 failures (repository open-weight model metadata includes weights links, DeepInfra preserves live modalities for new base models, and two snapshot module-resolution failures) reproduce identically on an unmodified dev.
  • bun run validate — passes.
  • tsc --noEmit in packages/sdk — clean.

No data uses the new variant yet. If this lands I'll follow up by converting the DeepSeek entries to carry both rates.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] packages/core/src/sync/index.ts:950 - Check: New cost.tiers variants must round-trip through sync TOML serialization. Why: formatToml only writes a tier header when tier.tier.size is set, so a type = "time" tier drops type/windows and is rewritten as bare prices. Any provider that preserves existing?.cost?.tiers (OpenRouter, Kilo, Anthropic, DigitalOcean, etc.) will corrupt hand-authored peak/off-peak data on the next sync. Action: Serialize time tiers (e.g. tier = { type = "time", windows = [...] }), keep context-tier output unchanged, and add a formatToml regression test that round-trips a time tier.

@guillaumegay13

Copy link
Copy Markdown
Author

Good catch from the review bot — this was a real hole, not a false positive. formatToml only wrote the tier = { ... } header when tier.tier.size was set, so a time tier round-tripped as a bare [[cost.tiers]] block with prices and no discriminator:

[[cost.tiers]]
input = 0.44
output = 1.32

Since ~13 syncers carry existing?.cost?.tiers forward (OpenRouter, Kilo, Anthropic, DigitalOcean, and others), any hand-authored time tier on an auto-synced provider would have been destroyed on the next hourly run. DeepSeek itself has no syncer, so the motivating case wasn't exposed — but the resellers that serve the same models are.

Fixed in the latest commit: time tiers serialize as tier = { type = "time", windows = [...] }, context-tier output is byte-identical to before, and there's a round-trip regression test through Bun.TOML.parse. I verified the test fails without the writer fix and passes with it.

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Aug 17, 2026
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Aug 17, 2026
@guillaumegay13

Copy link
Copy Markdown
Author

Follow-up after an adversarial second-pass review of the branch:

Verified clean: full catalog generation at merge-base vs this branch over all 364 tiered provider TOMLs produces byte-identical JSON — the output shape of existing data is provably unchanged, not just test-covered.

Addressed in the latest commit: the schema didn't define what happens when a context tier and a time tier both match a request (>200k tokens during a peak window). Now documented in both the schema comment and the SDK types: the context tier wins; tiers replace the base cost and never compose. A provider charging a combined long-context peak rate can't be expressed with one tier of each kind — if that ever materializes it needs its own schema discussion.

Two caveats worth recording for reviewers, no code change:

  1. deepPartial (zod 3.24) doesn't recurse into ZodUnion, so in the sync-side ExistingModel the fields inside tier are now effectively required rather than deep-partial. No on-disk TOML has a partial tier header (checked), so nothing breaks today — but a hypothetical size-less tier = { type = "context" } file would now fail the sync parse rather than pass it.
  2. Syncers that rebuild tiers wholesale from API data (openrouter, crossmodel, deepinfra) would drop a hand-authored time tier on their models on the next sync. The pass-through syncers (anthropic, kilo, nano-gpt, etc.) preserve them, and DeepSeek's own provider has no syncer, so the motivating case is unaffected. If time tiers ever get authored on an aggregator's models, those syncs need the same preserve-authored-tiers treatment.

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@rekram1-node

Copy link
Copy Markdown
Collaborator

ill have to merge something like this soon, openrouter has been thinking about how to model these things so ill look at them

@guillaumegay13

Copy link
Copy Markdown
Author

@rekram1-node happy to work on it if you have any feedback!

@xyzs996

xyzs996 commented Aug 22, 2026 •

Copy link
Copy Markdown
Contributor

openrouter has been thinking about how to model these things so ill look at them

They've already shipped it. It's live in /api/v1/models right now, and the shape is close enough to this PR that it's worth putting the actual wire format in the thread rather than waiting on it.

I censused the full model list this morning (2026-08-22 14:42 UTC, 421 models):

curl -s https://openrouter.ai/api/v1/models \
| jq '[.data[] | select(.pricing.overrides)] | length'                                    # 60
curl -s https://openrouter.ai/api/v1/models \
| jq -r '.data[] | select(any(.pricing.overrides[]?; .utc_start)) | .id'                  # 1

60 of 421 models carry pricing.overrides. 59 of them use min_prompt_tokens. Exactly one uses time of day, and it's the model this PR exists for:

"pricing": {
  "prompt": "0.00000022", "completion": "0.00000066", "input_cache_read": "0.000000007",
  "overrides": [
    { "utc_start": 1000, "utc_end":  100, "prompt": "0.00000022", "completion": "0.00000066", "input_cache_read": "0.000000007" },
    { "utc_start":  100, "utc_end":  400, "prompt": "0.00000044", "completion": "0.00000132", "input_cache_read": "0.000000014" },
    { "utc_start":  400, "utc_end":  600, "prompt": "0.00000022", "completion": "0.00000066", "input_cache_read": "0.000000007" },
    { "utc_start":  600, "utc_end": 1000, "prompt": "0.00000044", "completion": "0.00000132", "input_cache_read": "0.000000014" }
  ]
}

(deepseek/deepseek-v4-flash-vision-exp. Same numbers DeepSeek publishes: peak 01:00–04:00 and 06:00–10:00 UTC, everything else off-peak.)

Five things in there that bear on this PR.

1. The wrap is real, not hypothetical. utc_start: 1000, utc_end: 100 is 10:00 through 01:00 — it crosses midnight. Your "an end that precedes its start wraps past midnight" rule is the only reading under which the first entry means anything, and the four windows then sum to exactly 1440 minutes. Under a naive start <= t < end the first window is dead and 15 hours of the day fall through to base. So the one production example of this field in existence requires the rule you already picked.

2. min_prompt_tokens and utc_start live in the same array. Not two lists, not a discriminator field — one overrides[], each entry carrying whichever keys apply. That maps onto cost.tiers[] with tier.type rather than onto a separate cost.time_windows, which is the choice this PR made.

3. Precedence is undefined upstream too. Zero models carry both kinds today (both kinds: 0 in the census), so OpenRouter has not had to answer "long context during a peak window" either. Your "context tier wins, tiers never compose" is a decision you get to make; it isn't in conflict with anything shipped.

4. This turns caveat #2 into a fix rather than a caveat. You noted that openrouter/crossmodel/deepinfra rebuild tiers wholesale and would drop a hand-authored time tier. For the openrouter syncer specifically that's now avoidable — the windows are in the API response, so the syncer can derive the time tier instead of preserving or dropping it. The conversion is small but has one trap: these are HHMM integers, so 100 is 01:00 and 1000 is 10:00. Anything that stringifies before slicing needs String(v).padStart(4, "0"), and utc_end: 0 (midnight) has to survive the same path.

5. The gap neither schema covers: day of week. From midnight Beijing time on 2026-08-23 — which is 16:00 UTC today, a little over an hour from now — DeepSeek bills off-peak all day on Saturdays and Sundays. windows = ["01:00-04:00", "06:00-10:00"] has no way to say that, and neither does OpenRouter's overrides (their entry above still shows the weekday split). So the first hours where every catalog carrying V4 peak rates is wrong are 2026-08-23 01:00–04:00 and 06:00–10:00 UTC, and it's a clean 2× over-estimate rather than a random error.

Worth noting the timezone edge if a weekdays field ever gets added: the rule is stated in Beijing time while the windows are UTC, so "is it the weekend" is not a property of the UTC date. 2026-08-23 00:30 UTC is Sunday in Beijing and bills off-peak; 2026-08-24 00:30 UTC is Monday in Beijing and bills peak. Same UTC hour, same UTC-weekend-looking timestamp, different price.

I'd suggest not blocking this PR on that. A 2× overstatement on 6 of 168 hours a week is a much smaller error than the flat number #4891 had to ship, and it's one-directional, so it can go in notes and be fixed later without any data being wrong in a new way. Flattening peak and off-peak into one number, which is the status quo, is wrong 24 hours a day.

@xyzs996

xyzs996 commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

@guillaumegay13 you asked for feedback — I read the branch against the live OpenRouter response and DeepSeek's own pricing page. Three notes, and the first two are validation rather than complaints.

1. The midnight wrap is required, not defensive. The only production example of time-of-day pricing on OpenRouter today is deepseek/deepseek-v4-flash-vision-exp, and its first window is utc_start: 1000, utc_end: 100 — 10:00 through 01:00, wrapping. So toIntervals splitting a wrap into [start, 1440] and [0, end] isn't a hypothetical case; drop it and the one model that exercises this feature can't be authored. (Its four windows sum to exactly 1440 minutes, and its base rate equals the off-peak rate, which is also the shape your TimeTier comment assumes: peak in the tier, off-peak in [cost].)

2. "Context tier wins, tiers never compose" is safe against today's data. I censused all 421 models this afternoon — 68 override entries across 60 models:

entries with min_prompt_tokens only : 64   (context-tiered)
entries with utc_start/utc_end only :  4   (time-windowed)
entries with both                   :  0
entries with neither                :  0

No provider currently ships a combined long-context peak rate, so the restriction in your comment costs nothing today. It's worth asserting both === 0 in the syncer rather than trusting it, though — that assertion is the fail-loud signal for the day someone does, and it fires on a genuinely new pricing dimension instead of on every routine new field.

3. The one gap I'd flag before merge: day of week. DeepSeek's page says that from midnight Beijing time on 2026-08-23 (= 2026-08-22 16:00 UTC) both V4 models bill off-peak all day on Saturdays and Sundays. windows: string[] can't say that, so every weekend request resolves to the peak rate — an upper bound, 2x too high on 6 of 168 hours a week. It's one-directional and bounded, so it's fine to ship and document, but it should be a stated decision rather than something discovered later.

Two things make it nastier than it looks: the rule is stated in Beijing time while the windows are UTC, so "is it the weekend" is not a property of the UTC date — 2026-08-23 00:30 UTC is Beijing Sunday (off-peak) while 2026-08-24 00:30 UTC is Beijing Monday (peak). If days ever gets added, it needs a timezone alongside it, not just a weekday list.

Importer detail, if the openrouter syncer starts deriving these: OpenRouter's utc_start/utc_end are HHMM integers, not strings — 100 is 01:00 and 1000 is 10:00. Anything that stringifies before slicing needs String(v).padStart(4, "0"), and utc_end: 0 (00:00) has to survive the same path rather than being treated as absent. Your HH:MM-HH:MM string form is the better authored shape; it's the conversion that has the trap in it.

@rekram1-node on "openrouter has been thinking about how to model these things" — the shape they shipped is the table above: both tier kinds in one pricing.overrides array, discriminated by which keys an entry carries, no explicit type field. Full census with the raw JSON is in my earlier comment if it's useful for the syncer side.

@github-actions

Copy link
Copy Markdown
Contributor

Closing this pull request as stale because it has not been updated in 30 days. Feel free to reopen it or submit a new pull request if the work is resumed.

@github-actions github-actions Bot closed this Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants