Skip to content

feat(examples): Baseten Model APIs recipe on traceai-openai (TH-8310) - #234

Open
nik13 wants to merge 5 commits into
devfrom
feat/th-8310-baseten
Open

nik13 wants to merge 5 commits into
devfrom
feat/th-8310-baseten

Conversation

@nik13

@nik13 nik13 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR adds a Baseten Model APIs recipe at python/examples/baseten/ (Linear TH-8310, parent TH-8103). Customers who call Baseten's OpenAI-compatible Chat Completions endpoint (https://inference.baseten.co/v1) with the official openai SDK get traces through the existing traceai-openai instrumentor. It adds no package and no instrumentor (decision D1), and changes nothing outside python/examples/baseten/.

  • src/app.py: calls register() and instruments with OpenAIInstrumentor before creating the client. The Baseten key goes only to the OpenAI client; the Future AGI keys go only to the tracer. check_base_url() refuses two Baseten URLs that need a different recipe and never rewrites a URL:

    • the Anthropic Messages beta root, https://inference.baseten.co with no /v1;
    • dedicated-deployment hosts, model-*.api.baseten.co.

    main() returns 2 before tracing or any network call when it refuses a URL.

  • requirements.txt: the tested pins openai==3.24.0, traceAI-openai==0.1.10 and fi-instrumentation-otel==1.1.0.

  • README.md: install, configure (which key goes where), run, code, what you see in Future AGI, Baseten specifics (x-session-affinity is not a Future AGI session id; using_session is), privacy, limits, test command and tested versions.

  • tests/: a loopback contract on the shared harness Receiver (test(harness): shared OTLP harness for TH-8103 contract tests #203), plus a socket guard for subprocesses and a fake OpenAI server.

What the tests pin

Every test runs offline. In-process tests use httpx.MockTransport at the documented host, and subprocess tests use a fake on 127.0.0.1 under a guard that refuses non-loopback DNS and connects. The guard has a negative control.

Request:

  • the request goes to exactly https://inference.baseten.co/v1/chat/completions;
  • it carries Authorization: Bearer <Baseten key> and no Future AGI header.

Exported span:

  • exactly one ChatCompletion LLM span, exported to /tracer/v1/traces with X-Api-Key/X-Secret-Key, the project_name resource and project_type=observe;
  • gen_ai.request.model is the model id the response returns;
  • token usage is copied when present and omitted entirely, not set to 0, when the response has none;
  • gen_ai.provider.name is openai (D2). The instrumentor labels every non-OpenAI host this way, and the test pins it so that a later shared fix shows up;
  • the Baseten key appears nowhere in span attributes, events, status, resource or OTLP request headers.

Behaviour:

  • FI_HIDE_INPUTS/FI_HIDE_OUTPUTS remove a unique marker from the export, and a control run without hiding finds it;
  • using_session("s-1") sets session.id, and x-session-affinity does not;
  • a 401 produces openai.AuthenticationError, a span with ERROR status and an exception event, with no key in the recorded text;
  • the app runs end to end in a subprocess, chat and --stream, with the guard log empty and a sentinel showing the guard was installed;
  • refused URLs, including trailing-dot, upper-case, userinfo and port spellings of the refused hosts, raise ValueError without being rewritten; main() exits 2 before a client or tracer is created.

Current traceai-openai behaviour, pinned and stated in the README (D-F5):

  • streamed and failed calls produce spans with no gen_ai.request.model, because the instrumentor takes the model only from a non-streamed response. The requested model is still inside gen_ai.request.parameters;
  • the default stream fixture has no usage, so it exports no usage attributes.

The shared fix is tracked separately (Linear TH-8402) and is not copied into this recipe.

Tests (exact head 34ed9de8daab7fde731450ac106859e7a5316d34, clean tree, actual exit codes + JUnit)

Cell Python openai traceAI packages Result
1–4 3.10 / 3.11 / 3.12 / 3.13 3.24.0 repository source 43 passed each, rc 0
5 3.11 1.69.0 (floor) repository source 43 passed, rc 0
6 3.11 3.24.0 published traceAI-openai==0.1.10 + fi-instrumentation-otel==1.1.0 (imported from site-packages, no repo source on the path) 43 passed, rc 0

The command is in the README. No files outside python/examples/baseten/ changed, and the head did not move during the run. The first commit, 0141574, had the same 6/6 matrix with 32 tests.

Decisions (approved: Nikhil 2026-10-03 blanket)

  • D1: a recipe on traceai-openai, not a traceai-baseten package.
  • D2: the provider label stays openai. A host-to-provider mapping would be one shared change in traceai-openai, not a per-logo change.
  • D-F3: refuse the documented wrong or out-of-scope URLs; never normalize them.
  • D-F5: pin the missing model on streamed and failed spans here, and fix it once in traceai-openai (TH-8402).
  • The vendor docs (https://docs.baseten.co/inference/model-apis/overview) were re-checked on 2026-10-06; the base URL and BASETEN_API_KEY are unchanged.

Not covered

  • Dedicated deployments, the Anthropic Messages beta and Anthropic SDK calls, the Baseten CLI, and historical import.
  • No live Baseten call was made (no key, no spend). Live model behaviour, fi-collector auth and storage, and the trace view are not exercised here.
  • CI: none configured on this repository for this path.

Stacked on #203 (shared OTLP test harness, 3eaadc8), which must land first. Not for merge until review is complete.

Review status

  • Review r1 (pr-reviewer t_6453b6cd, claude-opus-5-5, at 0141574): APPROVED, with no P1 or P2 findings. Six P3 items were raised. Fixed in 34ed9de:

    • R1: a trailing-dot host bypassed both refusals. The tests failed before the fix (4 RED) and pass after.
    • R2: test_hide_outputs had no non-hidden control.
    • R3: the README now names gen_ai.request.model and notes the requested model stays in gen_ai.request.parameters.
    • R4: the requirements.txt comment mentioned a worktree.
    • R6: a vacuous fake-server check is replaced by a client-constructor spy.

    Kept as P3: R5, optional app hardening (FI key check, a force_flush() result warning). The reviewer's context notes also stand: recipe tests are not wired into CI, and register()'s signal handlers exit 0.

  • Verification r1 (pr-verifier t_dab442fa, claude-opus-5-5, at 34ed9de): VERIFIED, with no blocking findings. R1–R4 and R6 are confirmed fixed, and R5 is accepted as P3. New P3 items, kept as follow-ups with no third fix round:

    • V1: check_base_url is a best-effort guard. It compares the URL as urllib parses it, so dot-segment spellings (/v1/.., /.) and IDNA label separators can still reach a refused surface after httpx normalizes the URL. It is a guard, not a security boundary.
    • V2: no fixture returns a response model that differs from the request model, so the README's "the model id the provider returns" is source-derived rather than test-pinned.
    • V3: the published-wheel import origin is recorded by Rick's matrix runner, not asserted inside the suite.
    • F4: the requested model in gen_ai.request.parameters is dropped when FI_HIDE_LLM_INVOCATION_PARAMETERS=true.

    Context notes from the verifier:

    • the OpenAI SDK reads OPENAI_ORG_ID/OPENAI_PROJECT_ID/OPENAI_CUSTOM_HEADERS from the environment and would send them to Baseten;
    • a final stream chunk with an explicit "usage": null may raise inside the shared accumulator (trigger unverified; noted on TH-8402);
    • the recipe tests are not wired into CI.

Video demo

A narrated CLI walkthrough, 6:21, recorded at the verified head 34ed9de (1080p H.264/AAC, burned captions, 10 embedded chapters). The video, captions, transcript, chapters, preview and both media-verification notes are private attachments on the Linear issue: TH-8310 (Future AGI workspace access required). Every command chapter runs ./run_demo.sh <chapter> live in a real terminal and shows [exit 0] on screen.

Time Chapter
00:00 Problem: Baseten through the OpenAI SDK (card)
00:49 head: PR #234 at 34ed9de, the six recipe files and the env variable names
01:20 run: chat and stream against the local fake; what Future AGI and the provider each received
02:30 host: the real documented base URL through MockTransport, 0 external DNS
03:05 refusals: eight URLs refused (exit 2, 0 requests, 0 spans), two returned unchanged
03:57 error: a 401 from the provider; ERROR span, model in request parameters, no key exported
04:24 privacy: FI_HIDE_INPUTS / FI_HIDE_OUTPUTS versus control; the provider still receives the prompt
04:49 tests: LIVE Python 3.11 suite (43 passed) and the RECORDED six-cell exact-head matrix
05:24 limits
05:43 Review status (card)

Not shown: a live Baseten call, a running fi-collector, authentication, storage or the trace view. All runs use loopback fakes, and the narration says so. Static waits of 2 s or more were cut; nothing was sped up and no output was edited.

nik13 added 5 commits October 3, 2026 18:58
The Node OTLP exporter streams with Transfer-Encoding: chunked and sends
no Content-Length. The Receiver read Content-Length only, so every Node
export came back 400 with zero spans; two TH-8103 children had to put a
de-chunking relay in front of it. Receiver now accepts both framings.

It also records one entry per accepted export (path, lower-cased
headers, flattened resource attributes) via requests(), so a contract
test can assert X-Api-Key/X-Secret-Key and project_name through the
shared harness instead of a private recorder.

Verified: 6 harness tests pass, and the TanStack example's real Node
exporter delivered 4 chunked exports (4 spans, collector path, both
auth headers, project_name/project_type) with no relay.

Refs: TH-8339, TH-8103
Add python/examples/baseten: trace Baseten Model APIs Chat Completions made
with the official openai SDK (base_url https://inference.baseten.co/v1) by
using the existing traceai-openai instrumentor. No Baseten package.

- src/app.py: register() + OpenAIInstrumentor before the client; the
  Baseten key goes only to the OpenAI client, Future AGI keys only to the
  tracer. check_base_url() refuses the Anthropic-beta root (no /v1) and
  dedicated-deployment hosts (model-*.api.baseten.co); it never rewrites.
- tests: loopback contract on the shared harness Receiver: request URL and
  bearer key at the documented host (MockTransport), one LLM span, model,
  usage (omitted when absent), provider label openai (D2), project resource
  and FI headers, no vendor key in any exported field, streaming, 401 error
  span, FI_HIDE_INPUTS/OUTPUTS with control, using_session, app subprocess
  under a socket guard, refused URLs exit 2 before tracing.
- Pins current traceai-openai behaviour: streamed and failed calls have no
  model attribute; the default stream has no usage attributes.
- README: install, configure, run, code, what you see, Baseten specifics,
  privacy, limits, tests and tested versions.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8310
…EADME)

From pr-reviewer t_6453b6cd (APPROVED, P3 follow-ups):
- R1: check_base_url compares the host without an absolute-FQDN trailing
  dot, so https://inference.baseten.co. and model-*.api.baseten.co. are
  refused too (the URL is still returned/refused unchanged, never
  rewritten). Tests pin trailing-dot, upper-case, userinfo and port
  variants; the two trailing-dot cases failed before this change.
- R2: test_hide_outputs gains a non-hidden control.
- R3: README names gen_ai.request.model and notes the requested model stays
  in gen_ai.request.parameters on streamed and failed spans; tests pin it.
- R4: requirements.txt comment no longer mentions a worktree.
- R6: test_main_rejects_url_before_tracing replaces a vacuous fake-server
  check with a client-constructor spy.

approved: Nikhil 2026-10-03 blanket
Refs: TH-8310
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant