diff --git a/CHANGELOG.md b/CHANGELOG.md index 0f74d727..18160799 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] - 2026-09-28 +### Added +- Agent Forge keeps the traces of its runs (M5b, #227): every run is written with `fileTraceExporter()` to `.lousho/agents//traces` (the files `lousho traces` reads), and the Trace tab gets a list of the agent's past runs (time, duration, model calls, tokens, cost, status) with a "Live" entry while a run is active; choosing one opens its spans in the waterfall, with span kind and error status. New server routes `GET /agents/:id/traces?limit=N` and `GET /agents/:id/traces/:traceId`; the server reads only inside the agent's trace folder. The live `span` WebSocket messages now carry `kind` and `status` (optional fields of `SpanEvent`). See [Agent Forge](docs/agent-forge.md). + ### Changed - CI peer matrix (LOU-M8, #231): the `typecheck-ai7` and `typecheck-zod4` jobs are replaced by one `peers` job with five entries, each installed for real on top of the default install and run through `tsc`, `test:types`, both builds and `npx vitest run`: `ai4-zod4`, `ai6-zod3`, `ai6-zod4`, `ai7-zod3`, `ai7-zod4`. The two zod 4 entries on `ai` 6/7 also install `ollama-ai-provider-v2` (3.x / 4.x), and the new `src/providers/ollamaV2.contract.test.ts` runs `OllamaProvider` against the real package and a local fake Ollama server (`generate()`, `stream()`, a tool-call turn through `createAgent().send()` and `.stream()`). The `ai-v6` dev alias replaces the hand-made `ai` 6 stand-in in `aiMajorPeers.test.ts`. - `lousho init --provider ollama` now scaffolds `ai@^7.0.0` with `ollama-ai-provider-v2@^4.0.0` and `zod@^4.0.0` (it was `ai@^4.3.19` with `ollama-ai-provider@^1.2.0` and zod 3). Existing projects are not touched; to stay on the old pairing keep `ai@^4.3.19` and `ollama-ai-provider@^1.2.0`. diff --git a/apps/agent-forge/server/__tests__/runRegistry.test.ts b/apps/agent-forge/server/__tests__/runRegistry.test.ts index 252a9c06..56faaa50 100644 --- a/apps/agent-forge/server/__tests__/runRegistry.test.ts +++ b/apps/agent-forge/server/__tests__/runRegistry.test.ts @@ -148,6 +148,10 @@ describe('RunManager', () => { expect(runSpans).toHaveLength(2); // start + end const [startSpan, endSpan] = runSpans; expect(startSpan.endTime).toBeUndefined(); + // M5b: span kind is forwarded, and status when the SDK set one (only a failed span has it). + expect(startSpan.kind).toBe('internal'); + expect(spans.every((s) => s.status === undefined)).toBe(true); + expect(spans.find((s) => s.attributes['gen_ai.operation.name'] === 'chat')?.kind).toBe('client'); expect(endSpan.endTime).toBeGreaterThanOrEqual(endSpan.startTime); expect(spans.some((s) => s.attributes['gen_ai.operation.name'] === 'chat')).toBe(true); expect( diff --git a/apps/agent-forge/server/__tests__/traces.test.ts b/apps/agent-forge/server/__tests__/traces.test.ts new file mode 100644 index 00000000..c3602f15 --- /dev/null +++ b/apps/agent-forge/server/__tests__/traces.test.ts @@ -0,0 +1,149 @@ +/** + * M5b: persisted trace history - a run writes the SDK's trace files under the + * agent's folder and `GET /agents/:id/traces[/:traceId]` reads them back. + */ +import { describe, it, expect, beforeEach, afterEach } from 'vitest'; +import * as fs from 'node:fs'; +import * as path from 'node:path'; +import * as os from 'node:os'; +import request from 'supertest'; +import { listTraces } from '@lousho/build-ai-agent/traces'; +import type { AgentSpec } from '@lousho/build-ai-agent'; +import { createApp } from '../app'; +import { RunManager } from '../runRegistry'; +import { FileCheckpointStore } from '../checkpointStore'; +import { FileApprovalStore } from '../approvalStore'; +import { createFsAgentStore } from '../../src/persistence/fsAgentStore'; +import type { AgentRunStatusPayload, TraceDetailPayload, TraceSummaryPayload } from '../../shared/wireTypes'; + +const SPEC: AgentSpec = { + name: 'trace-agent', + prompt: 'You are a helpful agent.', + provider: { type: 'mock', model: 'mock-1' }, +}; + +function stopped(runManager: RunManager, agentId: string): Promise { + if (runManager.status(agentId).status === 'stopped') return Promise.resolve(runManager.status(agentId)); + return new Promise((resolve, reject) => { + const timer = setTimeout(() => reject(new Error(`Timed out waiting for ${agentId}`)), 3000); + const onStatus = (payload: AgentRunStatusPayload) => { + if (payload.agentId !== agentId || payload.status !== 'stopped') return; + clearTimeout(timer); + runManager.off('status', onStatus); + resolve(payload); + }; + runManager.on('status', onStatus); + }); +} + +describe('M5b trace history routes', () => { + let baseDir: string; + let runManager: RunManager; + let app: ReturnType; + + beforeEach(() => { + baseDir = fs.mkdtempSync(path.join(os.tmpdir(), 'lou-m5b-test-')); + const agentStore = createFsAgentStore(baseDir); + runManager = new RunManager({ + baseDir, + checkpointStore: new FileCheckpointStore(baseDir), + approvalStore: new FileApprovalStore(baseDir), + loadSpec: (id) => agentStore.load(id), + saveSpec: (id, spec) => agentStore.save(id, spec), + }); + app = createApp({ agentStore, runManager, baseDir }); + }); + + afterEach(() => { + fs.rmSync(baseDir, { recursive: true, force: true }); + }); + + async function finishedRun(agentId: string, input = 'What is the weather?'): Promise { + await request(app).post(`/agents/${agentId}/run`).send({ input, spec: SPEC }).expect(202); + await stopped(runManager, agentId); + } + + it('a run writes a trace file under the agent folder, readable with the SDK reader', async () => { + await finishedRun('weather'); + const dir = path.join(baseDir, '.lousho', 'agents', 'weather', 'traces'); + expect(fs.existsSync(dir)).toBe(true); + const summaries = await listTraces({ dir }); + expect(summaries).toHaveLength(1); + expect(summaries[0].file.startsWith(dir)).toBe(true); + }); + + it('GET /agents/:id/traces lists the runs newest first, without file paths', async () => { + await finishedRun('weather'); + await finishedRun('weather', 'And tomorrow?'); + const res = await request(app).get('/agents/weather/traces').expect(200); + const traces = (res.body as { traces: TraceSummaryPayload[] }).traces; + expect(traces).toHaveLength(2); + expect(traces[0].startTime).toBeGreaterThanOrEqual(traces[1].startTime); + expect(traces[0].status).toBe('ok'); + expect(traces[0].modelCalls).toBeGreaterThan(0); + expect(traces[0]).not.toHaveProperty('file'); + expect(JSON.stringify(res.body)).not.toContain('lou-m5b-test-'); + + const limited = await request(app).get('/agents/weather/traces?limit=1').expect(200); + expect(limited.body.traces).toHaveLength(1); + }); + + it('GET /agents/:id/traces/:traceId returns the spans with kind and status', async () => { + await finishedRun('weather'); + const [summary] = (await request(app).get('/agents/weather/traces').expect(200)).body.traces as TraceSummaryPayload[]; + const res = await request(app).get(`/agents/weather/traces/${summary.traceId}`).expect(200); + const body = res.body as TraceDetailPayload; + expect(body.traceId).toBe(summary.traceId); + const root = body.spans.find((s) => s.id === summary.traceId); + expect(root).toMatchObject({ kind: 'internal' }); + expect(body.spans.find((s) => s.attributes['gen_ai.operation.name'] === 'chat')).toMatchObject({ kind: 'client' }); + }); + + it('a failed span keeps its error status in the list and the detail', async () => { + const dir = path.join(baseDir, '.lousho', 'agents', 'broken', 'traces', '2026-10-02'); + fs.mkdirSync(dir, { recursive: true }); + const root = { v: 1, traceId: 'r1', id: 'r1', name: 'invoke_agent broken', kind: 'internal', startTime: 1000, endTime: 1500, attributes: {}, status: { code: 'error', message: 'boom' } }; + const tool = { v: 1, traceId: 'r1', id: 't1', parentId: 'r1', name: 'execute_tool x', kind: 'internal', startTime: 1100, endTime: 1200, attributes: {}, status: { code: 'error', message: 'boom' } }; + const lines = [tool, root].map((line) => `${JSON.stringify(line)} +`).join(''); + fs.writeFileSync(path.join(dir, 'r1.jsonl'), lines); + expect((await request(app).get('/agents/broken/traces').expect(200)).body.traces[0]).toMatchObject({ traceId: 'r1', status: 'error' }); + const detail = (await request(app).get('/agents/broken/traces/r1').expect(200)).body as TraceDetailPayload; + expect(detail.spans.map((s) => s.status)).toEqual([{ code: 'error', message: 'boom' }, { code: 'error', message: 'boom' }]); + }); + + it('an agent without runs has an empty list; an unknown agent is 404', async () => { + await request(app).put('/agents/fresh').send(SPEC).expect(204); + expect((await request(app).get('/agents/fresh/traces').expect(200)).body).toEqual({ traces: [] }); + await request(app).get('/agents/nobody/traces').expect(404); + await request(app).get('/agents/nobody/traces/abc').expect(404); + }); + + it('an unknown trace is 404', async () => { + await finishedRun('weather'); + await request(app).get('/agents/weather/traces/00000000-0000-0000-0000-000000000000').expect(404); + }); + + it('a path-like trace or agent id is 400 and never reaches the filesystem', async () => { + await finishedRun('weather'); + for (const bad of ['..%2F..%2Fagent.yaml', '..', 'a%2Fb', 'a%5Cb', '%2E%2E', 'x.jsonl']) { + const res = await request(app).get(`/agents/weather/traces/${bad}`); + expect(res.status, bad).not.toBe(200); + expect(res.status, bad).toBeLessThan(500); + } + await request(app).get('/agents/weather/traces/..%2F..%2Fagent.yaml').expect(400); + await request(app).get('/agents/weather/traces/a%5Cb').expect(400); + await request(app).get('/agents/..%2Fweather/traces').expect(400); + }); + + it('an ambiguous id prefix is 400', async () => { + await finishedRun('weather'); + const dir = path.join(baseDir, '.lousho', 'agents', 'weather', 'traces'); + const [day] = fs.readdirSync(dir); + for (const name of ['ab-1', 'ab-2']) { + const line = { v: 1, traceId: name, id: name, name: 'invoke_agent x', startTime: 1, endTime: 2, attributes: {} }; + fs.writeFileSync(path.join(dir, day, `${name}.jsonl`), `${JSON.stringify(line)}\n`); + } + await request(app).get('/agents/weather/traces/ab').expect(400); + }); +}); diff --git a/apps/agent-forge/server/app.ts b/apps/agent-forge/server/app.ts index 06461daa..db97b6d1 100644 --- a/apps/agent-forge/server/app.ts +++ b/apps/agent-forge/server/app.ts @@ -10,6 +10,8 @@ * POST /agents/:id/stop -> abort the in-flight run, if any * GET /agents/:id/status -> AgentRunStatusPayload * POST /agents/:id/approve -> body: { approvalId, approved, note? } + * GET /agents/:id/traces?limit=N -> { traces: TraceSummaryPayload[] } (M5b, newest first) + * GET /agents/:id/traces/:traceId -> TraceDetailPayload | 404 (M5b) * GET /runs/:id/history -> RunHistoryPayload (LOU-D45 time travel) * POST /runs/:id/fork -> body: ForkRunRequest; starts the fork, 202 ForkRunResponse * GET /runs/compare?a=&b= -> RunComparisonPayload (compareTrajectories) @@ -45,6 +47,14 @@ import { isValidAgentId } from './types'; import { SecretsStore, isSecretProvider } from './secretsStore'; import { SettingsStore } from './settingsStore'; import type { ForkRunRequest, ForkRunResponse, SettingsProfile } from '../shared/wireTypes'; +import { + DEFAULT_TRACE_LIMIT, + MAX_TRACE_LIMIT, + isValidTraceId, + listAgentTraces, + readAgentTrace, + traceDirExists, +} from './traceStore'; import { DEPLOY_ADAPTERS, isDeployAdapter, runDeploy } from './deployRunner'; export interface CreateAppOptions { @@ -216,6 +226,54 @@ function parseForkRequest(body: Partial | undefined): { fromStep return { fromStep: fromStep as number, patch }; } +/** M5b: the persisted traces of an agent's runs (files under `.lousho/agents//traces`, the ones `lousho traces` reads). */ +function registerTraceRoutes(app: Express, agentStore: AgentStore, baseDir: string): void { + const known = async (id: string) => traceDirExists(baseDir, id) || (await agentStore.load(id)) !== undefined; + + app.get( + '/agents/:id/traces', + asyncRoute(async (req, res) => { + const raw = Number.parseInt(String(req.query.limit ?? ''), 10); + const limit = Number.isInteger(raw) && raw > 0 ? Math.min(raw, MAX_TRACE_LIMIT) : DEFAULT_TRACE_LIMIT; + if (!(await known(paramId(req)))) { + res.status(404).json({ error: `No saved agent '${paramId(req)}'` }); + return; + } + res.json({ traces: await listAgentTraces(baseDir, paramId(req), limit) }); + }) + ); + + app.get( + '/agents/:id/traces/:traceId', + asyncRoute(async (req, res) => { + const traceId = Array.isArray(req.params.traceId) ? req.params.traceId[0] : req.params.traceId; + if (!isValidTraceId(traceId)) { + res.status(400).json({ error: 'Invalid trace id' }); + return; + } + if (!(await known(paramId(req)))) { + res.status(404).json({ error: `No saved agent '${paramId(req)}'` }); + return; + } + let spans; + try { + spans = await readAgentTrace(baseDir, paramId(req), traceId); + } catch (error) { + if (error instanceof SDKError) { + res.status(400).json({ error: error.message }); + return; + } + throw error; + } + if (spans.length === 0) { + res.status(404).json({ error: `No trace '${traceId}' for agent '${paramId(req)}'` }); + return; + } + res.json({ traceId, spans }); + }) + ); +} + /** LOU-D45 time travel: a run's step history, forking it from a step, and comparing two runs. */ function registerTimeTravelRoutes(app: Express, runManager: RunManager): void { app.get( @@ -553,6 +611,7 @@ export function createApp({ registerChatRoutes(app, runManager, triggerRegistry); registerDebugRoutes(app, runManager); registerTimeTravelRoutes(app, runManager); + registerTraceRoutes(app, agentStore, baseDir); registerProviderKeyRoutes(app, secrets); registerProfileRoutes(app, settings); registerDeployRoutes(app, agentStore, baseDir); diff --git a/apps/agent-forge/server/runRegistry.ts b/apps/agent-forge/server/runRegistry.ts index a757b85a..bed4c7a1 100644 --- a/apps/agent-forge/server/runRegistry.ts +++ b/apps/agent-forge/server/runRegistry.ts @@ -10,6 +10,8 @@ */ import { EventEmitter } from 'node:events'; import { randomUUID } from 'node:crypto'; +import { fileTraceExporter } from '@lousho/build-ai-agent/traces'; +import { agentTraceDir, fanOutExporter } from './traceStore'; import { AgentExecutor, FlowExecutor, @@ -313,6 +315,8 @@ export class RunManager extends EventEmitter { startTime: span.startTime, endTime: span.endTime, attributes: span.attributes, + ...(span.kind !== undefined && { kind: span.kind }), + ...(span.status !== undefined && { status: span.status }), }); } @@ -320,13 +324,15 @@ export class RunManager extends EventEmitter { * O2: builds a `TraceExporter` (src/execution/tracing.ts) that forwards * every real span start/end notification for this run over the existing * WS channel (as `{type:'span', ...}` messages, see wsServer.ts) rather - * than a second tracing pipeline. + * than a second tracing pipeline. M5b: it also writes the run to the + * agent's trace folder (see traceStore.ts), so the Trace tab can reopen it. */ private makeTraceExporter(agentId: string): TraceExporter { - return { + const live: TraceExporter = { onSpanStart: (span) => this.emitSpan(agentId, span), onSpanEnd: (span) => this.emitSpan(agentId, span), }; + return fanOutExporter(live, fileTraceExporter({ dir: agentTraceDir(this.opts.baseDir, agentId) })); } /** O3: (re)creates the debug session for a fresh run(), seeded from this agent's persisted breakpoints. */ diff --git a/apps/agent-forge/server/traceStore.ts b/apps/agent-forge/server/traceStore.ts new file mode 100644 index 00000000..6e450574 --- /dev/null +++ b/apps/agent-forge/server/traceStore.ts @@ -0,0 +1,59 @@ +/** + * M5b: where Agent Forge keeps run traces, and how it reads them back. + * + * Every run writes the SDK's trace files (`fileTraceExporter`, the same + * format `lousho traces` reads) to `/.lousho/agents//traces`. + * The server only ever reads inside that folder: an agent id is validated + * before it becomes a path segment, a trace id is matched against the file + * names found there (never joined into a path), and the absolute file path is + * not sent to the browser. + */ +import * as fs from 'node:fs'; +import * as path from 'node:path'; +import type { TraceExporter } from '@lousho/build-ai-agent'; +import { listTraces, readTrace, type TraceSummary } from '@lousho/build-ai-agent/traces'; +import type { SpanEvent, TraceSummaryPayload } from '../shared/wireTypes'; + +/** A trace id is a UUID in practice; this is the strictest shape that still allows prefixes. */ +const TRACE_ID_RE = /^[A-Za-z0-9-]{1,128}$/; + +export const DEFAULT_TRACE_LIMIT = 20; +export const MAX_TRACE_LIMIT = 200; + +export function isValidTraceId(id: string): boolean { + return TRACE_ID_RE.test(id); +} + +/** The agent's trace folder; `agentId` must already have passed `isValidAgentId`. */ +export function agentTraceDir(baseDir: string, agentId: string): string { + return path.join(baseDir, '.lousho', 'agents', agentId, 'traces'); +} + +export function traceDirExists(baseDir: string, agentId: string): boolean { + return fs.existsSync(agentTraceDir(baseDir, agentId)); +} + +/** A TraceExporter that calls every exporter in turn. */ +export function fanOutExporter(...exporters: TraceExporter[]): TraceExporter { + return { + onSpanStart: (span) => exporters.forEach((exporter) => exporter.onSpanStart?.(span)), + onSpanEnd: (span) => exporters.forEach((exporter) => exporter.onSpanEnd?.(span)), + }; +} + +function toPayload(summary: TraceSummary): TraceSummaryPayload { + const { file, ...rest } = summary; + void file; // the server-side path stays on the server + return rest; +} + +/** The agent's recent traces, newest first. */ +export async function listAgentTraces(baseDir: string, agentId: string, limit: number): Promise { + const summaries = await listTraces({ dir: agentTraceDir(baseDir, agentId), limit }); + return summaries.map(toPayload); +} + +/** One trace's spans; `[]` when the agent has no such trace. */ +export async function readAgentTrace(baseDir: string, agentId: string, traceId: string): Promise { + return (await readTrace(traceId, { dir: agentTraceDir(baseDir, agentId) })) as SpanEvent[]; +} diff --git a/apps/agent-forge/shared/wireTypes.ts b/apps/agent-forge/shared/wireTypes.ts index eeec025d..5c05f011 100644 --- a/apps/agent-forge/shared/wireTypes.ts +++ b/apps/agent-forge/shared/wireTypes.ts @@ -98,6 +98,31 @@ export interface SpanEvent { startTime: number; endTime?: number; attributes: Record; + /** M5b: `internal` or `client` (a model call). */ + kind?: 'internal' | 'client'; + /** M5b: `error` when the span failed. */ + status?: { code: 'ok' | 'error'; message?: string }; +} + +/** M5b: one persisted run in `GET /agents/:id/traces` (the SDK's `TraceSummary` without the server-side file path). */ +export interface TraceSummaryPayload { + traceId: string; + name: string; + agent?: string; + startTime: number; + durationMs: number; + status: 'ok' | 'error'; + modelCalls: number; + toolCalls: number; + inputTokens: number; + outputTokens: number; + costUsd?: number; +} + +/** M5b: `GET /agents/:id/traces/:traceId`. */ +export interface TraceDetailPayload { + traceId: string; + spans: SpanEvent[]; } /** diff --git a/apps/agent-forge/src/components/debug/TraceHistory.tsx b/apps/agent-forge/src/components/debug/TraceHistory.tsx new file mode 100644 index 00000000..9db06323 --- /dev/null +++ b/apps/agent-forge/src/components/debug/TraceHistory.tsx @@ -0,0 +1,68 @@ +import type { TraceSummaryPayload } from '../../../shared/wireTypes'; + +/** The list entry that stands for the run in progress. */ +export const LIVE_ENTRY = 'live'; + +function formatDuration(ms: number): string { + return ms < 1000 ? `${Math.round(ms)}ms` : `${(ms / 1000).toFixed(2)}s`; +} + +function formatCost(costUsd: number | undefined): string { + return costUsd === undefined ? '-' : `$${costUsd.toFixed(costUsd < 0.01 ? 6 : 4)}`; +} + +/** One line per persisted run: when, how long, model calls, tokens in/out, cost and status. */ +export function TraceHistory({ + traces, + selectedId, + liveActive, + error, + onSelect, +}: { + traces: TraceSummaryPayload[]; + /** `LIVE_ENTRY` or a trace id. */ + selectedId: string; + /** A run is in progress: the list starts with a "Live" entry. */ + liveActive: boolean; + error?: string; + onSelect: (id: string) => void; +}) { + return ( +
+ {liveActive && ( + + )} + {error &&
{error}
} + {!error && traces.length === 0 && !liveActive && ( +
No saved traces yet. Each run is saved under .lousho/agents/<agent>/traces.
+ )} + {traces.map((trace) => ( + + ))} +
+ ); +} diff --git a/apps/agent-forge/src/components/debug/TracePanel.tsx b/apps/agent-forge/src/components/debug/TracePanel.tsx index f058220a..9af75d37 100644 --- a/apps/agent-forge/src/components/debug/TracePanel.tsx +++ b/apps/agent-forge/src/components/debug/TracePanel.tsx @@ -1,9 +1,12 @@ -import { useMemo, useState } from 'react'; +import { useEffect, useMemo, useState } from 'react'; import { useAppState } from '../../state/AppState'; import { orderSpansForWaterfall } from '../../state/spanReducer'; -import type { SpanEvent } from '../../../shared/wireTypes'; +import type { SpanEvent, TraceSummaryPayload } from '../../../shared/wireTypes'; import type { AgentGraphSpec } from '../../graph/types'; +import { runtimeClient } from '../../runtime/runtimeClient'; +import { errorMessage } from '../errorMessage'; import { findLlmNodeId, findToolNodeId } from './nodeLookup'; +import { LIVE_ENTRY, TraceHistory } from './TraceHistory'; function nodeIdForSpan(graph: AgentGraphSpec, span: SpanEvent): string | undefined { // SDK spans follow the OpenTelemetry GenAI conventions (LOU-D9): the operation is @@ -30,6 +33,24 @@ function spanGeometry(span: SpanEvent, minStart: number, totalMs: number) { return { startPct, durationMs, widthPct }; } +/** The tooltip of a span row: its attributes, led by the failure message of an errored span. */ +function spanTitle(span: SpanEvent): string { + const attributes = JSON.stringify(span.attributes); + const message = span.status?.code === 'error' ? span.status.message : undefined; + return message ? `${message} +${attributes}` : attributes; +} + +/** The duration column: `error` for a failed span, `running` until it ends. */ +function spanDuration(span: SpanEvent, durationMs: number): string { + if (span.status?.code === 'error') return 'error'; + return span.endTime === undefined ? 'running' : `${durationMs}ms`; +} + +function classNames(...names: (string | false)[]): string { + return names.filter(Boolean).join(' '); +} + function SpanRow({ span, minStart, @@ -44,29 +65,25 @@ function SpanRow({ onSelect: () => void; }) { const { startPct, durationMs, widthPct } = spanGeometry(span, minStart, totalMs); - const running = span.endTime === undefined; + const failed = span.status?.code === 'error'; return ( -
+
{span.name} + {span.kind && {span.kind}} - + - {running ? 'running' : `${durationMs}ms`} + {spanDuration(span, durationMs)}
); } -/** - * O2: real span waterfall, driven by `{type:'span'}` WS messages forwarded - * from the SDK's `TraceExporter.onSpanStart`/`onSpanEnd` hooks (see - * server/runRegistry.ts's `makeTraceExporter()`) - bar position/width are - * real `Date.now()` timestamps, not synthetic ones. Clicking a span - * highlights the corresponding canvas node (`llm`/`tool.call`'s attribute - * carries the tool name) and jumps the Logs tab's highlight to the same - * node, via the shared `AppState.highlightedNodeId`. - */ -export function TracePanel() { - const { spans, highlightedNodeId, setHighlightedNodeId, graph, setDrawerTab } = useAppState(); +/** The waterfall of one trace's spans; clicking a span highlights its canvas node. */ +function Waterfall({ spans, jumpToLogs }: { spans: SpanEvent[]; jumpToLogs: boolean }) { + const { highlightedNodeId, setHighlightedNodeId, graph, setDrawerTab } = useAppState(); const [selectedSpanId, setSelectedSpanId] = useState(undefined); const ordered = useMemo(() => orderSpansForWaterfall(spans), [spans]); @@ -76,7 +93,7 @@ export function TracePanel() { function handleSelect(span: SpanEvent) { setSelectedSpanId(span.id); setHighlightedNodeId(nodeIdForSpan(graph, span)); - setDrawerTab('logs'); + if (jumpToLogs) setDrawerTab('logs'); } if (ordered.length === 0) { @@ -101,3 +118,96 @@ export function TracePanel() {
); } + +/** + * O2: real span waterfall, driven by `{type:'span'}` WS messages forwarded + * from the SDK's `TraceExporter.onSpanStart`/`onSpanEnd` hooks (see + * server/runRegistry.ts's `makeTraceExporter()`) - bar position/width are + * real `Date.now()` timestamps, not synthetic ones. Clicking a span + * highlights the corresponding canvas node (`llm`/`tool.call`'s attribute + * carries the tool name) and jumps the Logs tab's highlight to the same + * node, via the shared `AppState.highlightedNodeId`. + * + * M5b: a list on the left holds the agent's persisted traces (the files + * `lousho traces` reads, served by `GET /agents/:id/traces`); choosing one + * shows its spans in the same waterfall. While a run is active the list + * starts with "Live", which is the WebSocket feed above. The list reloads + * when the agent changes and each time a run ends. + */ +export function TracePanel() { + const { spans, agentId, runStatus } = useAppState(); + const [traces, setTraces] = useState([]); + const [listError, setListError] = useState(undefined); + const [selectedId, setSelectedId] = useState(LIVE_ENTRY); + const [pastSpans, setPastSpans] = useState([]); + const status = runStatus?.status; + const liveActive = status === 'running' || status === 'paused'; + + // A run starts: follow it. + useEffect(() => { + if (liveActive) setSelectedId(LIVE_ENTRY); + }, [liveActive]); + + // Another agent: back to its live view. + useEffect(() => { + setSelectedId(LIVE_ENTRY); + setPastSpans([]); + }, [agentId]); + + // Load the list for this agent, and again when a run ends. + useEffect(() => { + if (liveActive) return; + let cancelled = false; + runtimeClient + .listTraces(agentId) + .then((list) => { + if (cancelled) return; + setTraces(list); + setListError(undefined); + }) + .catch((error: unknown) => { + if (!cancelled) setListError(errorMessage(error)); + }); + return () => { + cancelled = true; + }; + }, [agentId, liveActive, status]); + + // A past trace is loaded when it is chosen. + useEffect(() => { + if (selectedId === LIVE_ENTRY) return; + let cancelled = false; + runtimeClient + .readTrace(agentId, selectedId) + .then((loaded) => { + if (!cancelled) setPastSpans(loaded); + }) + .catch((error: unknown) => { + if (cancelled) return; + setPastSpans([]); + setListError(errorMessage(error)); + }); + return () => { + cancelled = true; + }; + }, [agentId, selectedId]); + + const showingLive = selectedId === LIVE_ENTRY; + return ( +
+ { + setListError(undefined); + setSelectedId(id); + }} + /> +
+ +
+
+ ); +} diff --git a/apps/agent-forge/src/components/debug/__tests__/TraceHistory.test.tsx b/apps/agent-forge/src/components/debug/__tests__/TraceHistory.test.tsx new file mode 100644 index 00000000..4ed14492 --- /dev/null +++ b/apps/agent-forge/src/components/debug/__tests__/TraceHistory.test.tsx @@ -0,0 +1,182 @@ +import { describe, it, expect, vi, afterEach, beforeEach } from 'vitest'; +import { act } from 'react'; +import { createRoot, type Root } from 'react-dom/client'; +import type { ReactElement } from 'react'; +import { TraceHistory, LIVE_ENTRY } from '../TraceHistory'; +import { TracePanel } from '../TracePanel'; +import type { AgentRunStatusPayload, SpanEvent, TraceSummaryPayload } from '../../../../shared/wireTypes'; + +(globalThis as { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true; + +const mocks = vi.hoisted(() => ({ + listTraces: vi.fn(), + readTrace: vi.fn(), + appState: {} as Record, +})); + +vi.mock('../../../runtime/runtimeClient', () => ({ + runtimeClient: { listTraces: mocks.listTraces, readTrace: mocks.readTrace }, +})); +vi.mock('../../../state/AppState', () => ({ useAppState: () => mocks.appState })); + +// What `GET /agents/:id/traces` returns for an agent that ran twice, the newer run failing. +const TRACES: TraceSummaryPayload[] = [ + { + traceId: 'trace-new', + name: 'invoke_agent weather', + agent: 'weather', + startTime: Date.parse('2026-10-02T10:05:00Z'), + durationMs: 1650, + status: 'error', + modelCalls: 2, + toolCalls: 1, + inputTokens: 159, + outputTokens: 32, + costUsd: 0.000043, + }, + { + traceId: 'trace-old', + name: 'invoke_agent weather', + startTime: Date.parse('2026-10-02T10:00:00Z'), + durationMs: 420, + status: 'ok', + modelCalls: 1, + toolCalls: 0, + inputTokens: 10, + outputTokens: 5, + }, +]; + +// What `GET /agents/:id/traces/trace-new` returns: a root, a model call and a failed tool call. +const PAST_SPANS: SpanEvent[] = [ + { id: 'trace-new', name: 'invoke_agent weather', startTime: 1000, endTime: 2650, kind: 'internal', status: { code: 'error', message: 'boom' }, attributes: { 'gen_ai.operation.name': 'invoke_agent' } }, + { id: 'c1', parentId: 'trace-new', name: 'chat gpt-4o-mini', startTime: 1000, endTime: 1900, kind: 'client', attributes: { 'gen_ai.operation.name': 'chat' } }, + { id: 't1', parentId: 'trace-new', name: 'execute_tool get_weather', startTime: 1900, endTime: 1920, kind: 'internal', status: { code: 'error', message: 'boom' }, attributes: { 'gen_ai.operation.name': 'execute_tool' } }, +]; + +const LIVE_SPANS: SpanEvent[] = [ + { id: 'live-root', name: 'invoke_agent live-run', startTime: 5000, attributes: {} }, +]; + +const statusOf = (status: AgentRunStatusPayload['status']): AgentRunStatusPayload => ({ agentId: 'weather', status, updatedAt: '2026-10-02T10:00:00.000Z' }); + +let root: Root | undefined; +afterEach(() => { + act(() => root?.unmount()); + root = undefined; + vi.clearAllMocks(); +}); + +function render(element: ReactElement): HTMLElement { + const container = document.createElement('div'); + root = createRoot(container); + act(() => root?.render(element)); + return container; +} + +async function flush(): Promise { + await act(async () => { + await Promise.resolve(); + }); +} + +describe('TraceHistory (M5b)', () => { + it('renders a row per trace with time, duration, model calls, tokens, cost and status', () => { + const container = render( undefined} />); + const rows = [...container.querySelectorAll('[role=option]')]; + expect(rows).toHaveLength(2); + expect(rows[0].textContent).toContain('1.65s'); + expect(rows[0].textContent).toContain('2 model'); + expect(rows[0].textContent).toContain('159/32 tokens'); + expect(rows[0].textContent).toContain('$0.000043'); + expect(rows[0].textContent).toContain('error'); + expect(rows[1].textContent).toContain('420ms'); + expect(rows[1].textContent).toContain('-'); + expect(rows[1].getAttribute('aria-selected')).toBe('true'); + }); + + it('shows the Live entry first only while a run is active', () => { + const active = render( undefined} />); + expect(active.querySelector('[role=option]')?.textContent).toContain('Live'); + act(() => root?.unmount()); + const idle = render( undefined} />); + expect(idle.textContent).not.toContain('Live'); + }); + + it('calls onSelect with the trace id, and explains an empty list', () => { + const onSelect = vi.fn(); + const container = render(); + act(() => (container.querySelectorAll('[role=option]')[1] as HTMLElement).click()); + expect(onSelect).toHaveBeenCalledWith('trace-old'); + act(() => root?.unmount()); + const empty = render(); + expect(empty.textContent).toContain('No saved traces yet'); + }); +}); + +describe('TracePanel with trace history (M5b)', () => { + beforeEach(() => { + mocks.listTraces.mockResolvedValue(TRACES); + mocks.readTrace.mockResolvedValue(PAST_SPANS); + mocks.appState.agentId = 'weather'; + mocks.appState.spans = []; + mocks.appState.runStatus = statusOf('stopped'); + mocks.appState.highlightedNodeId = undefined; + mocks.appState.setHighlightedNodeId = vi.fn(); + mocks.appState.setDrawerTab = vi.fn(); + mocks.appState.graph = { nodes: [], edges: [] }; + }); + + it('lists the agent traces and shows the selected one in the waterfall with kind and error status', async () => { + const container = render(); + await flush(); + expect(mocks.listTraces).toHaveBeenCalledWith('weather'); + expect(container.querySelectorAll('[role=option]')).toHaveLength(2); + + act(() => (container.querySelectorAll('[role=option]')[0] as HTMLElement).click()); + await flush(); + expect(mocks.readTrace).toHaveBeenCalledWith('weather', 'trace-new'); + const rows = [...container.querySelectorAll('.trace-row')]; + expect(rows.map((row) => row.querySelector('.trace-name')?.textContent)).toEqual([ + 'invoke_agent weather', + 'chat gpt-4o-mini', + 'execute_tool get_weather', + ]); + expect(rows[1].querySelector('.trace-kind')?.textContent).toBe('client'); + expect(rows[2].classList.contains('trace-row-error')).toBe(true); + expect(rows[2].querySelector('.trace-bar')?.classList.contains('error')).toBe(true); + expect(rows[1].classList.contains('trace-row-error')).toBe(false); + }); + + it('shows the live spans under a Live entry while a run is active', async () => { + mocks.appState.runStatus = statusOf('running'); + mocks.appState.spans = LIVE_SPANS; + const container = render(); + await flush(); + const options = [...container.querySelectorAll('[role=option]')]; + expect(options[0].textContent).toContain('Live'); + expect(options[0].getAttribute('aria-selected')).toBe('true'); + expect(container.querySelector('.trace-name')?.textContent).toBe('invoke_agent live-run'); + expect(mocks.listTraces).not.toHaveBeenCalled(); + }); + + it('reloads the list when a run ends', async () => { + mocks.appState.runStatus = statusOf('running'); + const container = render(); + await flush(); + expect(mocks.listTraces).not.toHaveBeenCalled(); + + mocks.appState.runStatus = statusOf('stopped'); + act(() => root?.render()); + await flush(); + expect(mocks.listTraces).toHaveBeenCalledTimes(1); + expect(container.querySelectorAll('[role=option]')).toHaveLength(2); + }); + + it('says why the list is empty when the server cannot be read', async () => { + mocks.listTraces.mockRejectedValue(new Error('No saved agent')); + const container = render(); + await flush(); + expect(container.textContent).toContain('No saved agent'); + }); +}); diff --git a/apps/agent-forge/src/components/layout.css b/apps/agent-forge/src/components/layout.css index 07356cc3..f780e889 100644 --- a/apps/agent-forge/src/components/layout.css +++ b/apps/agent-forge/src/components/layout.css @@ -793,6 +793,8 @@ /* ---------- O2: Trace tab ---------- */ .trace-panel { + flex: 1; + min-width: 0; display: flex; flex-direction: column; height: 100%; @@ -1072,3 +1074,23 @@ .chat-input:disabled { opacity: 0.6; } + +/* M5b: persisted trace history next to the waterfall */ +.trace-layout { display: flex; height: 100%; min-height: 0; } +.trace-main { flex: 1; min-width: 0; height: 100%; } +.trace-history { display: flex; flex-direction: column; gap: 4px; width: 240px; flex: none; padding: 8px; overflow-y: auto; border-right: 1px solid var(--border); } +.trace-item { display: flex; flex-direction: column; gap: 2px; text-align: left; padding: 6px 8px; border: 1px solid var(--border); border-radius: 6px; background: transparent; color: var(--text); font: inherit; font-size: 12px; cursor: pointer; } +.trace-item:hover, .trace-item.selected { background: var(--surface-2); } +.trace-item.selected { border-color: var(--accent); } +.trace-item-head { display: flex; justify-content: space-between; gap: 8px; } +.trace-item-meta { color: var(--text-muted); } +.trace-item-error .trace-item-status, .trace-history-error { color: var(--danger); } +.trace-history-note { padding: 4px; color: var(--text-faint); font-size: 12px; } +.trace-kind { flex: none; color: var(--text-faint); font-size: 11px; } +.trace-bar.error { background: var(--danger); } +.trace-row-error .trace-name { color: var(--danger); } +@media (max-width: 700px) { + .trace-layout { flex-direction: column; } + .trace-history { width: auto; flex-direction: row; overflow-x: auto; overflow-y: hidden; border-right: 0; border-bottom: 1px solid var(--border); } + .trace-item { min-width: 200px; } +} diff --git a/apps/agent-forge/src/runtime/runtimeClient.ts b/apps/agent-forge/src/runtime/runtimeClient.ts index 95f47ded..f8026e57 100644 --- a/apps/agent-forge/src/runtime/runtimeClient.ts +++ b/apps/agent-forge/src/runtime/runtimeClient.ts @@ -22,6 +22,9 @@ import type { SettingsFile, SettingsProfile, StreamMessage, + SpanEvent, + TraceDetailPayload, + TraceSummaryPayload, } from '../../shared/wireTypes'; /** Same-origin default: `lousho studio` prints the API server's own URL, but in dev the Vite server proxies to it (see vite.config.ts). */ @@ -165,6 +168,25 @@ class RuntimeClient { return this.request(`/runs/${encodeURIComponent(runId)}/history`, { method: 'GET' }); } + /** M5b: the agent's persisted traces (the files `lousho traces` reads), newest first. */ + async listTraces(agentId: string, limit?: number): Promise { + const query = limit === undefined ? '' : `?${new URLSearchParams({ limit: String(limit) })}`; + const body = await this.request<{ traces: TraceSummaryPayload[] }>( + `/agents/${encodeURIComponent(agentId)}/traces${query}`, + { method: 'GET' } + ); + return body.traces; + } + + /** M5b: the spans of one persisted trace. */ + async readTrace(agentId: string, traceId: string): Promise { + const body = await this.request( + `/agents/${encodeURIComponent(agentId)}/traces/${encodeURIComponent(traceId)}`, + { method: 'GET' } + ); + return body.spans; + } + /** LOU-D45: forks run `runId` at a step, patched, and starts the fork - its status streams on `subscribe(response.runId)`. */ async forkRun(runId: string, body: ForkRunRequest): Promise { return this.request(`/runs/${encodeURIComponent(runId)}/fork`, { method: 'POST', body: JSON.stringify(body) }); diff --git a/docs/agent-forge.md b/docs/agent-forge.md index 4a48d025..8ed6dbdf 100644 --- a/docs/agent-forge.md +++ b/docs/agent-forge.md @@ -98,6 +98,26 @@ exists. message array" there to inspect the in-flight message list. **Output** shows the full `ExecutionResult` (messages, tool calls, usage, steps) as a collapsible JSON tree once the run finishes or pauses. + + **Trace history.** Every run is also saved as a trace file, so the + **Trace** tab keeps the runs of the agent after the page is reloaded or the + server restarted. A list on the left shows the agent's past runs, newest + first (start time, duration, model calls, tokens in/out, cost, status), + with a **Live** entry on top while a run is active; choose one to open its + spans in the same waterfall, with each span's kind (`internal` or + `client`) and failed spans in the error colour. The list refreshes when + a run ends. The files are the ones `lousho traces` reads, in + `.lousho/agents//traces//.jsonl` under the + directory `lousho studio` was started in (see + [Local traces](observability.md#local-traces)): + `npx lousho traces --dir .lousho/agents//traces` shows the same + runs in the terminal, and deleting the agent's folder deletes its traces. + The tab shows what the spans hold: model and tool names, token counts, + cost, durations and error messages, and with the span's tooltip the + attributes, which carry prompt and tool content. The server reads traces + only from that folder (an id with a path separator is a 400) and never + sends the file path to the browser. Keep `.lousho/` out of version + control. 5. **Chat with it.** The **Chat** tab is a separate conversational transport (`POST /agents/:id/message`, streamed back over the same WebSocket as everything else) - send it a message and it replies in the @@ -196,12 +216,15 @@ and replay it next to the original, with A run id is the run's checkpoint session id: the agent id for the agent's own runs, `.fork-` for a fork of run ``. The runtime control -server has three routes for it (the History tab uses them): +server has three routes for it (the History tab uses them), and two more for +an agent's saved traces (the Trace tab uses them): | Route | What it does | | --- | --- | | `GET /runs/:id/history` | `{ runId, steps }`: one entry per step of the run's latest execution (the newest checkpoint saved at that step), oldest first, with `status`, `savedAt`, `finishReason`, `toolCalls` (`{ id, name, args, result? }`), and `tokens` / `costUsd` of the step's model call when known. 404 when the run has no checkpoints. | | `POST /runs/:id/fork` | Body `{ fromStep, patch? }`, where `patch` is `{ toolResult?: { toolCallId, result }, appendInput?, businessState? }`. Forks the run at `fromStep`, then starts the fork through the run registry with the original run's spec. Returns 202 `{ runId, fromStep, status }`; the fork streams on `WS /agents//stream` and reports on `GET /agents//status` like any run. 400 for a bad body or an unknown `toolCallId`, 404 for an unknown run or a step with no checkpoint. | +| `GET /agents/:id/traces?limit=N` | `{ traces }`: the agent's saved runs from `.lousho/agents//traces`, newest first (`limit` defaults to 20, at most 200), each with `traceId`, `name`, `startTime`, `durationMs`, `status`, `modelCalls`, `toolCalls`, `inputTokens`, `outputTokens` and `costUsd` when known. An agent with no runs gives an empty list; 404 for an unknown agent. | +| `GET /agents/:id/traces/:traceId` | `{ traceId, spans }`: the spans of one saved run by start time, each with `id`, `parentId`, `name`, `kind`, `status`, `startTime`, `endTime` and `attributes`. `traceId` may be a unique prefix. 400 for an id that is not letters, digits and dashes or a prefix that matches several runs, 404 for an unknown agent or trace. | | `GET /runs/compare?a=&b=` | `compareTrajectories(a, b)` of two runs' latest checkpoints: each run's model turns, `divergedAt` (the first turn that differs) and the `drift` entries. 404 when either run has no checkpoint. | The flow, for an agent `weather` that has run once: diff --git a/llms-full.txt b/llms-full.txt index 364aef75..212101d3 100644 --- a/llms-full.txt +++ b/llms-full.txt @@ -1670,6 +1670,26 @@ exists. message array" there to inspect the in-flight message list. **Output** shows the full `ExecutionResult` (messages, tool calls, usage, steps) as a collapsible JSON tree once the run finishes or pauses. + + **Trace history.** Every run is also saved as a trace file, so the + **Trace** tab keeps the runs of the agent after the page is reloaded or the + server restarted. A list on the left shows the agent's past runs, newest + first (start time, duration, model calls, tokens in/out, cost, status), + with a **Live** entry on top while a run is active; choose one to open its + spans in the same waterfall, with each span's kind (`internal` or + `client`) and failed spans in the error colour. The list refreshes when + a run ends. The files are the ones `lousho traces` reads, in + `.lousho/agents//traces//.jsonl` under the + directory `lousho studio` was started in (see + [Local traces](https://github.com/LinuxDevil/agent-sdk/blob/main/docs/observability.md#local-traces)): + `npx lousho traces --dir .lousho/agents//traces` shows the same + runs in the terminal, and deleting the agent's folder deletes its traces. + The tab shows what the spans hold: model and tool names, token counts, + cost, durations and error messages, and with the span's tooltip the + attributes, which carry prompt and tool content. The server reads traces + only from that folder (an id with a path separator is a 400) and never + sends the file path to the browser. Keep `.lousho/` out of version + control. 5. **Chat with it.** The **Chat** tab is a separate conversational transport (`POST /agents/:id/message`, streamed back over the same WebSocket as everything else) - send it a message and it replies in the @@ -1768,12 +1788,15 @@ and replay it next to the original, with A run id is the run's checkpoint session id: the agent id for the agent's own runs, `.fork-` for a fork of run ``. The runtime control -server has three routes for it (the History tab uses them): +server has three routes for it (the History tab uses them), and two more for +an agent's saved traces (the Trace tab uses them): | Route | What it does | | --- | --- | | `GET /runs/:id/history` | `{ runId, steps }`: one entry per step of the run's latest execution (the newest checkpoint saved at that step), oldest first, with `status`, `savedAt`, `finishReason`, `toolCalls` (`{ id, name, args, result? }`), and `tokens` / `costUsd` of the step's model call when known. 404 when the run has no checkpoints. | | `POST /runs/:id/fork` | Body `{ fromStep, patch? }`, where `patch` is `{ toolResult?: { toolCallId, result }, appendInput?, businessState? }`. Forks the run at `fromStep`, then starts the fork through the run registry with the original run's spec. Returns 202 `{ runId, fromStep, status }`; the fork streams on `WS /agents//stream` and reports on `GET /agents//status` like any run. 400 for a bad body or an unknown `toolCallId`, 404 for an unknown run or a step with no checkpoint. | +| `GET /agents/:id/traces?limit=N` | `{ traces }`: the agent's saved runs from `.lousho/agents//traces`, newest first (`limit` defaults to 20, at most 200), each with `traceId`, `name`, `startTime`, `durationMs`, `status`, `modelCalls`, `toolCalls`, `inputTokens`, `outputTokens` and `costUsd` when known. An agent with no runs gives an empty list; 404 for an unknown agent. | +| `GET /agents/:id/traces/:traceId` | `{ traceId, spans }`: the spans of one saved run by start time, each with `id`, `parentId`, `name`, `kind`, `status`, `startTime`, `endTime` and `attributes`. `traceId` may be a unique prefix. 400 for an id that is not letters, digits and dashes or a prefix that matches several runs, 404 for an unknown agent or trace. | | `GET /runs/compare?a=&b=` | `compareTrajectories(a, b)` of two runs' latest checkpoints: each run's model turns, `divergedAt` (the first turn that differs) and the `drift` entries. 404 when either run has no checkpoint. | The flow, for an agent `weather` that has run once: