AI Software Engineer · AI Agents, Evaluation & Reliability
I build software and platforms, with a focus on reliable AI systems. My background spans distributed systems, cloud platforms, and services operated in production. My open source projects explore how to evaluate AI agents, trace their actions, and understand failures.
Independent research on evaluating AI agents and reconstructing their decisions. Papers and code are linked below.
- Giskard — Authored a merged fix that preserves partial interaction traces when input generation fails.
- OpenTelemetry GenAI — Authored a proposal for external references to AI inference content. Open pull request.
- inspect-evidence-sufficiency — Checks whether AI agent traces contain the evidence needed for a release decision. Includes an optional CI gate.
- decision-trace-reconstructor — Reconstructs AI agent decisions from recorded traces and reports which facts are evidenced, missing, or opaque.
- operational-evidence-plane — Reference implementation connecting release manifests, agent actions, permissions, traces, and replay data.
- AI Agents & Evaluation — agent workflows, evaluation pipelines and release checks.
- Reliability & Observability — tracing, failure analysis, safe deployment and rollback.
- Decision Reconstruction — recovering decision context from traces and runtime records.
- Platform Engineering — distributed services, cloud infrastructure and data pipelines.
- Languages: Python, Go.
- Platforms and data: Kubernetes, Terraform, AWS, GCP, Kafka, Flink, ClickHouse, PostgreSQL.
- Observability and controls: OpenTelemetry, Prometheus, Grafana, OPA/Rego, Vault, OAuth2/OIDC.
- Agent integration: MCP.
- Decision Evidence Maturity Model for Agentic AI: A Property-Level Method Specification
- Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes