Applied Research Engineer at Abundant (YC)
Research on agent evals — benchmarks, harnesses, and environments for frontier coding agents.
I design agent benchmarks and the sandboxed environments that grade them deterministically — clean reinforcement-learning signal and rigorous evaluation for frontier coding agents. My current focus is long-horizon autonomy: multi-hour library reproductions, full-stack product clones, and ML builds.
Co-author of SWE-Marathon, a benchmark of 20 multi-hour software-engineering tasks (code · paper).
- [08/2026] ⭐ Featured on the GLM 5.3 model card!
- [08/2026] ⚙️ Featured on the Grok 4.6 model card!
- [07/2026] ➕ SWE-Marathon v1.1 is out with a fresh leaderboard!
- [07/2026] 🌙 Featured on the Kimi K3 model card!
- [07/2026] 🔥 Featured on the Grok 4.5 model card!
- [06/2026] 🚀 Featured on the GLM 5.2 model card!
agent evaluation RL environments long-horizon tasks execution sandboxes model reliability



