I am focused on Agentic RL and LLM agent engineering: tool-use trajectories, reward and feedback design, evaluation, memory, and observability. I care about the part where research ideas become systems that can be tested, debugged, and improved.
|
Agentic RL Tool-use trajectories, reward signals, exploration, credit assignment, and agent evaluation. |
LLM Agents Planning, memory, retrieval, tool calling, multi-step reasoning, and human feedback. |
Systems Tracing, reports, dashboards, and interfaces that make agent behavior easier to inspect. |
| Layer | What I am exploring |
|---|---|
| Policy and planning | Agent loops, tool selection, reflective planning, task decomposition |
| Reward and feedback | Trajectory scoring, rubrics, preference signals, report-quality evaluation |
| Memory and knowledge | RAG, evidence graphs, long-running context, domain knowledge curation |
| Observability | Agent traces, async job dashboards, failure analysis, experiment reports |
| Product surface | Practical web apps and bots that make agent behavior visible and useful |
| Project | Direction | Notes |
|---|---|---|
| lingshu-nexus | LLM knowledge system | Research evidence platform for acupuncture and tVNS/taVNS scenarios. |
| decision-twin.skill | Agentic decision support | Codex skill for building a decision twin from context, constraints, values, and history. |
| feishu-project-bot | LLM workflow automation | Feishu bot that parses project updates, tracks progress, and generates reports. |
| GitCommit2Report | Developer productivity | Turns Git commit history into structured weekly reports with LLM assistance. |
| transformer-explore | Model understanding | Lightweight visual exploration around transformer concepts. |
| arq_dashboard | Agent infrastructure | Redis-backed dashboard for ARQ background jobs and async task visibility. |
| meme-maker | AI web product | Lightweight AI meme generator with multi-image upload, AI copywriting, and export flow. |
- Building LLM agents with clearer feedback, tracing, and evaluation loops.
- Studying how reinforcement learning ideas can improve tool-use agents.
- Turning research evidence and project context into structured systems that agents can use.
The best way to reach me is through GitHub: open an issue in a relevant repository or start from Deep-Octopus.


