Fine-tuning small language models (SLMs), engineering agentic systems, and optimizing LLM inference.
Pinned Loading
-
-
tunnel-engine
tunnel-engine PublicA production-grade LLM infra engine that uses vLLM for fast inference, LMCache for smart context caching, and LiteLM to handle multi-model routing and load balancing.
Python 4
-
Tachikoma-Observatory
Tachikoma-Observatory PublicA lightweight evaluation framework for testing Small Language Models (SLMs) on precise tool-calling capabilities to ensure they are production-ready.
Python 2
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.




