Bench360 is a modular benchmarking suite for local LLM deployments. It offers a full-stack, extensible pipeline to evaluate the latency, throughput, quality, and cost of LLM inference on consumer and enterprise GPUs. Bench360 supports flexible backends, tasks and scenarios, enabling fair and reproducible comparisons for researchers & practitioners.
benchmark performance framework energy deployment local optimization engine inference quantization energy-consumption tgi llm vllm llm-inference sglang lmdeploy bench360
-
Updated
Feb 18, 2026 - Python