This is a uer-friendly Python codebase designed for stress testing of Nvidia GPUs, intel CPUs, and AMD CPUs in various modes
-
Updated
Mar 2, 2025 - Python
This is a uer-friendly Python codebase designed for stress testing of Nvidia GPUs, intel CPUs, and AMD CPUs in various modes
nvProbe — Open-source NVIDIA GPU benchmark suite for CUDA workload automation, Slurm HPC cluster profiling, and MLPerf reporting
PCBench is a versatile Python-based system performance benchmarking tool designed to empower users with insights into their hardware's capabilities. Whether you're a tech enthusiast, a PC gamer, or a developer optimizing your code, PCBench provides comprehensive benchmarking for both CPUs and GPUs.
Benchmark CPU, Benchmark GPU, Storage, RAM using Python
Advanced benchmark harness for AI agent inference workloads on AMD GPU and ROCm cloud infrastructure
Reproducible SRAM surrogate simulation benchmark with CPU/CUDA lanes, fidelity validation, and GPU portability architecture.
A code to benchmark GPU performance on different models
🏆 Which 3DGS renderer is fastest? Which compression is best? We measured them all — on the same GPU, same scenes, same protocol.
Poor Paul's Benchmark as an MCP server. Queryable GPU inference data — quantization, throughput, VRAM, concurrent users — for Claude Desktop, Cursor, Windsurf, Cline, and any MCP client. Self-host or use the free hosted endpoint.
benchHUB is a Python-based project to parse, aggregate, and visualize system and performance benchmarks. It includes a Streamlit dashboard to display and compare results.
PyTorch scripts to benchmark CPU vs GPU matrix multiplication and monitor CPU/GPU stats during a live CIFAR100 training run
GPU vs CPU performance benchmarking for PyTorch and JAX. Works on AMD ROCm, DirectML, CUDA, MPS, CPU. Optimized for RX 5700 XT in WSL2.
Benchmark your GPU against any GGUF model and contribute to the public leaderboard. Measures throughput, TTFT, ITL, and VRAM limits across quantizations and context sizes.
GPU benchmark suite for AI inference workloads. Test throughput, latency, and power efficiency across NVIDIA, AMD, and Apple Silicon. By Petronella Technology Group.
A reproducible GPU benchmarking lab that compares FP16 vs FP32 training on MNIST using PyTorch, CuPy, and Nsight profiling tools. This project blends performance engineering with cinematic storytelling—featuring NVTX-tagged training loops, fused CuPy kernels, and a profiler-driven README that narrates the GPU’s inner workings frame by frame.
🎬 Explore GPU training efficiency with FP32 vs FP16 in this modular lab, utilizing Tensor Core acceleration for deep learning insights.
Local LLM benchmarks on an RTX 5070 Ti 12GB laptop GPU. Speed, coding, reasoning, and tool call accuracy across GGUF models with llama.cpp.
benchmark scaled LLM inference and training
AMD MI300X vs NVIDIA H200: a controlled single-GPU benchmark of inference and training. AMD's 1.32x paper FLOPs advantage inverts on achieved efficiency, but its 1.36x memory becomes a 1.84x KV-cache advantage that wins large-model serving outright.
Benchmarking several NVIDIA GPUs with dorado
Add a description, image, and links to the gpu-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the gpu-benchmark topic, visit your repo's landing page and select "manage topics."