Skip to content
@flashrt-project

FlashRT

Popular repositories Loading

  1. FlashRT FlashRT Public

    FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST.…

    C++ 564 76

  2. FlashRT-HF-kernels FlashRT-HF-kernels Public

    FlashRT-HF-kernels contains standalone FlashRT CUDA/CUTLASS kernels prepared for the Hugging Face kernels community, focusing on small-batch, low-latency inference paths for LLM, VLA, and physical …

    Python 23 7

  3. kernels kernels Public

    Forked from huggingface/kernels

    Build compute kernels and load them from the Hub.

    Python 2

  4. FlashRT-Structures FlashRT-Structures Public

    Attach FlashRT's verified acceleration structures to an unmodified PyTorch host (lerobot, Isaac-GR00T, openpi, transformers, diffusers, vLLM, SGLang) with no fork and no edit to the host. Kernels a…

    Python 1 1

  5. FlashRT-llama.cpp FlashRT-llama.cpp Public

    The FlashRT layer for llama.cpp hosts: fused-structure windows over ggml-cuda driven by FlashRT kernels. Mounted by a host at ggml/src/ggml-cuda/flashrt; vendors the FlashRT kernels it needs, carri…

    Cuda 1

  6. FlashRT-assets FlashRT-assets Public

    FlashRT-assets

Repositories

Showing 6 of 6 repositories
  • FlashRT Public

    FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B

    flashrt-project/FlashRT's past year of commit activity
    C++ 564 Apache-2.0 76 12 9 Updated Sep 12, 2026
  • FlashRT-Structures Public

    Attach FlashRT's verified acceleration structures to an unmodified PyTorch host (lerobot, Isaac-GR00T, openpi, transformers, diffusers, vLLM, SGLang) with no fork and no edit to the host. Kernels arrive from the Hugging Face kernel hub.

    flashrt-project/FlashRT-Structures's past year of commit activity
    Python 1 Apache-2.0 1 0 3 Updated Sep 11, 2026
  • FlashRT-llama.cpp Public

    The FlashRT layer for llama.cpp hosts: fused-structure windows over ggml-cuda driven by FlashRT kernels. Mounted by a host at ggml/src/ggml-cuda/flashrt; vendors the FlashRT kernels it needs, carries CUTLASS, and pins each host's patch series (upstream llama.cpp on RTX 5090, Jetson-PI-Edge on Jetson AGX Thor).

    flashrt-project/FlashRT-llama.cpp's past year of commit activity
    Cuda 1 Apache-2.0 0 0 0 Updated Sep 12, 2026
  • FlashRT-assets Public

    FlashRT-assets

    flashrt-project/FlashRT-assets's past year of commit activity
    0 0 0 0 Updated Aug 31, 2026
  • FlashRT-HF-kernels Public

    FlashRT-HF-kernels contains standalone FlashRT CUDA/CUTLASS kernels prepared for the Hugging Face kernels community, focusing on small-batch, low-latency inference paths for LLM, VLA, and physical AI workloads.

    flashrt-project/FlashRT-HF-kernels's past year of commit activity
    Python 23 MIT 7 0 6 Updated Aug 21, 2026
  • kernels Public Forked from huggingface/kernels

    Build compute kernels and load them from the Hub.

    flashrt-project/kernels's past year of commit activity
    Python 2 Apache-2.0 131 0 0 Updated Aug 14, 2026

Top languages

Loading…

Most used topics

Loading…