Skip to main content
Back to trending

flashinfer

FlashInfer: Kernel Library for LLM Serving

  • gpu
  • large-large-models
  • cuda
  • pytorch
  • llm-inference
  • jit
  • attention
  • nvidia
  • distributed-inference
  • moe
View on GitHub
Stars
6,545
Forks
1,531
+ today
+1
Created
3y

Ranking data as of October 5, 2026 (UTC).

Star History

Today, hour by hour

6,545 stars
01

Overview

FlashInfer: Kernel Library for LLM Serving It ranks #1685 on GitTiger, gaining +1 star on October 5, 2026 (UTC).

The project is written in Cuda and has 1,531 forks. It was created 3y ago.

Installation
git clone https://github.com/flashinfer-ai/flashinfer.git
cd flashinfer
# see README for setup