flashinfer
FlashInfer: Kernel Library for LLM Serving
- gpu
- large-large-models
- cuda
- pytorch
- llm-inference
- jit
- attention
- nvidia
- distributed-inference
- moe
- Stars
- 6,545
- Forks
- 1,531
- + today
- +1
- Created
- 3y
Ranking data as of October 5, 2026 (UTC).
Star History
Today, hour by hour
01
Overview
FlashInfer: Kernel Library for LLM Serving It ranks #1685 on GitTiger, gaining +1 star on October 5, 2026 (UTC).
The project is written in Cuda and has 1,531 forks. It was created 3y ago.
Installation
git clone https://github.com/flashinfer-ai/flashinfer.git
cd flashinfer
# see README for setup