
flashinfer
Pilot it+198 this week★ 6,268 · 1,342 forks
LLM serving kernel library
Worth a timeboxed spike before you bet on it.
Who it's for
small teams with PyTorch and GPU needs
What it replaces
custom LLM serving code
The catch
high ops burden and CUDA dependency
Your first hour
Evaluate FlashInfer's PyTorch integration
The numbers
Stars6,268
Forks1,342
Stars added (7d)+198
Open issues852
LanguagePython
LicenceApache-2.0
Last pushUpdated today
Project age3 years old
Maintainers describe it as: “FlashInfer: Kernel Library for LLM Serving”
attentioncudadistributed-inferencegpujitlarge-large-modelsllm-inferencemoenvidiapytorch
flashinfer, in short
- Should a small team use flashinfer?
- Pilot it. Worth a timeboxed spike before you bet on it. small teams with PyTorch and GPU needs
- What does flashinfer actually do?
- LLM serving kernel library
- What does flashinfer replace?
- custom LLM serving code
- What is the downside of flashinfer?
- high ops burden and CUDA dependency
- Can flashinfer be used in a commercial product?
- Its licence is Apache-2.0, which is permissive and generally fine for commercial use. Confirm against the LICENSE file in the repository.