vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

90.8k
Stars
+28.0k
Gained
44.7%
Growth
Python
Language

🎭 Best For

⚖️ Compare With

🏷️ Topics & Ecosystem

amd blackwell cuda deepseek deepseek-v3 gpt gpt-oss inference kimi llama llm llm-serving model-serving moe openai pytorch qwen qwen3 tpu transformer

📊 Activity

Latest commit: 2026-09-02. Over the past 294 days, this repository gained 28.0k stars (+44.7% growth). Activity data is based on daily RepoPi snapshots of the GitHub repository.