The 12 best GitHub repos for LLM serving
Run, serve, and route models yourself. Local inference, gateways, and the proxies that let you swap providers without rewriting your app.
Browse all top repos for LLM serving
Ollama
★ 180k · Go · ActiveGet up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
llama.cpp
★ 125k · C++ · ActiveA few options to get llama.cpp installed on your machine: Visit https://llama.
vLLM
★ 89k · Python · ActiveA high-throughput and memory-efficient inference and serving engine for LLMs
Unsloth
★ 74k · Python · ActiveLocal UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
GPT4Free
★ 67k · Python · ActiveThe official gpt4free repository. various collection of powerful language models. opus 4.6 gpt 5.3 kimi 2.5 deepseek v3.2 gemini 3
LiteLLM
★ 57k · Python · ActiveThe fastest, litest AI Gateway. Rust core with Python SDK.
AirLLM
★ 32k · Jupyter Notebook · ActiveAirLLM 70B inference with single 4GB GPU
Gitleaks
★ 29k · Go · ActiveFind secrets with Gitleaks
Dyad
★ 21k · TypeScript · ActiveLocal, open-source AI app builder for power users v0 / Lovable / Replit / Bolt alternative Star if you like it!
Skypilot
★ 11k · Python · ActiveThe AI Compute Platform for frontier teams.
ClawRouter
★ 6.6k · TypeScript · ActiveThe agent-native LLM router for autonomous agents. 66 models (8 free), <1ms local routing, USDC payments on Base & Solana via x402.
LM Studio CLI
★ 5.2k · TypeScript · Activelms - Command Line Tool for LM Studio Built with lmstudio.
FAQs about repos for LLM serving
Have questions about GitHub repositories and which ones are worth using? Find answers to the most common ones below.