- 3.8kFreeTokenllm-inference
Edge-native MoE serving engine designed for running 290B+ frontier models locally on consumer GPUs and CPUs
3.8kApache-2.0 - 2400kKimi K3llm-models
text-generation
2400kApache-2.0 - 1850kGLM-5llm-models
text-generation
1850kMIT - 950kGemma 4llm-models
image-text-to-text
950kgemma - 177kOllamallm-inference
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
177kMIT - 121kllama.cppllm-inference
LLM inference in C/C++
121kMIT - 87kvLLMllm-inference
A high-throughput and memory-efficient inference and serving engine for LLMs
87kApache-2.0 - 48kLocalAIllm-inference
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
48kMIT - 31kSGLangllm-inference
SGLang is a high-performance serving framework for large language models and multimodal models.
31kApache-2.0 - 28kMLXllm-inference
MLX: An array framework for Apple silicon
28kMIT - 25kllamafilellm-inference
Distribute and run LLMs with a single file.
25kNot specified - 23kMLC LLMllm-inference
Universal LLM Deployment Engine with ML Compilation
23kApache-2.0 - 21kCandlellm-inference
Minimalist ML framework for Rust
21kApache-2.0 - 14kTensorRT-LLMllm-inference
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
14kNot specified - 11kkoboldcppllm-inference
Run GGUF models easily with a KoboldAI UI. One File. Zero Install.
11kAGPL-3.0 - 11kText Generation Inferencellm-inference
Large Language Model Text Generation Inference
11kApache-2.0 - 4.6kExLlamaV2llm-inference
A fast inference library for running LLMs locally on modern consumer-class GPUs
4.6kMIT - 1.3kTabbyAPIllm-inference
The official API server for Exllama. OAI compatible, lightweight, and fast.
1.3kAGPL-3.0 - SponsorReach 50,000+ buyers
Enterprise buyers looking for private AI solutions see your brand here.
- 11998kQwen3llm-models
text-generation
11998kApache-2.0 - 9245kDeepSeek-R1llm-models
text-generation
9245kMIT - 1618kLlama 3llm-models
text-generation
1618kllama3 - 1155kDeepSeek-V3llm-models
text-generation
1155kNot specified - 829kGemmallm-models
image-text-to-text
829kgemma - 248kPhi-3 Mini 128K Instructllm-models
text-generation
248kMIT - 162kSmolLMllm-models
text-generation
162kApache-2.0
Stop paying for AI APIs. Everything here runs on your hardware.
Enterprise buyers looking for private AI solutions see your brand here.
No fluff. No spam. Join 12,000+ builders.