Tracked Repositories

175 open-source AI inference repositories across 69 organizations.

175 repositories across 69 organizations

HuggingFace

5 repos·231.1k·124 commits this week

HuggingFace Transformers — state-of-the-art NLP/ML model library (~140K stars)

164.3k
72

HuggingFace Diffusers — diffusion model inference & training (Stable Diffusion, Flux, etc.)

34.3k
34

Minimalist Rust ML framework for inference — targets browser WASM and GPU, zero Python dependency

20.9k
10

HuggingFace TGI — LLM serving (archived March 2026, read-only)

10.9k

One-command HuggingFace Hub → OpenVINO IR export; INT4/INT8 quantization via NNCF for Intel NPU deployment

614
8

TensorFlow

3 repos·206.5k·120 commits this week

Industry-standard deep learning framework with XLA compilation backend

197.1k
112

TensorFlow Serving — high-performance gRPC/REST serving for TF models (multi-version, canary, batching)

6.4k
4

TensorFlow Lite for microcontrollers and embedded devices

3.1k
4

Ollama

1 repo·179.0k·17 commits this week

User-friendly local LLM runner built on llama.cpp (~167K stars)

179.0k
17

ggml-org

2 repos·177.9k·172 commits this week

High-performance LLM inference in C/C++ (CPU + GPU)

124.8k
125

High-performance Whisper speech recognition in C/C++

53.1k
47

Open WebUI

1 repo·149.3k

Self-hosted ChatGPT alternative with built-in RAG, offline-capable (~104K stars)

149.3k

Meta / PyTorch

3 repos·111.8k·343 commits this week

Primary ML framework; torch.compile + AOTInductor for production inference optimization

102.5k
247

PyTorch's portable execution framework for on-device inference

4.9k
96

TorchServe — production PyTorch model serving (archived August 2025)

4.3k

DeepSeek AI

1 repo·104.4k

Reference inference code for DeepSeek-V3 (671B MoE); includes FP8 training framework

104.4k

vLLM Project

3 repos·95.8k·380 commits this week

Most widely adopted open-source LLM serving engine; PagedAttention, continuous batching

89.5k
280

vLLM omni-modality inference (audio, vision, speech)

6.2k
84

vLLM community plugin for Intel Gaudi accelerators

52
16

Google AI Edge

18 repos·83.8k·231 commits this week

Cross-platform ML pipeline framework (vision, audio, NLP)

36.7k
7

AI Edge model gallery

24.5k
5

LiteRT for language model inference

6.2k
40

Official Gemma model cookbook — recipes, fine-tuning, deployment guides

4.0k

Google's Lite Runtime (successor to TensorFlow Lite)

3.3k
95

Sample apps using MediaPipe

2.8k
11

Highly optimized neural network operators library (ARM, x86, WASM)

2.4k
52

Model visualization and exploration tool

1.5k
1

LiteRT integration with PyTorch

1.1k
6

Python API for Coral Edge TPU inference (archived)

405

Sample code for LiteRT

401
3

Quantization tooling for AI Edge models

192
10

C++ API for Coral Edge TPU inference (archived)

95

Web samples for MediaPipe

69

Command-line tooling for LiteRT

39
1

Sample models for AI Edge

25

Evaluation tooling for AI Edge models

10

Google AI Edge documentation site

Nomic AI

1 repo·77.4k

Desktop AI app + SDK for running LLMs locally (~73K stars)

77.4k

Miscellaneous

8 repos·67.0k·45 commits this week

MLC's universal LLM deployment engine (multi-backend)

23.1k
1

DwarfStar native DeepSeek V4 inference engine for Metal, CUDA, and ROCm

21.6k

Smartphone/PC LLM inference exploiting activation sparsity (PowerInfer-2: Qualcomm NPU, Mixtral MoE up to 47B); org transferred from SJTU-IPADS

9.7k

Tile-based ML language and compiler

7.3k
20

Tile-based runtime for ultra-low-latency LLM inference

1.7k

Community on-device LLM project

1.6k

vLLM-style inference on Apple silicon via MLX

1.5k
18

OpenVINO-based OpenAI-compatible inference server for Intel CPU/GPU/NPU

504
6

Apple / ML-Explore

13 repos·64.3k·116 commits this week

Array framework for ML on Apple silicon (Python)

28.1k
60

Example models and applications using MLX

8.9k

Reverse-engineered Apple Neural Engine (ANE) — hardware ops, memory layout, firmware interactions

7.2k

LLM inference and fine-tuning with MLX

6.7k
7

Tools for converting & running models with Core ML

5.4k
6

Example apps using MLX Swift

2.6k

Swift bindings for MLX

2.0k
4

Model export recipes, Python primitives, and Swift runtime utilities for on-device AI (Core AI)

1.6k
14

LLM inference in Swift via MLX

783
15

Efficient data loading for MLX

480

C bindings for MLX

231

Bridges PyTorch and Core AI — convert models to Core AI IR, composite ops, custom lowerings, inline Metal kernels

130
6

PyTorch model compression and optimizations for deployment via Core AI on Apple silicon

105
4

BerriAI

1 repo·56.8k·908 commits this week

Unified OpenAI-compatible proxy for 100+ LLM providers (vLLM, Ollama, Bedrock, Azure, etc.)

56.8k
908

Mudler (LocalAI)

1 repo·48.6k·85 commits this week

Free, open-source OpenAI drop-in replacement — runs locally, no GPU required (~36K stars)

48.6k
85

Microsoft / ONNX

4 repos·47.6k·112 commits this week

Microsoft's cross-platform, high-performance ONNX inference engine

21.4k
71

Open Neural Network Exchange format specification

21.3k
21

Microsoft on-device / local model runtime

2.5k
9

Model optimization tool (HF → ONNX → quantize → NPU deployment); used under the hood by Microsoft Foundry Local

2.4k
11

Oobabooga

1 repo·47.6k·1 commits this week

Gradio web UI for LLMs — multi-backend (llama.cpp, ExLlamaV2, transformers) (~43K stars)

47.6k
1

Exo Explore

1 repo·46.9k

Run LLMs distributed across heterogeneous devices (Mac, iPhone, etc.)

46.9k

Ray Project

1 repo·43.6k·64 commits this week

Distributed AI compute engine; Ray Serve handles online and async batch inference (~39K stars)

43.6k
64

DeepSpeed AI

1 repo·43.0k·12 commits this week

Microsoft DeepSpeed — distributed training and inference (ZeRO, MII, FastGen)

43.0k
12

LM-Sys

1 repo·39.5k

LLM serving framework and home of Chatbot Arena (~37K stars)

39.5k

NVIDIA

4 repos·38.5k·231 commits this week

NVIDIA's optimized LLM inference library (GPU)

14.4k
226

NVIDIA's high-performance deep learning inference SDK (GPU)

13.3k
1

CUDA C++ templates for high-performance matrix-multiply (GEMM) and convolution kernels

10.3k
3

C++ LLM/VLM inference runtime for Jetson and NVIDIA edge devices

512
1

JAX (Google DeepMind)

1 repo·36.2k·145 commits this week

Composable NumPy transformations (JIT, grad, vmap) compiled via XLA to GPUs and TPUs — primary DeepMind research/production runtime

36.2k
145

SGLang

2 repos·33.0k·420 commits this week

High-throughput LLM/VLM serving with RadixAttention and structured generation

32.2k
365

SGLang serving for TTS, ASR, speech, and omni-modal models

875
55

Tencent

2 repos·28.4k·3 commits this week

High-performance neural network inference for mobile (Android/iOS)

23.7k
3

Tencent Neural Network — mobile and edge inference

4.6k

Modular

1 repo·27.6k·389 commits this week

Modular Platform monorepo — MAX inference server/framework + Mojo programming language for portable, high-performance AI on CPUs and GPUs

27.6k
389

Mozilla AI

1 repo·25.7k

Single-file LLM executables via Cosmopolitan Libc — zero install, all platforms (~21K stars)

25.7k

KVCache AI

2 repos·25.6k·69 commits this week

CPU-GPU hybrid inference; runs DeepSeek 671B on 14GB VRAM + 382GB DRAM with massive speedup over llama.cpp

19.3k
6

Kimi serving platform — KV-cache and disaggregated LLM serving

6.3k
63

Dao AI Lab

1 repo·24.7k·1 commits this week

Official FlashAttention — fast, memory-efficient exact attention (FA-2/FA-3) kernels for GPUs

24.7k
1

BentoML

2 repos·21.3k

Run open-source LLMs as OpenAI-compatible API endpoints

12.5k

Unified serving framework: real-time APIs, task queues, batching, multi-model chains

8.8k

jundot

1 repo·20.0k·167 commits this week

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

20.0k
167

Triton Language (OpenAI)

1 repo·20.0k·52 commits this week

Python-like GPU kernel language used by vLLM FlashAttention and PyTorch inductor

20.0k
52

MLC AI

1 repo·18.6k

High-performance LLM inference in web browsers via WebGPU

18.6k

Alibaba

2 repos·17.2k·24 commits this week

Alibaba's neural network inference framework for mobile & edge

15.9k
16

Alibaba high-performance LLM inference engine

1.3k
8

Blaizzy (Community MLX)

5 repos·14.6k·97 commits this week

Audio models (TTS, ASR) with MLX

7.8k
18

Vision-language models on Apple silicon via MLX

5.4k
78

Swift audio inference using MLX

758
1

Text embedding models with MLX

428

Video model inference with MLX

290

K2 / Next-gen ASR

1 repo·14.3k·5 commits this week

ONNX-based runtime for ASR, TTS, VAD, and keyword spotting

14.3k
5

Cactus Compute

6 repos·14.2k·24 commits this week

Tiny on-device foundation model for phones, wearables, and robots

8.0k
20

Cactus core edge inference framework

5.9k
4

React Native bindings for Cactus

178

Kotlin/Android bindings for Cactus

74

Flutter bindings for Cactus

71

Demo chat app using Cactus

28

Apache

2 repos·14.1k·14 commits this week

Apache TVM ML compiler — auto-tunes models for any hardware target

13.7k
13

Apache TVM Foreign Function Interface for deep learning compilation

447
1

ArgMax

5 repos·13.0k·2 commits this week

On-device speech AI for Apple Silicon (WhisperKit sibling)

6.3k
1

On-device Whisper inference for Apple platforms (Swift)

6.3k
1

Python tooling for WhisperKit model optimization

246

On-device AI benchmarking framework

88

Swift playground for ArgMax SDK

21

OpenVINO Toolkit / Intel

3 repos·12.5k·93 commits this week

Intel's toolkit for optimizing & deploying deep learning on Intel hardware

10.7k
74

Neural Network Compression Framework — quantization, pruning, sparsity for OpenVINO

1.2k
1

OpenVINO GenAI — generative AI layer with speculative decoding & KV-cache opt

577
18

RunAnywhere

8 repos·11.9k·103 commits this week

RunAnywhere SDKs for on-device inference deployment

10.3k
58

RunAnywhere CLI tool

1.5k

Swift iOS/macOS starter example for the RunAnywhere SDK

12
10

Browser WASM reference app for the RunAnywhere Web SDK

10
9

React Native starter app using the RunAnywhere SDK

9
2

Android starter example for the RunAnywhere SDK

8
9

Flutter starter app for the RunAnywhere on-device SDK

6
2

Electron desktop app for on-device LLM, VLM, STT, TTS, VAD, and RAG

1
13

AMD Ryzen AI (XDNA NPU)

7 repos·11.8k·107 commits this week

Lemonade SDK — high-level, multi-vendor LLM inference SDK (OGA + llama.cpp), OpenAI-compatible server mode

5.4k
61

Xilinx/AMD AI Engine dev stack — source of the Vitis AI Execution Provider for XDNA

1.8k

AMD-backed Linux NPU runtime, used alongside Lemonade Server on Ryzen AI XDNA 2 (moved to ROCm)

1.8k
26

Open-source RAG/chat/agent app for Ryzen AI NPU, built on Lemonade SDK

1.5k
19

Official Ryzen AI Software stack — Vitis AI EP, quantization, LLM deployment; contains an experimental XDNA-NPU llama.cpp fork

872
1

Ingests any PyTorch model, optimizes, benchmarks/deploys across hardware targets incl. XDNA NPU

244

AMD's quantization toolkit (INT4/INT8 PTQ + QAT) for Ryzen AI NPU deployment

156

Intel

2 repos·11.6k·5 commits this week

Intel IPEX-LLM — local LLM acceleration on Intel hardware (archived Jan 2026, read-only)

8.9k

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity

2.7k
5

Triton Inference Server

1 repo·10.9k·5 commits this week

NVIDIA Triton — production multi-model inference server (HTTP/gRPC, multi-backend)

10.9k
5

Mistral AI

1 repo·10.8k

Official minimal inference library for all Mistral models (7B, Mixtral, Pixtral)

10.8k

Qualcomm

3 repos·9.9k·101 commits this week

On-device LLM/VLM SDK for Snapdragon NPU, GPU, and CPU (formerly Nexa AI / nexa-sdk; acquired by Qualcomm)

8.3k
64

State-of-the-art ML models optimized for Qualcomm Snapdragon NPU/DSP/QNN deployment

1.2k
36

Sample apps and tutorials for deploying models on Qualcomm hardware (TFLite, ONNX, QNN)

442
1

Dusty-NV (NVIDIA Jetson)

1 repo·9.0k

DNN inference library & tutorials for NVIDIA Jetson

9.0k

InternLM / Shanghai AI Lab

1 repo·8.0k·12 commits this week

High-throughput LLM serving with TurboMind engine (C++/CUDA)

8.0k
12

AI Dynamo (NVIDIA)

1 repo·7.8k·115 commits this week

Datacenter-scale distributed inference serving framework (Rust + Python, disaggregated prefill/decode, engine-agnostic)

7.8k
115

Osaurus

1 repo·7.7k·32 commits this week

Native macOS AI agent harness in Swift — any model, persistent memory, autonomous execution, MCP server, MLX + Apple Neural Engine, fully offline

7.7k
32

PaddlePaddle (Baidu)

1 repo·7.3k

Lightweight inference engine for mobile & embedded from PaddlePaddle

7.3k

FlashInfer

1 repo·6.2k·69 commits this week

High-performance GPU kernel library for LLM serving — attention, sampling, and KV-cache primitives

6.2k
69

TurboDeRP (ExLlamaV2)

2 repos·5.8k

High-performance EXL2-quantized inference for consumer NVIDIA GPUs

4.6k

Successor to ExLlamaV2 — quantized local LLM inference on consumer NVIDIA GPUs

1.1k

OpenNMT

1 repo·4.6k·2 commits this week

Fast C++ inference for Transformer models; INT8/INT16 CPU quantization, multi-platform

4.6k
2

OpenXLA

1 repo·4.5k·218 commits this week

Compiler for JAX, TF, PyTorch targeting GPU, TPU, and CPU from a unified IR

4.5k
218

ModelTC

1 repo·4.2k·3 commits this week

Lightweight, high-throughput Python-based LLM inference and serving framework

4.2k
3

Predibase

1 repo·3.8k

Multi-LoRA inference server — serve thousands of fine-tuned adapters on a single GPU

3.8k

Liquid AI

5 repos·3.3k

Examples, tutorials and apps for Liquid AI LFM + LEAP SDK

2.4k

Speech-to-Speech audio models by Liquid AI

562

Minimal fine-tuning repo for LFM2, fully open-source

198

Example apps for LeapSDK

74

Liquid AI documentation

30

Luminal AI

1 repo·2.9k·13 commits this week

Rust-based deep learning compiler with a small static graph IR for fast, portable inference (CUDA, Metal, CPU)

2.9k
13

Fluid Inference

3 repos·2.8k·37 commits this week

On-device audio inference framework

2.7k
29

Fluid Inference core runtime

85
8

Rust text processing library for inference

46

Try Mirai

2 repos·1.8k·29 commits this week

Mirai's on-device inference runtime

1.7k
23

Mirai's LLaMA-based on-device model

88
6

UbiquitousLearning

1 repo·1.6k·1 commits this week

Multimodal LLM inference framework for mobile & edge

1.6k
1

ARM Software

1 repo·1.3k

ARM Neural Network SDK for ARM & Mali devices

1.3k

AMD ROCm

4 repos·1.3k·130 commits this week

AI Tensor Engine for ROCm — centralized repo for high-perf AI operators on AMD Instinct GPUs

533
76

AMD's graph inference engine for MI-series GPUs

325
10

ROCm fork of FlashAttention with Composable Kernel (CK) and Triton backends

238

AiTer Optimized Model — lightweight vLLM-like server built on AITER kernels for ROCm

159
44

Rudrank Riyam

1 repo·1.2k

iOS/macOS workbench for Apple's Foundation Models framework — recipes, guided labs, prompt playground, AFM CLI, FMFBench eval suite, and reusable Swift packages for on-device Apple Intelligence

1.2k

NimbleEdge

2 repos·527

NimbleEdge's deliteAI on-device inference framework

525

NimbleEdge fork of ExecuTorch with edge optimizations

2

ThunderAgent

1 repo·418

A simple, fast and robust program-aware agentic inference system

418

Picovoice

1 repo·317

Picovoice's on-device LLM inference engine

317

Zetic AI

5 repos·78·2 commits this week

MLange sample applications

70

MLange extension library

5

iOS framework for MLange

2
2

iOS extension framework for MLange

1

MLange SDK documentation

0

raullenchai

1 repo·0

The fastest local AI engine for Apple Silicon — 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling, drop-in OpenAI-compatible replacement