Tracked Repositories

175 open-source AI inference repositories across 69 organizations.

175 repositories across 69 organizations

HuggingFace

5 repos·231.7k·140 commits this week

HuggingFace Transformers — state-of-the-art NLP/ML model library (~140K stars)

164.7k
69

HuggingFace Diffusers — diffusion model inference & training (Stable Diffusion, Flux, etc.)

34.4k
43

Minimalist Rust ML framework for inference — targets browser WASM and GPU, zero Python dependency

21.0k
24

HuggingFace TGI — LLM serving (archived March 2026, read-only)

10.9k

One-command HuggingFace Hub → OpenVINO IR export; INT4/INT8 quantization via NNCF for Intel NPU deployment

617
4

TensorFlow

3 repos·208.2k·388 commits this week

Industry-standard deep learning framework with XLA compilation backend

198.8k
381

TensorFlow Serving — high-performance gRPC/REST serving for TF models (multi-version, canary, batching)

6.4k
3

TensorFlow Lite for microcontrollers and embedded devices

3.1k
4

ggml-org

2 repos·180.3k·135 commits this week

High-performance LLM inference in C/C++ (CPU + GPU)

126.9k
133

High-performance Whisper speech recognition in C/C++

53.4k
2

Ollama

1 repo·180.0k·20 commits this week

User-friendly local LLM runner built on llama.cpp (~167K stars)

180.0k
20

Open WebUI

1 repo·150.8k·108 commits this week

Self-hosted ChatGPT alternative with built-in RAG, offline-capable (~104K stars)

150.8k
108

Meta / PyTorch

3 repos·112.1k·524 commits this week

Primary ML framework; torch.compile + AOTInductor for production inference optimization

102.7k
371

PyTorch's portable execution framework for on-device inference

5.0k
153

TorchServe — production PyTorch model serving (archived August 2025)

4.3k

DeepSeek AI

1 repo·104.4k

Reference inference code for DeepSeek-V3 (671B MoE); includes FP8 training framework

104.4k

vLLM Project

3 repos·97.5k·462 commits this week

Most widely adopted open-source LLM serving engine; PagedAttention, continuous batching

90.9k
343

vLLM omni-modality inference (audio, vision, speech)

6.6k
110

vLLM community plugin for Intel Gaudi accelerators

53
9

Google AI Edge

18 repos·84.3k·300 commits this week

Cross-platform ML pipeline framework (vision, audio, NLP)

36.8k
21

AI Edge model gallery

24.6k
30

LiteRT for language model inference

6.4k
77

Official Gemma model cookbook — recipes, fine-tuning, deployment guides

4.0k

Google's Lite Runtime (successor to TensorFlow Lite)

3.4k
98

Sample apps using MediaPipe

2.8k
2

Highly optimized neural network operators library (ARM, x86, WASM)

2.4k
39

Model visualization and exploration tool

1.6k
7

LiteRT integration with PyTorch

1.1k
11

Sample code for LiteRT

416
14

Python API for Coral Edge TPU inference (archived)

405

Quantization tooling for AI Edge models

196
1

C++ API for Coral Edge TPU inference (archived)

95

Web samples for MediaPipe

74

Command-line tooling for LiteRT

41

Sample models for AI Edge

25

Evaluation tooling for AI Edge models

10

Google AI Edge documentation site

Nomic AI

1 repo·77.4k

Desktop AI app + SDK for running LLMs locally (~73K stars)

77.4k

Miscellaneous

8 repos·67.7k·162 commits this week

MLC's universal LLM deployment engine (multi-backend)

23.1k

DwarfStar native DeepSeek V4 inference engine for Metal, CUDA, and ROCm

22.0k
121

Smartphone/PC LLM inference exploiting activation sparsity (PowerInfer-2: Qualcomm NPU, Mixtral MoE up to 47B); org transferred from SJTU-IPADS

9.8k

Tile-based ML language and compiler

7.3k
39

Tile-based runtime for ultra-low-latency LLM inference

1.7k

Community on-device LLM project

1.7k

vLLM-style inference on Apple silicon via MLX

1.6k
2

OpenVINO-based OpenAI-compatible inference server for Intel CPU/GPU/NPU

513

Apple / ML-Explore

13 repos·65.3k·86 commits this week

Array framework for ML on Apple silicon (Python)

28.3k
28

Example models and applications using MLX

8.9k

Reverse-engineered Apple Neural Engine (ANE) — hardware ops, memory layout, firmware interactions

7.3k

LLM inference and fine-tuning with MLX

6.9k
17

Tools for converting & running models with Core ML

5.4k
1

Example apps using MLX Swift

2.7k

Model export recipes, Python primitives, and Swift runtime utilities for on-device AI (Core AI)

2.0k
16

Swift bindings for MLX

2.0k
3

LLM inference in Swift via MLX

796
15

Efficient data loading for MLX

483

C bindings for MLX

235
3

Bridges PyTorch and Core AI — convert models to Core AI IR, composite ops, custom lowerings, inline Metal kernels

143
1

PyTorch model compression and optimizations for deployment via Core AI on Apple silicon

114
2

BerriAI

1 repo·57.9k·1258 commits this week

Unified OpenAI-compatible proxy for 100+ LLM providers (vLLM, Ollama, Bedrock, Azure, etc.)

57.9k
1258

Mudler (LocalAI)

1 repo·48.8k·57 commits this week

Free, open-source OpenAI drop-in replacement — runs locally, no GPU required (~36K stars)

48.8k
57

Microsoft / ONNX

4 repos·48.1k·109 commits this week

Microsoft's cross-platform, high-performance ONNX inference engine

21.8k
49

Open Neural Network Exchange format specification

21.4k
50

Microsoft on-device / local model runtime

2.5k
6

Model optimization tool (HF → ONNX → quantize → NPU deployment); used under the hood by Microsoft Foundry Local

2.4k
4

Oobabooga

1 repo·47.6k

Gradio web UI for LLMs — multi-backend (llama.cpp, ExLlamaV2, transformers) (~43K stars)

47.6k

Exo Explore

1 repo·47.2k

Run LLMs distributed across heterogeneous devices (Mac, iPhone, etc.)

47.2k

Ray Project

1 repo·43.7k·75 commits this week

Distributed AI compute engine; Ray Serve handles online and async batch inference (~39K stars)

43.7k
75

DeepSpeed AI

1 repo·43.1k·33 commits this week

Microsoft DeepSpeed — distributed training and inference (ZeRO, MII, FastGen)

43.1k
33

LM-Sys

1 repo·39.5k

LLM serving framework and home of Chatbot Arena (~37K stars)

39.5k

NVIDIA

4 repos·38.8k·225 commits this week

NVIDIA's optimized LLM inference library (GPU)

14.5k
223

NVIDIA's high-performance deep learning inference SDK (GPU)

13.3k

CUDA C++ templates for high-performance matrix-multiply (GEMM) and convolution kernels

10.4k
2

C++ LLM/VLM inference runtime for Jetson and NVIDIA edge devices

534

JAX (Google DeepMind)

1 repo·36.2k·136 commits this week

Composable NumPy transformations (JIT, grad, vmap) compiled via XLA to GPUs and TPUs — primary DeepMind research/production runtime

36.2k
136

SGLang

2 repos·34.9k·454 commits this week

High-throughput LLM/VLM serving with RadixAttention and structured generation

33.8k
398

SGLang serving for TTS, ASR, speech, and omni-modal models

1.1k
56

Modular

1 repo·29.5k·315 commits this week

Modular Platform monorepo — MAX inference server/framework + Mojo programming language for portable, high-performance AI on CPUs and GPUs

29.5k
315

Tencent

2 repos·28.4k·7 commits this week

High-performance neural network inference for mobile (Android/iOS)

23.8k
7

Tencent Neural Network — mobile and edge inference

4.6k

KVCache AI

2 repos·25.9k·79 commits this week

CPU-GPU hybrid inference; runs DeepSeek 671B on 14GB VRAM + 382GB DRAM with massive speedup over llama.cpp

19.5k
7

Kimi serving platform — KV-cache and disaggregated LLM serving

6.5k
72

Mozilla AI

1 repo·25.9k

Single-file LLM executables via Cosmopolitan Libc — zero install, all platforms (~21K stars)

25.9k

Dao AI Lab

1 repo·24.8k·1 commits this week

Official FlashAttention — fast, memory-efficient exact attention (FA-2/FA-3) kernels for GPUs

24.8k
1

jundot

1 repo·21.4k·71 commits this week

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

21.4k
71

BentoML

2 repos·21.3k·2 commits this week

Run open-source LLMs as OpenAI-compatible API endpoints

12.5k

Unified serving framework: real-time APIs, task queues, batching, multi-model chains

8.8k
2

Triton Language (OpenAI)

1 repo·20.1k·45 commits this week

Python-like GPU kernel language used by vLLM FlashAttention and PyTorch inductor

20.1k
45

MLC AI

1 repo·18.9k·4 commits this week

High-performance LLM inference in web browsers via WebGPU

18.9k
4

Alibaba

2 repos·17.3k·10 commits this week

Alibaba's neural network inference framework for mobile & edge

16.0k
5

Alibaba high-performance LLM inference engine

1.3k
5

Cactus Compute

6 repos·16.5k·11 commits this week

Tiny on-device foundation model for phones, wearables, and robots

10.2k
11

Cactus core edge inference framework

6.0k

React Native bindings for Cactus

178

Kotlin/Android bindings for Cactus

74

Flutter bindings for Cactus

71

Demo chat app using Cactus

28

Blaizzy (Community MLX)

5 repos·14.8k·123 commits this week

Audio models (TTS, ASR) with MLX

7.8k
39

Vision-language models on Apple silicon via MLX

5.5k
80

Swift audio inference using MLX

770
4

Text embedding models with MLX

439

Video model inference with MLX

295

K2 / Next-gen ASR

1 repo·14.6k·11 commits this week

ONNX-based runtime for ASR, TTS, VAD, and keyword spotting

14.6k
11

Apache

2 repos·14.2k·63 commits this week

Apache TVM ML compiler — auto-tunes models for any hardware target

13.7k
49

Apache TVM Foreign Function Interface for deep learning compilation

456
14

ArgMax

5 repos·13.1k

On-device speech AI for Apple Silicon (WhisperKit sibling)

6.4k

On-device Whisper inference for Apple platforms (Swift)

6.4k

Python tooling for WhisperKit model optimization

246

On-device AI benchmarking framework

87

Swift playground for ArgMax SDK

21

OpenVINO Toolkit / Intel

3 repos·12.6k·106 commits this week

Intel's toolkit for optimizing & deploying deep learning on Intel hardware

10.8k
83

Neural Network Compression Framework — quantization, pruning, sparsity for OpenVINO

1.2k
3

OpenVINO GenAI — generative AI layer with speculative decoding & KV-cache opt

582
20

AMD Ryzen AI (XDNA NPU)

7 repos·12.1k·142 commits this week

Lemonade SDK — high-level, multi-vendor LLM inference SDK (OGA + llama.cpp), OpenAI-compatible server mode

5.6k
25

AMD-backed Linux NPU runtime, used alongside Lemonade Server on Ryzen AI XDNA 2 (moved to ROCm)

1.8k
51

Xilinx/AMD AI Engine dev stack — source of the Vitis AI Execution Provider for XDNA

1.8k

Open-source RAG/chat/agent app for Ryzen AI NPU, built on Lemonade SDK

1.5k
66

Official Ryzen AI Software stack — Vitis AI EP, quantization, LLM deployment; contains an experimental XDNA-NPU llama.cpp fork

879

Ingests any PyTorch model, optimizes, benchmarks/deploys across hardware targets incl. XDNA NPU

246

AMD's quantization toolkit (INT4/INT8 PTQ + QAT) for Ryzen AI NPU deployment

162

RunAnywhere

8 repos·11.9k·35 commits this week

RunAnywhere SDKs for on-device inference deployment

10.3k
16

RunAnywhere CLI tool

1.5k
9

Swift iOS/macOS starter example for the RunAnywhere SDK

12
4

Browser WASM reference app for the RunAnywhere Web SDK

10

React Native starter app using the RunAnywhere SDK

9

Android starter example for the RunAnywhere SDK

9

Flutter starter app for the RunAnywhere on-device SDK

6

Electron desktop app for on-device LLM, VLM, STT, TTS, VAD, and RAG

2
6

Intel

2 repos·11.6k·2 commits this week

Intel IPEX-LLM — local LLM acceleration on Intel hardware (archived Jan 2026, read-only)

8.9k

SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity

2.7k
2

Triton Inference Server

1 repo·11.0k·8 commits this week

NVIDIA Triton — production multi-model inference server (HTTP/gRPC, multi-backend)

11.0k
8

Mistral AI

1 repo·10.8k

Official minimal inference library for all Mistral models (7B, Mixtral, Pixtral)

10.8k

Qualcomm

3 repos·10.0k·110 commits this week

On-device LLM/VLM SDK for Snapdragon NPU, GPU, and CPU (formerly Nexa AI / nexa-sdk; acquired by Qualcomm)

8.3k
63

State-of-the-art ML models optimized for Qualcomm Snapdragon NPU/DSP/QNN deployment

1.2k
46

Sample apps and tutorials for deploying models on Qualcomm hardware (TFLite, ONNX, QNN)

449
1

Dusty-NV (NVIDIA Jetson)

1 repo·9.0k

DNN inference library & tutorials for NVIDIA Jetson

9.0k

InternLM / Shanghai AI Lab

1 repo·8.0k·13 commits this week

High-throughput LLM serving with TurboMind engine (C++/CUDA)

8.0k
13

AI Dynamo (NVIDIA)

1 repo·8.0k·150 commits this week

Datacenter-scale distributed inference serving framework (Rust + Python, disaggregated prefill/decode, engine-agnostic)

8.0k
150

Osaurus

1 repo·7.8k·81 commits this week

Native macOS AI agent harness in Swift — any model, persistent memory, autonomous execution, MCP server, MLX + Apple Neural Engine, fully offline

7.8k
81

PaddlePaddle (Baidu)

1 repo·7.3k

Lightweight inference engine for mobile & embedded from PaddlePaddle

7.3k

FlashInfer

1 repo·6.3k·68 commits this week

High-performance GPU kernel library for LLM serving — attention, sampling, and KV-cache primitives

6.3k
68

TurboDeRP (ExLlamaV2)

2 repos·5.9k·44 commits this week

High-performance EXL2-quantized inference for consumer NVIDIA GPUs

4.6k

Successor to ExLlamaV2 — quantized local LLM inference on consumer NVIDIA GPUs

1.3k
44

OpenNMT

1 repo·4.7k·4 commits this week

Fast C++ inference for Transformer models; INT8/INT16 CPU quantization, multi-platform

4.7k
4

OpenXLA

1 repo·4.5k·202 commits this week

Compiler for JAX, TF, PyTorch targeting GPU, TPU, and CPU from a unified IR

4.5k
202

ModelTC

1 repo·4.3k·5 commits this week

Lightweight, high-throughput Python-based LLM inference and serving framework

4.3k
5

Predibase

1 repo·3.8k

Multi-LoRA inference server — serve thousands of fine-tuned adapters on a single GPU

3.8k

raullenchai

1 repo·3.6k·188 commits this week

The fastest local AI engine for Apple Silicon — 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling, drop-in OpenAI-compatible replacement

3.6k
188

Liquid AI

5 repos·3.3k·11 commits this week

Examples, tutorials and apps for Liquid AI LFM + LEAP SDK

2.4k
3

Speech-to-Speech audio models by Liquid AI

567

Minimal fine-tuning repo for LFM2, fully open-source

206
7

Example apps for LeapSDK

75

Liquid AI documentation

31
1

Luminal AI

1 repo·3.0k·9 commits this week

Rust-based deep learning compiler with a small static graph IR for fast, portable inference (CUDA, Metal, CPU)

3.0k
9

Fluid Inference

3 repos·2.9k·2 commits this week

On-device audio inference framework

2.7k
1

Fluid Inference core runtime

86

Rust text processing library for inference

46
1

Try Mirai

2 repos·1.8k·15 commits this week

Mirai's on-device inference runtime

1.7k
13

Mirai's LLaMA-based on-device model

89
2

UbiquitousLearning

1 repo·1.6k·2 commits this week

Multimodal LLM inference framework for mobile & edge

1.6k
2

ARM Software

1 repo·1.3k

ARM Neural Network SDK for ARM & Mali devices

1.3k

AMD ROCm

4 repos·1.3k·172 commits this week

AI Tensor Engine for ROCm — centralized repo for high-perf AI operators on AMD Instinct GPUs

551
98

AMD's graph inference engine for MI-series GPUs

329
30

ROCm fork of FlashAttention with Composable Kernel (CK) and Triton backends

238

AiTer Optimized Model — lightweight vLLM-like server built on AITER kernels for ROCm

169
44

Rudrank Riyam

1 repo·1.2k

iOS/macOS workbench for Apple's Foundation Models framework — recipes, guided labs, prompt playground, AFM CLI, FMFBench eval suite, and reusable Swift packages for on-device Apple Intelligence

1.2k

NimbleEdge

2 repos·526

NimbleEdge's deliteAI on-device inference framework

524

NimbleEdge fork of ExecuTorch with edge optimizations

2

ThunderAgent

1 repo·433

A simple, fast and robust program-aware agentic inference system

433

Picovoice

1 repo·317

Picovoice's on-device LLM inference engine

317

Zetic AI

5 repos·78

MLange sample applications

70

MLange extension library

5

iOS framework for MLange

2

iOS extension framework for MLange

1

MLange SDK documentation

0