Model Serving Open-Source Projects
Browse selected Model Serving open-source AI projects from the current public snapshot, with verified use cases, deployment notes, limitations, and sources.
Showing 40 indexable projects from the current snapshot.
ollama/ollama
Ollama simplifies downloading and running open large language models locally on macOS, Windows, or Linux without requiring paid services. Users can interact with models via a CLI, a REST API, or client libraries.
open-webui/open-webui
A self-hostable web interface for managing and interacting with LLM runners and AI features. It is designed to operate entirely offline and supports various model providers alongside built-in retrieval augmented generation.
huggingface/transformers
Transformers centralizes state-of-the-art model definitions for text, vision, audio, video, and multimodal tasks, providing a unified API for inference and training. It targets AI engineers, researchers, and developers who need to share tr…
Robbyant/lingbot-map
LingBot-Map is a feed-forward 3D foundation model for reconstructing scenes from streaming image or video data. It unifies coordinate grounding, dense geometric cues, and long-range drift correction to support efficient streaming 3D recons…
AUTOMATIC1111/stable-diffusion-webui
AUTOMATIC1111 Stable Diffusion WebUI is a Gradio-based interface for local Stable Diffusion image generation and management. It provides text-to-image, image-to-image, inpainting, upscaling, and training capabilities, with documented suppo…
vllm-project/vllm
vLLM provides high-throughput, memory-efficient inference and serving for large language models. It targets developers who need to serve models locally or in self-hosted environments with efficient attention memory management.
ggml-org/llama.cpp
llama.cpp is a dependency-free C/C++ library for running large language model (LLM) inference locally or in the cloud. It provides tools for model quantization, grammar-constrained generation, and API serving, requiring models to be in the…
BerriAI/litellm
LiteLLM provides a unified gateway and Python SDK for calling 100+ LLM providers using the OpenAI format, addressing the operational complexity of managing disparate SDKs and authentication patterns. It includes production-ready features s…
2noise/ChatTTS
ChatTTS is a text-to-speech model optimized for dialogue scenarios, providing fine-grained prosodic control and multi-speaker support. It is released for academic purposes only under a CC BY-NC 4.0 license and supports English and Chinese.
hiyouga/LlamaFactory
A framework that provides zero-code fine-tuning for over 100 large language models and vision-language models through command-line and web interfaces. It enables continuous pre-training, supervised fine-tuning, and preference alignment usi…
unslothai/unsloth
Unsloth provides a local web UI and Python library for running and training large language models with reduced VRAM usage and faster speeds. It supports text, audio, embedding, and vision models, and operates via local or self-hosted execu…
xtekky/gpt4free
This project aggregates multiple LLM and media generation providers behind a unified OpenAI-compatible interface. It provides Python and JavaScript clients, a local GUI, a REST API, and an MCP server under a community-first license.
keras-team/keras
Keras 3 is a high-level deep learning framework that lets you build, train, and run models across JAX, TensorFlow, PyTorch, and OpenVINO. It is designed for developers and researchers who need to accelerate model development and scale trai…
ultralytics/ultralytics
Ultralytics provides SOTA YOLO models for real-time computer vision across object detection, segmentation, classification, pose estimation, and tracking. It is accessible via a unified CLI and Python API for model training, evaluation, and…
songquanpeng/one-api
This project consolidates API keys and distributes LLM API requests across multiple model providers using a standard OpenAI-compatible API format. It provides load balancing across configured channels.
jingyaogong/minimind
An open-source educational project that enables individuals to train and understand a 64M-parameter large language model from scratch in approximately 2 hours. It provides a full PyTorch-based training pipeline, model architecture code, op…
mudler/LocalAI
LocalAI is a self-hosted AI engine that runs large language, vision, audio, image, and video models behind a single API without requiring a GPU. It provides drop-in OpenAI, Anthropic, and ElevenLabs API compatibility while keeping data wit…
oobabooga/textgen
A desktop application for running local large language models privately, without external APIs or cloud services, and with zero telemetry. Users can chat with text and vision models, call custom functions, train LoRAs, generate images, and…
microsoft/qlib
Qlib is a modular framework for applying AI and machine learning to quantitative investment research. It supports building complete workflows from data collection through model training and backtesting, with configurable components that ca…
coqui-ai/TTS
A deep learning toolkit for Text-to-Speech generation that provides pretrained models in over 1100 languages, voice cloning, and model training tools. It is intended for developers, AI engineers, and researchers building speech synthesis a…
Kong/kong
Kong is a cloud-native, platform-agnostic gateway designed to orchestrate microservices, conventional API traffic, and agentic LLM and MCP traffic. It provides centralized proxy functionality, AI governance, and extensible deployment model…
janhq/jan
Jan is a desktop application that lets users download and run large language models locally or connect to cloud providers, maintaining data control without mandatory cloud dependence. It supports text prompts and provides an OpenAI-compati…
ray-project/ray
Ray provides a unified framework for scaling Python and AI applications from a laptop to a cluster. It offers distributed abstractions for tasks and actors, alongside libraries for data processing, training, tuning, reinforcement learning,…
gradio-app/gradio
Gradio lets Python developers build interactive web applications for machine learning models and Python functions without requiring JavaScript, CSS, or web hosting experience. It provides UI components, sharing features, and client APIs fo…
Alishahryar1/free-claude-code
A local proxy and Admin UI that routes Claude Code, Codex, and Pi agent traffic to free, paid, or local model providers. It enables developers to select, validate, and manage multiple providers from a single interface.
chatanywhere/GPT_API_free
This project offers free and paid API access to multiple large language models through a unified OpenAI-compatible protocol. It uses dynamic network acceleration so users in regions with connectivity restrictions can connect directly witho…
TencentARC/GFPGAN
GFPGAN is a Python library and CLI tool for real-world blind face restoration, leveraging generative priors from a pretrained face GAN. It is operated locally via command-line inference, accepts image or folder inputs, and outputs restored…
google-ai-edge/mediapipe
MediaPipe provides cross-platform libraries and tools for building and deploying on-device machine learning pipelines that process images, video, and text locally on the user's device. It is intended for developers who need to integrate cu…
1Panel-dev/1Panel
1Panel is a web-based VPS control panel that enables server and Docker management through a visual GUI, replacing the need for CLI memorization. It natively integrates an AI agent runtime for deploying LLMs and includes an app marketplace…
musistudio/claude-code-router
Claude Code Router is a local control plane that exposes one stable endpoint for routing requests across multiple LLM providers and coding agents. It consolidates provider credentials, routing rules, and observability into a single self-ho…
shiyu-coder/Kronos
A pre-trained foundation model specialized for financial candlestick (K-line) sequences, supporting forecasting and fine-tuning for quantitative data analysis. It is designed for data teams and researchers who need to process high-noise fi…
OpenBMB/VoxCPM
VoxCPM is a tokenizer-free text-to-speech model that outputs 48kHz audio, supports 30 languages, and creates or clones voices from text descriptions or short reference clips. It is suitable for developers and creators who need self-hosted,…
mckaywrigley/chatbot-ui
An open-source AI chat application that connects to any model and can be deployed locally or in the cloud. It requires a Supabase backend for database and authentication and relies on terminal usage for setup and configuration.
fishaudio/fish-speech
Fish Speech is a multilingual text-to-speech and voice cloning model that supports over 80 languages and multi-speaker generation. It accepts text and short reference audio to generate speech audio without requiring phoneme preprocessing.
datawhalechina/self-llm
An open-source tutorial collection that helps learners configure environments, deploy, apply, and fine-tune more than 50 mainstream large language models across Linux systems and multiple hardware platforms. It targets Chinese-speaking beg…
onyx-dot-app/onyx
Onyx is a self-hostable application layer and chat interface that brings RAG, web search, deep research, and custom agents to a hosted environment. Teams can deploy it via Docker, Kubernetes, source, or use a managed cloud service, provide…
sgl-project/sglang
SGLang is a serving framework for large language and multimodal models, providing runtime features such as RadixAttention and tensor parallelism across diverse hardware. It is designed for teams that need low-latency, high-throughput infer…
eugeneyan/applied-ml
A curated directory that organizes real-world machine learning papers and engineering blogs into 31 specific categories. It serves data teams and researchers investigating how organizations implement ML in production.
cheahjs/free-llm-api-resources
A curated, automatically updated list of legitimate free LLM API resources and their associated limits. Developers and researchers can use it to discover available models and rate limits from providers offering free access or trial credits.
invoke-ai/InvokeAI
InvokeAI is a self-hosted creative engine for generating and refining visual media using Stable Diffusion and other generative models. It provides a web interface, a Unified Canvas for in/out-painting, and node-based workflows for customiz…