LocalAI
With over 48,000 GitHub stars and monthly releases since March 2023, LocalAI is the self-hosted AI engine that replaces every OpenAI endpoint with a single Docker container running on your own infrastructure — serving chat completions, image generation, text-to-speech, speech-to-text, embeddings, vision, video generation, and function calling through identical API schemas that require zero application code changes. The composable backend architecture isolates each inference engine as a separate gRPC service running in its own OCI container, so llama.cpp, vLLM, SGLang, transformers, whisper.cpp, diffusers, MLX, Stable Diffusion, and Flux install on demand without touching the core, can run on separate machines, and a fault in one never affects others. Hardware acceleration spans NVIDIA CUDA 12 and 13, AMD ROCm, Intel oneAPI/SYCL, Apple Silicon Metal, Vulkan, and NVIDIA Jetson L4T — or runs entirely on CPU without any GPU. Built-in AI agents support autonomous tool use, retrieval-augmented generation, Model Context Protocol integration, and skill-based workflows directly in the web interface. The model gallery provides curated YAML configuration files for hundreds of models that install with a single command, while P2P federated inference distributes model shards across multiple machines for running models larger than any single node's memory. Multi-user API key authentication with quotas and role-based access enables team deployments. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
FreeLLMAPI
FreeLLMAPI collapses the chaos of 29 free LLM providers — Google AI, Cerebras, Groq, Mistral, OpenRouter, GitHub Models, Cohere, Cloudflare Workers AI, NVIDIA NIM, HuggingFace, SiliconFlow, Reka, Z.ai, and more — into a single /v1 endpoint that speaks both OpenAI and Anthropic protocols. The smart router selects the best available model for each request, automatically fails over to the next provider when rate limits hit, and tracks per-key token consumption so you never exceed a free-tier cap. Keys are stored with AES-256-GCM encryption and clients authenticate using a single unified bearer token, never exposing upstream provider credentials to downstream applications. The catalog tracks 251 model families across 358 provider/model endpoints with approximately 4 billion tokens per month of aggregate free-tier capacity, auto-refreshing from a signed manifest at freellmapi.co twice daily without requiring git pulls. Beyond chat completions, the proxy handles embedding, image generation, and audio/TTS endpoints, plus structured outputs with JSON schema forwarding, JSON healing, and format-ignore failover. An integrated MCP server at /mcp provides gateway introspection for coding agents, while the self-hosted OpenAPI reference at /v1/docs documents every route. Compatible with OpenAI SDKs, LangChain, LlamaIndex, Continue, Claude Code, and Hermes — just change base_url. Deploy via Docker, npm, or build from source. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Forge
Forge intercepts failing LLM tool calls and fixes them before they derail your agent workflow, applying rescue parsing, retry nudges, response validation, and step enforcement between your AI clients and local model backends. The proxy server mode drops in as a transparent intermediary speaking both the OpenAI chat-completions API and the Anthropic Messages API, so tools like Aider, Claude Code, Continue, and opencode connect through it without configuration changes. Under the hood, the WorkflowRunner provides a complete agentic loop manager with system prompt injection, tool execution, context compaction with configurable thresholds, and VRAM budgeting for consumer GPUs with 12-32 GB. SlotWorker enables priority-queued access to shared inference slots with automatic preemption for multi-agent architectures. The guardrails middleware exposes a two-method check-and-record API that wraps into any existing orchestration loop, providing malformed tool-call rescue parsing, retry nudge generation, required step enforcement, and prerequisite ordering without taking over execution control. Backend adapters support generic OpenAI-compatible endpoints, Ollama, llama-server, Llamafile, vLLM, and Anthropic with automatic model discovery and health checking. Architecture Decision Records document every design choice. Launched February 2026, already at 2,200+ GitHub stars. MIT licensed.
OpenLLM
OpenLLM serves any large language model as an OpenAI-compatible API endpoint from a single CLI command, handling model download, backend selection, quantization, and port binding automatically. It supports the full spectrum of popular models including Llama 3.3, Qwen2.5, DeepSeek, Mistral, and Phi3, choosing between vLLM and PyTorch inference backends based on hardware capabilities. When vLLM is available, continuous batching with PagedAttention achieves up to 23x throughput improvement over naive serving, while GPTQ and bitsandbytes quantization reduces memory requirements for GPU-constrained deployments. The server exposes a RESTful API on port 3000 with full OpenAI client library compatibility, enabling drop-in replacement for commercial providers in any application using the standard chat completions format. A built-in web chat UI at the /chat endpoint provides immediate interactive testing without external clients. Custom model repositories allow teams to maintain private catalogs of fine-tuned models alongside the default repository that tracks the latest releases. Deployment workflows generate production-ready Docker images automatically, with Kubernetes manifest support for orchestrated scaling. Native integration with LangChain and LlamaIndex supports RAG pipelines, Transformers Agents enables tool-calling workflows, and HuggingFace Hub handles model discovery. Server-Sent Events enable real-time token streaming across all API endpoints. Backed by BentoML's production ML infrastructure. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.