Open WebUI
Large language models get a polished front end that can run fully offline: Open WebUI is the self-hosted front end of choice. It talks to local model runners, primarily Ollama, and to any OpenAI-compatible API, so LM Studio, vLLM, Groq, Mistral, OpenRouter, and cloud providers all plug into the same chat interface and can be mixed per conversation. RAG is built in: upload files to knowledge bases or reference them in chat with the # command, backed by a choice of nine vector databases (ChromaDB and PGVector officially maintained) and multiple extraction engines including Tika and Docling, with hybrid BM25-plus-vector search and cross-encoder reranking. Web search results from providers like SearXNG, Brave, and Tavily inject directly into conversations. Extensibility comes from Python tools and functions that run inside the chat, a Pipelines plugin framework, and native MCP support. Multi-user features include RBAC, SSO, and group permissions, and the instance itself exposes an OpenAI-compatible API your own apps can call.
AnythingLLM
Chat with your own documents: AnythingLLM, from Mintplex Labs, wraps retrieval-augmented generation (RAG) in an open-source application anyone can run. You organize content into workspaces, each an isolated namespace with its own documents, vector embeddings, chat history, and settings, so one instance can hold several separate knowledge bases. Upload PDFs, DOCX, TXT, and other formats, or scrape web pages; the built-in collector parses and chunks them into a vector database (LanceDB by default, with Pinecone, Chroma, Qdrant, and others supported). Answers cite their source documents. It works with both cloud LLMs (OpenAI, Anthropic, Gemini) and local ones via Ollama or LM Studio, and the embedding model is separately configurable. Beyond RAG chat, it includes AI agents that can browse the web and run tools, an embeddable chat widget for your website, a developer API, and multi-user mode with admin, manager, and default roles plus per-workspace access control. Context assembly is smarter than naive RAG: pinned documents, attached files, vector search hits, and recent chat history are combined under a token budget so the model's context window is filled efficiently, and each workspace supports multiple independent conversation threads against the same knowledge base. Because the embedding model, vector store, and chat LLM are all independently swappable, you can move between providers without re-ingesting a single document. The stack is Node.js with a React frontend, MIT-licensed.
Hermes Agent
OpenRouter's most-used application by token volume — over 17 trillion tokens processed — Hermes Agent is an open-source autonomous agent built by Nous Research that lives on your server and gets more capable every day. Define a goal in natural language and Hermes plans sub-tasks, executes them through tool integrations, observes results, handles errors, and refines until the job is done or it genuinely needs your input. Persistent memory with full-text search and LLM summarization lets it recall context across sessions, and an agent-created skills system self-improves after complex tasks. A messaging gateway connects Telegram, Discord, Slack, WhatsApp, Signal, and 16 more platforms with cross-channel conversation continuity. A built-in cron scheduler runs daily reports, nightly backups, and weekly audits unattended. Subagent spawning parallelizes workstreams, and six terminal backends — local, Docker, SSH, Singularity, Modal, and Daytona — fit any infrastructure. Works with any LLM provider: Nous Portal, OpenRouter for 400+ models from 70+ providers, OpenAI, Anthropic, or your own endpoint. The API key you supply powers all LLM calls; billing goes through your own account. Running on a dedicated VPS with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Onyx
Formerly known as Danswer and now backed by over 31,000 GitHub stars with 253 releases, Onyx delivers a production-ready AI platform that turns any LLM into a context-aware enterprise assistant connected to your organization's actual knowledge. The agentic RAG pipeline combines BM-25 keyword search with prefix-aware embedding models in a hybrid index, then deploys AI agents to retrieve, verify, and synthesize answers with source citations from over 40 connected workplace tools including Google Drive, Confluence, Slack, Notion, Jira, SharePoint, GitHub, and Linear. Custom AI assistants with configurable prompts, backing knowledge sets, and document-level access control enable specialized agents for engineering, sales, support, and research workflows. The platform supports every major LLM provider — Anthropic Claude, OpenAI, Google Gemini, plus self-hosted options via Ollama, LiteLLM, and vLLM for fully air-gapped deployments. Beyond chat, Onyx provides web search with Serper, Google PSE, Brave, and SearXNG integration, an in-house web crawler, code execution, file creation, and multi-step deep research with report generation. Enterprise features include SSO via Google OAuth, OIDC, or SAML with SCIM provisioning, role-based access control, usage analytics by team and agent, query history auditing, PII removal through custom code hooks, and full whitelabeling. Deploy via Docker Compose on any infrastructure. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed (Community Edition).
LobeHub
With over 82,000 GitHub stars and 700,000+ downloads, LobeHub has evolved from its origins as LobeChat into a comprehensive multi-agent AI collaboration platform where humans and autonomous agent teams co-evolve. The platform's Agent Harness architecture functions as an operating system between AI models and applications, handling prompt presets, tool orchestration, lifecycle hooks, planning, filesystem access, and sub-agent management across 25+ model providers including OpenAI, Anthropic Claude, Google Gemini, DeepSeek, Mistral, Groq, AWS Bedrock, Azure OpenAI, and local models through Ollama. Agent Groups enable sophisticated collaboration with sequential, parallel, iterative, and debate orchestration modes, allowing multiple specialized agents to tackle complex workflows simultaneously. The Agent Builder creates production-ready agents from natural language descriptions with auto-configuration, drawing from a marketplace of 505+ pre-built agents and 10,000+ MCP-compatible skills and plugins. Pages provide collaborative document editing with multi-agent co-authoring, while Schedules automate agent runs around the clock without human supervision. The knowledge base leverages PostgreSQL with pgvector for RAG-powered retrieval, and Personal Memory gives agents transparent, editable context that evolves through continual learning. The full self-hosted stack deploys via Docker Compose with PostgreSQL, Redis, RustFS for S3-compatible storage, and SearXNG for private web search, all configurable through environment variables. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. LobeHub Community licensed.
Odysseus
Agents with tool use, deep research, a document editor, an IMAP/SMTP email client with AI triage, notes, tasks, and a CalDAV-synced calendar - Odysseus bundles all of it into one open-source, self-hosted AI workspace. It runs local models through Ollama, vLLM, or llama.cpp and cloud APIs like OpenAI and OpenRouter, with a hardware-aware Cookbook that scans your machine and recommends quantized models that fit. Persistent memory uses ChromaDB with hybrid vector-plus-keyword retrieval, web search runs through a bundled SearXNG instance, and agents can use MCP servers, files, and shell access with safety controls, plus custom skills and scheduled agent tasks. A blind Compare mode runs side-by-side model duels with identities hidden and accumulates Elo-style ratings from your votes, so model selection is based on your actual workloads rather than leaderboard claims. Deep research mode - adapted from the Tongyi DeepResearch approach - reads sources through SearXNG and produces cited reports, while the email client tags, summarizes, sets reminders, and drafts replies locally rather than through a third-party mail AI. The writing-first document editor adds AI edits, Markdown and HTML support, and version history. The stack is Python 3.11 with FastAPI, SQLite for state, and a vanilla JS frontend, licensed AGPL-3.0 with zero telemetry. Because agents can read email and execute commands, keep authentication enabled and never expose it as a public unauthenticated service.
Hermes Studio
The most comprehensive open-source control plane for Hermes Agent — a full workspace combining AI chat, visual workflows, multi-agent orchestration, coding agent management, and platform channel integration in one self-hosted dashboard. Real-time chat streaming over Socket.IO connects to any OpenAI-compatible backend including Ollama, OpenAI, Anthropic, and custom endpoints with multi-session management, tool call expansion, inline file previews for HTML, PDF, DOCX, images, and source code, plus profile-scoped uploads and workspace attachments. The visual workflow builder provides a Vue Flow canvas for constructing DAG-structured pipelines with directed edges, conditional routes, approval gates, loops, and live execution with per-node status updates. Platform channel integration configures Telegram, Discord, Slack, WhatsApp, Matrix, Feishu, DingTalk, QQBot, WeChat, and WeCom bots from one page with credential management and per-platform behavior settings. Multi-agent group chat rooms enable real-time messaging with @mention routing, automatic context compression, and SQLite message persistence. The coding agent panel installs, launches, and monitors Claude Code and Codex with built-in terminal, session history, and file diffs. Usage analytics track token consumption, estimated costs, cache hit rates, and 30-day daily trends with model distribution charts. Kanban boards plan and track agent work alongside cron job scheduling for recurring tasks. Deploy via Docker, npm CLI, or desktop installer for Windows, macOS, and Linux. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
OpenClaw VPS
A personal AI assistant that remembers what it learns and reaches you wherever you are — OpenClaw is an open-source agent gateway built by the OpenClaw Foundation with 247,000+ GitHub stars. It connects to 200+ LLM models through providers like Anthropic, OpenRouter, and OpenAI, and meets you on 21+ messaging channels: Telegram, Slack, Discord, WhatsApp, Signal, iMessage, Matrix, and more. Persistent memory with full-text search lets the agent recall context across sessions, and a self-improving skills system means it gets more capable the longer it runs. Voice wake words and talk mode enable hands-free interaction on macOS, iOS, and Android. A live canvas provides an agent-driven visual workspace. Built-in browser automation, cron scheduling for unattended tasks, and subagent spawning for parallel workstreams round out the toolset. The gateway architecture keeps all sessions, credentials, and conversation history on your own server — nothing transits a third-party cloud unless you choose to connect one. The API key you provide for your chosen LLM provider powers the underlying calls; billing goes through your own account. Running on a dedicated VPS with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Memoh
Memoh delivers an open-source multi-agent platform where every AI agent gets its own computer — not a chat window but a fully isolated container with dedicated filesystem, desktop environment, browser, network stack, and persistent long-term memory that survives across sessions, days, and platforms. The containerd-based runtime ensures each bot operates in complete isolation with snapshot and data import/export capabilities. The memory engine uses LLM-driven fact extraction with hybrid retrieval combining dense embeddings via Qdrant, sparse vectors, and BM25, plus 24-hour context loading and automatic compaction — with Mem0 and OpenViking as drop-in alternatives. Ten communication channels connect agents to users through Telegram, Discord, Lark, QQ, Matrix, WeCom, WeChat, Email, Web UI, and group chats with cross-platform identity binding. MCP tool calling enables agents to interact with external services, while browser automation drives GUI workflows for web research and data extraction. Agent hosting supports running external coding agents like Codex and Claude Code inside Memoh workspaces via ACP with per-bot configuration. Scheduled tasks run without human triggers, and agents proactively reach out when needed. The web dashboard built with Vue 3 and Tailwind CSS provides streaming chat, tool call visualization, file management, model and provider configuration, and bot lifecycle management. Deploy via Docker Compose with PostgreSQL, Qdrant, sparse service, and the Go backend server. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
LibreChat
Every major model provider behind one ChatGPT-style interface: LibreChat spans OpenAI, Anthropic, Google, Azure, AWS Bedrock, Vertex AI, Groq, Mistral, OpenRouter, DeepSeek, and any OpenAI-compatible endpoint including local Ollama. You can switch models mid-conversation and compare providers without changing tools. Its Agents framework builds no-code custom assistants with tool access via Model Context Protocol servers, file search over uploaded documents through an optional pgvector-backed RAG service, and a sandboxed Code Interpreter that executes Python, JavaScript, Go, C++, Java, PHP, and Rust. Artifacts render React components, HTML, and Mermaid diagrams directly in chat, and image generation works through DALL-E and other configured providers. Multi-user support is enterprise-grade, with OAuth, SAML, LDAP, and two-factor authentication, per-user conversation history in MongoDB, and Meilisearch-powered search across all messages and files, plus reusable presets, forkable threads, and persistent memory across conversations. The economics favor teams: instead of a ChatGPT Plus seat per person, everyone shares one instance billed per API token, with access to every provider rather than one - and providers see individual API calls, not your accumulated organizational knowledge. Deployment is Docker Compose; API keys and endpoints are configured through .env and librechat.yaml.
NanoClaw
NanoClaw delivers a radically simple alternative to OpenClaw — a single Node.js process and a handful of files that provide the same core functionality with true container-level security isolation. Agents execute inside Docker containers on Linux or Apple Containers on macOS, where even root access inside the sandbox cannot reach the host filesystem. The platform natively runs Claude Code via Anthropic's official Claude Agent SDK, with drop-in alternatives including OpenAI Codex, OpenRouter via OpenCode, Google, DeepSeek, and local open-weight models via Ollama — configurable per agent group. Multi-channel messaging connects WhatsApp, Telegram, Discord, Slack, Microsoft Teams, iMessage, Matrix, Google Chat, Webex, Linear, GitHub, WeChat, and email via Resend, installed on demand through skill commands. Each agent group receives its own CLAUDE.md memory file, isolated filesystem, container sandbox, and session state — a prompt injection in one group cannot exfiltrate data from another. The OneCLI Agent Vault handles credentials so agents never hold raw API keys, while approval-gated self-modification allows agents to request new packages or MCP servers that administrators must authorize. Scheduled tasks run recurring jobs inside containers with message delivery back to users. The setup script handles dependencies, authentication, and container configuration through Claude Code conversation. Deploy on any Docker-capable Linux server. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
OpenSquilla
Claiming 60-80% token cost reduction compared to flat single-model deployments and backed by 6,500+ GitHub stars, OpenSquilla delivers an intelligent AI agent runtime where a local ML classifier evaluates every turn on message length, code blocks, keyword patterns, and semantic embeddings before routing it to the optimal model tier from C0 through C3. The pluggable provider layer connects natively to TokenRhythm, OpenRouter, OpenAI, Anthropic, Ollama, DeepSeek, Gemini, DashScope, Moonshot, Mistral, Groq, Zhipu, SiliconFlow, vLLM, LM Studio, and additional compatible backends with primary-plus-fallback selection. The four-tier cognitive memory architecture spans working, episodic, semantic, and raw layers with vector-semantic and BM25 retrieval powered by on-device ONNX embeddings that never leave your infrastructure. Security isolation operates at the syscall level via Bubblewrap on Linux and Seatbelt on macOS, complemented by policy-based execution controls and prompt injection protections. The unified TurnRunner executes identically across the Vue-based control console Web UI, terminal CLI, and chat channel integrations including Slack and Discord, ensuring consistent tool dispatch, retry logic, and decision logging regardless of entry point. Built-in skills cover deep research, multi-search-engine queries, document generation for DOCX, PPTX, XLSX, and PDF formats, GitHub integration, cron scheduling, and bounded subagent delegation. Per-agent workspaces with durable session storage provide transcript replay, context state management, and per-call cost tracking with automatic quota enforcement. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache-2.0 licensed.
FastGPT
FastGPT lets you build production AI agents and knowledge base chatbots through a visual drag-and-drop workflow editor, connecting any LLM provider to your documents with retrieval-augmented generation that cites sources and reduces hallucination. The workflow canvas chains LLM calls, conditional branching, HTTP requests, code sandbox execution, and plugin nodes into complex conversation flows and agent skill pipelines without writing backend code. The knowledge base engine ingests documents in ten formats (TXT, Markdown, HTML, PDF, DOCX, PPTX, CSV, XLSX, URL scraping, and CSV batch import) then applies automatic chunking, hybrid vector retrieval with semantic reranking, and QA-pair splitting to deliver accurate, citation-backed answers. FastGPT connects to virtually any LLM provider through its AI Proxy aggregation layer: OpenAI GPT-4o, Anthropic Claude, Google Gemini, DeepSeek, Qwen, ERNIE Bot, and models hosted via Ollama all work through a unified OpenAI-compatible API. Bidirectional MCP support enables agents to call external tools and expose their own capabilities to other systems. Completed applications can be shared via login-free links, embedded as iframe widgets, or integrated with WeCom, Lark, DingTalk, and WeChat Official Accounts through the published REST API. Application operation logs, conversation annotation, and per-model usage analytics provide full lifecycle governance for compliance-sensitive deployments. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. FastGPT Open Source License (Apache 2.0 based) licensed.
GoRaven
GoRaven transforms AI chat from a question-answer window into a full engineering workstation where agents read files, write code, run shell commands, query databases via MCP tools, and deliver structured results — orchestrating across OpenAI, Claude, DeepSeek, Gemini, Qwen, GLM, and Ollama with task-based routing that allocates the right model for each job based on cost and capability. Built on a Go backend using the Freedom framework with Iris HTTP and a React/TypeScript frontend powered by Vite and Tailwind CSS, each user operates in an isolated workspace with team-shared project areas and centrally managed model quotas. The skill marketplace packages prompts, scripts, and workflows as reusable installable units with automatic dependency resolution and centralized versioning. MCP toolchain integration connects agents to internal APIs, databases, private services, and CLI tools so they query data, invoke services, and trigger actions directly. RAG-powered knowledge bases ingest policies, documentation, and business data for real-time retrieval during planning, coding, and Q&A with source attribution. Long-running task support decomposes complex work through a main agent coordinating sub-agents that execute in parallel across sessions. Plugin hooks inject custom logic at conversation start and end, tool calls, and SSE event streams without forking core code. The operations dashboard tracks usage metrics, model consumption, and team activity. Supports SQLite, MySQL, or PostgreSQL with Redis or local memory caching. Deploy with a single Docker command. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Khoj
A self-hosted "second brain": Khoj indexes your own files and answers questions from them, parsing Markdown (whole Obsidian vaults included), org-mode, PDF, Word, plain text, Notion pages, GitHub repositories, and images described by a vision model, then embedding everything with sentence-transformers into a vector index for semantic search and RAG with cited sources. Any LLM backend works: local models like Llama, Qwen, or Mistral via Ollama, or cloud models like GPT, Claude, and Gemini. You can build custom agents, each with its own persona, scoped knowledge base, chat model, and tools such as web search and code execution. Scheduled automations run recurring research and deliver newsletters or notifications to your inbox, and research mode performs multi-hop web searches with inline citations. Access it from a browser, the Obsidian plugin, Emacs, desktop, or WhatsApp - all clients connect to the same self-hosted instance, making Khoj one of the few AI assistants Emacs users can point at decades of org files. Semantic search means recall works without exact keywords: "that paper about forecasting with transformers" surfaces the right PDF even when you cannot remember its title. Switching LLM backends never requires re-indexing your documents, and with a local model via Ollama, even inference stays on hardware you control - journals, research, and private notes are never sent anywhere. Python/FastAPI stack, AGPL-licensed, with PostgreSQL storage.
Casibase
Casibase lets organizations build AI-powered knowledge bases that answer questions from their own documents, connecting to 30+ model providers through a unified admin interface with RAG retrieval and multi-agent orchestration via MCP and A2A protocols. The platform plugs into OpenAI GPT-4o, Anthropic Claude, Meta Llama, Google Gemini, DeepSeek, Ollama local models, HuggingFace, Azure OpenAI, and additional providers, while embedding APIs from OpenAI Ada and Baidu handle vector representation of ingested documents. Document ingestion parses TXT, Markdown, DOCX, PDF, CSV, XLSX, and PPTX files with intelligent chunking strategies for optimal retrieval accuracy. The built-in chat interface provides real-time AI conversations with manual session handover for human agent escalation, and comprehensive chat session logging enables audit trails for compliance. Enterprise identity management integrates Casdoor for Single Sign-On supporting GitHub, Google, WeChat, and OIDC providers with fine-grained access control via the Casbin permission engine. The multi-tenant architecture supports isolated knowledge bases per organization with role-based user management and configurable storage, model, and embedding providers per tenant. The React frontend with Ant Design v5 provides a polished admin dashboard for managing providers, knowledge stores, chat sessions, and user access, while the Go backend with Beego framework handles API logic with MySQL or MariaDB persistence. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
MindsHub
Backed by $50M+ from Benchmark, Y Combinator, and NVIDIA with 800+ contributors and 39,000+ GitHub stars, MindsHub Cowork is the unified AI workspace where open-source models handle entire projects — research, reporting, internal tools, scheduled operations — and return finished, shareable deliverables. The platform runs two interchangeable open-source agent harnesses, Anton and Hermes, swappable from a dropdown without losing context. A built-in Model Router pre-wires 25+ models spanning Anthropic Claude, OpenAI GPT, Google Gemini, DeepSeek, Qwen, Kimi, Grok, and MindsHub Air with automatic failover — no per-provider API keys required. A secure credentials vault connects BigQuery, PostgreSQL, Salesforce, HubSpot, Zendesk, Gong, Gmail, Google Drive, Notion, Linear, Stripe, and Slack, keeping secrets scoped per connection so agents never see raw keys. Agent output becomes publishable artifacts — documents, dashboards, apps, and code — each deployable to a live shareable URL. Cross-session persistent memory, a reusable skill library, and a background scheduler supporting hourly, daily, and weekly cadences enable autonomous recurring workflows. The architecture separates a React/Vite frontend (shipping as both Electron desktop app and web SPA) from a FastAPI backend with a versioned REST API at /api/v1 covering conversations, projects, artifacts, schedules, and connectors. Self-host via Docker Compose with nginx on port 3000 and the API on port 26866, or deploy on-prem, in a VPC, or air-gapped. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
PilotDeck
PilotDeck introduces a WorkSpace-first architecture where each project receives its own isolated file system, memory store, and skill set, preventing context bleed between parallel tasks. White-box memory makes generation, extraction, storage, and retrieval fully visible, letting users audit, edit, pin, and rollback individual entries when the agent misremembers, while Dream Mode consolidates memory fragments during idle windows. Smart Routing auto-detects task difficulty and sends complex calls to flagship models like Claude Sonnet or GPT-4o while routing simple requests to lighter models, achieving claimed 70% cost savings through on-device and cloud co-orchestration with TokenSaver tiering and sticky session binding. Always-on background execution keeps agents running after the user closes the browser, with Discovery and Cron-based scheduling for recurring workflows. The platform natively supports the Model Context Protocol for first-class MCP server integration, community skills via ClawHub on npm, lifecycle hooks intercepting PreToolUse and UserPromptSubmit events, and custom memory store providers. Multi-provider fallback automatically switches to backup providers on timeout or rate-limit errors. The WebSocket and HTTP gateway serves web, CLI, desktop, and Feishu IM channels from a single configuration. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.