831 applications
Pipelock screenshot thumbnail

Pipelock

Your AI coding agent has your API keys in its environment and unrestricted network access, which means one prompt injection away from sending those secrets anywhere. Pipelock closes that gap by sitting as a proxy between your agents and every outbound connection, scanning the actual content of HTTP, WebSocket, MCP, and Agent-to-Agent traffic before it leaves your server. An 11-layer scanner pipeline checks every request against 62 credential patterns covering AWS, GCP, Azure, GitHub, OpenAI, Anthropic, SSH keys, and database URLs, then inspects every response for prompt injection using 29 detection patterns with six-pass normalization that catches base64-encoded, leetspeak, and whitespace-obfuscated payloads. The MCP proxy wraps any Model Context Protocol server (stdio, HTTP, or WebSocket) with bidirectional scanning that detects tool description poisoning and mid-session rug-pull changes via SHA-256 fingerprinting. Every scanning decision produces a cryptographically signed action receipt that third parties can verify offline without trusting the agent or the vendor. The Operator Console provides a web dashboard for reviewing evidence scorecards, receipt timelines, agent sessions, enforcement decisions, and fleet posture at a glance. Cross-request taint tracking catches slow-drip exfiltration attempts that spread a secret across multiple calls. Canary tokens plant synthetic secrets that trip alerts the moment an agent tries to exfiltrate them. Pre-built Prometheus metrics and a Grafana dashboard provide real-time visibility into traffic volumes and block rates. Deploy on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
Dockhand screenshot thumbnail

Dockhand

Dockhand is a Docker management platforms, offering a modern alternative to Portainer with free OIDC SSO and vulnerability scanning that competitors gate behind paid tiers. Real-time container management provides start, stop, restart, and remove operations with live resource monitoring across CPU, memory, and network usage on a dashboard with real-time metrics. The visual Docker Compose editor enables stack creation and modification with syntax highlighting, while Git integration deploys stacks directly from repositories with webhooks and auto-sync for GitOps workflows. Vulnerability scanning powered by Grype and Trivy analyzes container images against CVE databases, with configurable auto-update scheduling that can trigger updates based on vulnerability severity criteria. The Hawser Go agent enables management of remote Docker hosts in Standard mode for LAN environments or Edge mode using outbound WebSocket connections for hosts behind NAT, firewalls, or dynamic IPs without exposing inbound ports. Interactive terminal sessions provide shell access into running containers, while the file browser enables uploading, downloading, and editing files directly within containers. Image management includes registry browsing, pull operations, and layer inspection alongside network and volume administration. The security-focused architecture builds its own OS layer from scratch using Wolfi packages via apko with every package explicitly declared. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. BSL 1.1 licensed, converting to Apache 2.0 in 2029.

Deploy
SiYuan screenshot thumbnail

SiYuan

Backed by over 45,000 GitHub stars and described as the tool that replaces Notion, Evernote, and Anki in a single Docker container, SiYuan is the privacy-first knowledge management system where every paragraph, heading, and list item is a uniquely addressable content block. The block-level architecture enables bidirectional links, transclusion embeds, and SQL query blocks that dynamically aggregate content across your entire workspace, while the knowledge graph visualization maps relationship networks between documents and blocks. Built-in databases support table views with relation and rollup columns, filter composition, sorting, and template-based calculations for structured data management alongside freeform notes. The FSRS spaced repetition engine turns any content block into a flashcard with scientifically calibrated review scheduling, eliminating the need for separate memorization tools. AI integration connects to OpenAI-compatible APIs for writing assistance, translation, summarization, and Q&A chat, with semantic search using embeddings and reranking for intelligent content retrieval. The Bazaar community marketplace delivers plugins, themes, templates, and widgets through a managed extension system with TypeScript plugin APIs. End-to-end encrypted synchronization works across S3-compatible storage, WebDAV servers, or SiYuan's own cloud service, while Tesseract OCR extracts searchable text from images and the web clipper captures pages from Chrome, Edge, and Firefox. Export targets include Markdown with assets, PDF, Word, and HTML. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Deploy
Coder screenshot thumbnail

Coder

With over 14,000 GitHub stars and enterprise adoption by security-conscious organizations, Coder transforms how development teams provision, manage, and secure their coding environments. Every workspace is defined as a Terraform template, meaning infrastructure engineers can standardize development environments across EC2 instances, Kubernetes pods, Docker containers, or any combination, while developers get self-service provisioning that launches in seconds rather than days of manual setup. The WireGuard-based networking layer establishes encrypted tunnels between developer machines and remote workspaces, providing low-latency access without exposing ports or configuring VPN concentrators. Automatic idle detection shuts down unused workspaces after configurable periods, directly reducing cloud compute costs for organizations running hundreds of developer environments. The Coder Agents feature introduces native AI coding capabilities where the agent loop executes entirely within the control plane on self-hosted infrastructure, keeping LLM API credentials out of individual workspaces and eliminating credential exfiltration risks. Centralized model governance allows platform teams to approve specific AI providers and models, set per-user spend limits, and maintain complete audit logs of all prompts, tool calls, and agent activity. IDE integration supports VS Code through a dedicated extension, JetBrains IDEs via Gateway and Toolbox plugins, and browser-based code-server for web access. The template registry provides pre-built configurations for common development stacks. DevContainer support builds environments from standard devcontainer.json specifications. Deploy on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Deploy
Gods Eye View screenshot thumbnail

Gods Eye View

God's Eye View turns your browser into a real-time spatial intelligence command center, rendering thousands of live aircraft, ships, satellites, earthquakes, traffic flows, and public cameras on a photorealistic 3D Earth powered by CesiumJS and Google Photorealistic 3D Tiles. Click any aircraft to see its transponder telemetry from OpenSky and adsb.lol, including route history, altitude, speed, and callsign; track live vessel positions worldwide through AIS beacon data from AISStream; or follow roughly 840 satellites color-coded by class using orbital elements from CelesTrak. A hands-free voice agent powered by the OpenAI Realtime API lets you ask the planet questions in natural language, and the globe annotates your answer directly in 3D space. Toggle FLIR mode for a thermal camera aesthetic, layer in NASA FIRMS wildfire data, switch between Google 3D, Bing aerial, and OpenStreetMap base layers, or tune into a geolocated world radio dial. Public CCTV cameras are projected into 3D city geometry with viewshed cones and direct-manipulation calibration. The cockpit mode provides a pilot-style briefing surface with mission-specific overlays. Entity inspection panels show detailed metadata for every tracked object, and shareable links let you send any scene configuration to a colleague. Ten of the thirteen live data layers work with zero API keys, while the required Google Maps key offers 1,000 free 3D tile sessions per month. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
OpenSquilla screenshot thumbnail

OpenSquilla

Claiming 60-80% token cost reduction compared to flat single-model deployments and backed by 6,500+ GitHub stars, OpenSquilla delivers an intelligent AI agent runtime where a local ML classifier evaluates every turn on message length, code blocks, keyword patterns, and semantic embeddings before routing it to the optimal model tier from C0 through C3. The pluggable provider layer connects natively to TokenRhythm, OpenRouter, OpenAI, Anthropic, Ollama, DeepSeek, Gemini, DashScope, Moonshot, Mistral, Groq, Zhipu, SiliconFlow, vLLM, LM Studio, and additional compatible backends with primary-plus-fallback selection. The four-tier cognitive memory architecture spans working, episodic, semantic, and raw layers with vector-semantic and BM25 retrieval powered by on-device ONNX embeddings that never leave your infrastructure. Security isolation operates at the syscall level via Bubblewrap on Linux and Seatbelt on macOS, complemented by policy-based execution controls and prompt injection protections. The unified TurnRunner executes identically across the Vue-based control console Web UI, terminal CLI, and chat channel integrations including Slack and Discord, ensuring consistent tool dispatch, retry logic, and decision logging regardless of entry point. Built-in skills cover deep research, multi-search-engine queries, document generation for DOCX, PPTX, XLSX, and PDF formats, GitHub integration, cron scheduling, and bounded subagent delegation. Per-agent workspaces with durable session storage provide transcript replay, context state management, and per-call cost tracking with automatic quota enforcement. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache-2.0 licensed.

Deploy
Flipt screenshot thumbnail

Flipt

Backed by 4,800+ GitHub stars and trusted by teams replacing LaunchDarkly and Split with a self-hosted solution, Flipt v2 delivers the first truly Git-native feature management platform that treats feature flags as code stored in your own repositories. The architecture eliminates all database dependencies by building immutable in-memory snapshots from YAML flag definitions on every Git commit, delivering sub-millisecond evaluation latency with zero external runtime dependencies beyond the single Go binary. Multi-environment support maps directly to Git abstractions — separate repositories per environment, different directories within the same repository, or different branches — enabling teams to use their existing branching strategy, pull request workflows, and code review processes for flag changes. The evaluation engine supports boolean flags, multivariate string and numeric variants, segment-based targeting with constraint rules, percentage rollouts, and namespace isolation. Native SCM integration with GitHub, GitLab, BitBucket, Azure DevOps, and Gitea creates merge proposals directly from the UI with GPG-signed commits. Flipt implements the OpenFeature Remote Evaluation Protocol with official providers for Go, Node.js, Python, Java, C#, Ruby, and Web SDKs, enabling vendor-agnostic flag evaluation across all services. The gRPC API with REST HTTP gateway exposes flag management, evaluation, and analytics endpoints. Offline mode continues serving flags when the source repository is temporarily unavailable. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Fair Core License (server) / MIT (client SDKs).

Deploy
Bifrost screenshot thumbnail

Bifrost

Bifrost is an open-source AI gateway that unifies 23+ LLM providers into a single OpenAI-compatible endpoint with automatic failover, semantic caching, and built-in cost governance, so one provider going down never takes your production AI application with it. Point your existing OpenAI or Anthropic SDK at Bifrost's local endpoint and gain access to OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Groq, Mistral, and Ollama without changing application code. Define fallback chains that automatically switch providers when one returns errors or exceeds latency thresholds, keeping response times stable during outages. The built-in web dashboard at port 8080 lets you configure providers, create virtual API keys, monitor live request traffic, and review analytics without editing configuration files. Semantic caching combines exact hash matching with vector similarity search via Weaviate, serving cached responses for identical or paraphrased prompts in sub-millisecond time to cut costs on repetitive workloads. The MCP gateway connects AI agents to external tools like filesystems, databases, and web APIs, exposing them to clients such as Claude Desktop and Cursor with per-key allow-lists. Four-tier budget hierarchy at customer, team, virtual key, and provider levels enforces spend caps, rate limits, and model restrictions across your organization. Extend functionality through custom Go plugins for analytics, monitoring, or security middleware. Native Prometheus metrics and OpenTelemetry distributed tracing give operations teams full production observability. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
OpenReplay screenshot thumbnail

OpenReplay

Backed by 12,400+ GitHub stars and positioned as the self-hosted alternative to FullStory and Hotjar, OpenReplay delivers the open-source session replay platform that keeps every byte of user behavior data on your own infrastructure. The JavaScript tracker captures pixel-perfect recordings of clicks, scrolls, form inputs, and navigation with automatic sensitive data masking, while simultaneously logging network requests, console errors, JavaScript exceptions, and Redux, VueX, MobX, NgRx, Pinia, and Zustand store state changes for complete technical context. DevTools mode reconstructs each session with full stack traces, GraphQL queries from Apollo and Relay, Fetch and Axios request payloads, CPU and memory metrics, and page speed waterfall charts — effectively giving developers a browser inspector tied to any user session. Product analytics surfaces conversion funnels, user journeys, click heatmaps, web vitals trends, and retention cohorts without requiring custom instrumentation. Co-browsing connects support agents to live user sessions with cursor control and WebRTC audio, enabling real-time assistance without third-party screen-sharing software. Integrations push session context into Sentry, Datadog, CloudWatch, Stackdriver, and Elastic for front-to-back debugging. Feature flags enable gradual rollouts with session-level targeting. The platform deploys to any cloud via Docker and Kubernetes with auto-scaling ingestion handling up to 50,000 sessions per month on the open-source edition. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPLv3 licensed.

Deploy
Kavita screenshot thumbnail

Kavita

Manga, comics, ebooks, and light novels get a streaming-service-style home in Kavita - a fast, cross-platform reading server for the DRM-free collection you share with family and friends. It natively serves CBZ, CBR, CB7, ZIP/RAR/7z archives, raw images, EPUB, and PDF, with hand-crafted web readers per format: webtoon scrolling, single and dual-page spreads with advanced caching for the comic reader, and a book reader with adjustable fonts, spacing, margins, color themes, and column modes. Reading progress tracks per user, so everyone resumes exactly where they stopped on any device. Metadata parses from filenames, ComicInfo.xml, and EPUB fields, feeding index-backed search, smart filters, collections, reading lists with CBL import, and Want to Read queues. Role-based user management covers age restrictions, per-library access, and OIDC authentication. An OPDS feed connects third-party clients - Panels on iOS, Librera on Android, KOReader on e-ink devices - and a comprehensive REST API supports custom integrations. EPUB annotation and highlight support, custom theming, and full localization round it out. Built with .NET and Angular, it handles 50,000+ file libraries without strain; optional Kavita+ adds AniList scrobbling, recommendations, and external metadata.

Deploy
Backstage screenshot thumbnail

Backstage

Adopted by over 3,400 companies and backed by 34,000 GitHub stars, Backstage is the open-source developer portal framework created by Spotify and now hosted by the Cloud Native Computing Foundation. The centralized Software Catalog registers every service, library, data pipeline, website, and ML model in your organization using YAML metadata files stored alongside code in GitHub, GitHub Enterprise, or GitLab, tracking ownership, lifecycle status, and dependency relationships across your entire ecosystem. Software Templates provide self-service infrastructure provisioning where developers fill out a form and Backstage automatically scaffolds new repositories, CI/CD pipelines, and cloud resources following your organization's standardized best practices. TechDocs renders Markdown documentation directly alongside the services it describes using a docs-like-code approach powered by MkDocs, with the TechDocs Addon Framework for extending the reading experience. The Search Platform indexes content across the catalog, TechDocs, Confluence, and Stack Overflow through configurable search backends. Kubernetes monitoring built specifically for service owners rather than cluster admins displays pod health, logs, and deployment status across any cloud provider or managed Kubernetes service. The plugin ecosystem includes over 230 open-source integrations covering CI/CD systems like GitHub Actions and GitLab Pipelines, monitoring platforms like Datadog and Grafana, cloud providers including AWS and Azure, plus specialized plugins for security scanning, cost management, API documentation, PagerDuty incident management, and Lighthouse website auditing. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache-2.0 licensed.

Deploy
Yuvomi screenshot thumbnail

Yuvomi

Replace a dozen household subscriptions with one app that keeps your family's tasks, calendars, meals, budget, and shopping in a single private place you control. Nineteen modules cover the daily rhythms of a household: assign chores with deadlines and priorities, plan the week's dinners and push the ingredients to a shared shopping list in one tap, track income and expenses by category with automatic subscription renewal warnings, and log health activities from vitals to medications. Two-way sync connects to Google Calendar via OAuth, iCloud and Nextcloud via CalDAV, and Outlook via Microsoft Graph, so every family member's phone shows the same schedule without switching apps. CardDAV keeps contacts in sync with external address books. The budget module handles accounts, loans, split expenses, and per-category planning with savings goals and month-over-month comparisons. Documents attach to tasks, receipts link to transactions, and a checked-off grocery item books itself back into the pantry inventory with its quantity. A Kanban board, recurring task schedules, and a rewards ledger that pays out points for completed chores keep kids and adults accountable. Wall mode turns a kitchen tablet into a glanceable family dashboard, and the Immich screensaver shows your own photos when the screen goes idle. API tokens and a built-in MCP endpoint let AI agents and third-party tools interact with every module programmatically. Deploy on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
MLflow screenshot thumbnail

MLflow

Trusted by thousands of organizations with over 30 million monthly downloads and 20,000+ GitHub stars, MLflow is the largest open-source AI engineering platform providing end-to-end lifecycle management for traditional ML models, LLMs, and AI agents. The OpenTelemetry-based tracing system captures complete request flows through any LLM provider or agent framework — including OpenAI, LangChain, DSPy, Vercel AI, PydanticAI, and smolagents — with one-line auto-instrumentation that tracks inputs, outputs, token usage, and costs at every intermediate step. MLflow's evaluation engine offers 50+ built-in metrics and LLM judges for systematic quality assessment, detecting issues across correctness, latency, adherence, relevance, and safety dimensions before code reaches production. The Prompt Registry versions, tests, and deploys prompts with full lineage tracking while automated optimization algorithms improve prompt performance using evaluation feedback. The AI Gateway provides a unified API endpoint for all LLM providers, enforcing rate limits, cost controls, and access policies across the organization. MLflow 3.0 introduces the LoggedModel abstraction linking traces, metrics, and prompts to specific model versions across Python, TypeScript, Java, and R SDKs. The model registry manages deployment workflows with automated quality gates, while experiment tracking records parameters, metrics, and artifacts across training runs. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache License 2.0 licensed.

Deploy
ToolJet screenshot thumbnail

ToolJet

Retool's job, self-hosted: ToolJet is an open-source low-code platform for building internal tools, dashboards, and admin panels. Apps are assembled in a drag-and-drop visual builder with 60+ responsive components, including tables, charts, forms, and lists, and connected to 80+ data sources: PostgreSQL, MySQL, MongoDB, REST and GraphQL APIs, cloud storage, and common SaaS tools. When visual configuration is not enough, you can run JavaScript or Python inline for queries and transformations. A built-in no-code database (ToolJet Database) covers apps that need their own tables without provisioning an external database, Workflows add node-based automation for background jobs with dedicated worker containers and a Redis-backed queue, and multi-page apps with multiplayer editing, inline comments, and mentions support team development. Security is designed for internal data: credentials are AES-256-GCM encrypted, data flows proxy-only through your server so database contents never reach a third-party cloud, and granular per-app access control plus SSO gate each tool. Where Retool-style platforms bill per builder and sometimes per end user, the self-hosted Community Edition serves unlimited builders and users at hosting cost, and full source availability means the platform itself can be forked, audited, and extended. The stack is Node.js and React on PostgreSQL, deployed via Docker.

Deploy
DataHub screenshot thumbnail

DataHub

DataHub maps your entire data ecosystem into a searchable, governed catalog where every table, pipeline, dashboard, and metric is discoverable and traceable from source to consumer. Originally built at LinkedIn to manage metadata at hyperscale and proven to handle 10 million+ assets and billions of relationships in production, the platform is now trusted by 3,000+ organizations including Netflix, Visa, Slack, and Pinterest. The Spring Java backend (GMS) exposes both GraphQL and OpenAPI REST endpoints, while the React frontend delivers an intuitive interface for searching, browsing, and governing data assets. The Python-based ingestion framework provides 80+ production-grade connectors extracting deep metadata from Snowflake, BigQuery, Redshift, Databricks, dbt, Airflow, Spark, Kafka, Looker, Tableau, Power BI, Superset, PostgreSQL, MySQL, Hive, Glue, S3, Iceberg, and Unity Catalog through pull-based scheduled crawls and push-based emission via Python and Java SDKs. Automatic table-level and column-level lineage detection uses SQL parsing with 97-99% accuracy, tracing data flows from ingestion pipelines through warehouses to BI dashboards. Real-time metadata streaming via Kafka keeps the catalog continuously synchronized as schemas evolve and pipelines execute. The governance layer provides business glossary management, tag propagation along lineage graphs, domain-based organization, and fine-grained access control policies. DataHub Actions triggers automated responses to metadata changes, enabling notifications, quality checks, and downstream workflows. Elasticsearch powers full-text search with faceted filtering across entities. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
LibreDesk screenshot thumbnail

LibreDesk

LibreDesk unifies live chat, email, and future channel integrations into a single agent inbox where every customer conversation converges regardless of origin, replacing per-seat-priced tools like Zendesk, Intercom, and Freshdesk with a zero-cost alternative that has surpassed 2,000 GitHub stars. Built on a Go backend with a Vue.js 3 and ShadcN UI frontend, it ships as a single binary requiring only PostgreSQL and Redis. The embeddable live chat widget drops onto any website with a snippet, while the AI assistant handles initial customer queries using answers grounded in your knowledge base before escalating to human agents when needed. Agent copilot drafts replies, summarizes conversation threads, and rewrites messages for tone adjustment directly within the inbox interface. Automation rules trigger on conversation events to tag, assign, and route tickets based on configurable conditions, while auto-assignment distributes workload by agent capacity or custom criteria. SLA management tracks response and resolution time targets with breach notifications, and automated CSAT surveys measure satisfaction after conversation closure. Macros save frequently sent responses as reusable templates that simultaneously set tags and assign conversations. Role-based access control provides granular per-action permissions for teams and individual agents, and SSO supports Google, Microsoft, and any OIDC provider. The HTTP/JSON API and webhook system enable custom integrations with external tools. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Deploy
CloudBeaver screenshot thumbnail

CloudBeaver

CloudBeaver puts a full-featured database management environment in your browser, connecting to PostgreSQL, MySQL, SQL Server, Oracle, ClickHouse, and over 100 additional engines through one unified interface that requires no desktop client installation. The Java server exposes a TypeScript/React frontend through a GraphQL API where teams can browse schemas, edit data, visualize relationships, and execute queries across all connected databases in a single workspace. The SQL Editor provides syntax highlighting, auto-completion with fuzzy search, AI-assisted SQL generation from natural language prompts, script management with save/download/upload capabilities, execution plan visualization, and multi-tab result display. The Data Editor enables direct cell editing, filtering, sorting, and bulk data modification with support for spatial GIS data rendering. Database administrators access a Navigator panel for browsing schemas, tables, views, foreign tables, triggers, dependencies, and stored procedures across all connected databases. ER Diagrams visualize table relationships and schema structure, while the Visual Query Builder constructs queries without hand-writing SQL. Multi-user administration provides role-based access control, connection sharing with configurable permissions, and session management. SSH tunneling secures remote database connections, and data can be exported or imported in multiple formats. Query History tracks all executed statements with timestamps and execution statistics. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
OpenLLM screenshot thumbnail

OpenLLM

OpenLLM serves any large language model as an OpenAI-compatible API endpoint from a single CLI command, handling model download, backend selection, quantization, and port binding automatically. It supports the full spectrum of popular models including Llama 3.3, Qwen2.5, DeepSeek, Mistral, and Phi3, choosing between vLLM and PyTorch inference backends based on hardware capabilities. When vLLM is available, continuous batching with PagedAttention achieves up to 23x throughput improvement over naive serving, while GPTQ and bitsandbytes quantization reduces memory requirements for GPU-constrained deployments. The server exposes a RESTful API on port 3000 with full OpenAI client library compatibility, enabling drop-in replacement for commercial providers in any application using the standard chat completions format. A built-in web chat UI at the /chat endpoint provides immediate interactive testing without external clients. Custom model repositories allow teams to maintain private catalogs of fine-tuned models alongside the default repository that tracks the latest releases. Deployment workflows generate production-ready Docker images automatically, with Kubernetes manifest support for orchestrated scaling. Native integration with LangChain and LlamaIndex supports RAG pipelines, Transformers Agents enables tool-calling workflows, and HuggingFace Hub handles model discovery. Server-Sent Events enable real-time token streaming across all API endpoints. Backed by BentoML's production ML infrastructure. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy