8 apps Voice
LiveKit screenshot thumbnail

LiveKit

With over 20,000 GitHub stars and adoption by companies building everything from telehealth platforms to AI voice agents, LiveKit is the most widely deployed open-source real-time communication server available. The Go-based Selective Forwarding Unit handles hundreds of concurrent participants per node with adaptive bitrate streaming, simulcast layers, SVC codec support for VP9 and AV1, and end-to-end encryption. Client SDKs span JavaScript, Swift, Kotlin, Flutter, React Native, Rust, Python, Unity, and ESP32 embedded devices, while server-side APIs cover Node.js, Go, Ruby, Java, Python, Rust, PHP, and .NET. The Agents framework enables building AI-powered voice and video applications — real-time speech-to-text, LLM-driven conversations, and computer vision pipelines — running as server-side participants in any room. Egress records sessions to S3-compatible storage or streams to RTMP endpoints, while Ingress pulls external feeds from OBS via RTMP, WHIP, or SRT into LiveKit rooms. The SIP bridge connects traditional telephony to WebRTC rooms for hybrid conferencing. JWT-based authentication, webhook notifications, room-level moderation APIs, and selective subscription give operators granular control. Deploy as a single binary for development, Docker Compose for production single-node, or Kubernetes with the official Helm chart for distributed multi-region clusters using Redis for state coordination. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
Dograh screenshot thumbnail

Dograh

Build a voice AI agent that answers calls, qualifies leads, books appointments, and transfers to a human when needed, all from a drag-and-drop workflow builder in your browser. Dograh ships as a Docker Compose stack (API, web UI, Postgres, Redis, MinIO) that you self-host on any Linux server with automatic HTTPS provisioning via Let's Encrypt. The visual workflow builder lets you design multi-turn conversation flows by connecting nodes for greetings, intent classification, tool calls, and handoffs; describe your use case in plain English and the platform generates the LLM prompt and node graph for you. Connect your own speech-to-text, LLM, and text-to-speech providers (OpenAI, Anthropic, Gemini, ElevenLabs, Deepgram, local Whisper, Kokoro, or any OpenAI-compatible endpoint) or use the built-in Speech-to-Speech mode with Gemini Flash Live and GPT-Realtime-2 for sub-200ms latency. Telephony plugs in through Twilio, Vonage, Vobiz, or Cloudonix for inbound and outbound calling, with live agent transfer when the conversation needs a human. Webhook tool calls connect to Salesforce, HubSpot, Google Calendar, Cal.com, or any REST API without writing orchestration code. The ClonedVoice feature mixes real human voice recordings for high-frequency phrases with neural TTS fallback for dynamic content, cutting costs while improving caller trust. A built-in MCP server lets AI coding agents like Claude Code or Cursor design, test, and edit workflows through natural language. Deploy on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. BSD 2-Clause licensed.

Deploy
Speakr screenshot thumbnail

Speakr

Speakr transforms audio recordings into organized, searchable, AI-enhanced notes with speaker recognition that identifies who said what across your entire recording library. The Python/Flask backend with Vue.js 3 and Tailwind CSS frontend deploys via Docker on port 8899, offering multiple transcription engines through auto-detected connectors: WhisperX for local processing with speaker diarization and voice embeddings, OpenAI Whisper and GPT-4o-transcribe, Mistral Voxtral for cloud diarization, AssemblyAI for multi-hour files, and any custom ASR webservice. Speaker voice profiles use embedding comparison to recognize individuals across different recordings automatically, while custom vocabulary biases the transcriber toward domain-specific jargon. The AI layer goes well beyond transcription: customizable summaries with per-recording, per-tag, and per-folder prompt templates; event extraction surfacing action items and calendar events; per-recording chat with streaming responses; and Inquire Mode for semantic search and natural-language queries across your entire library simultaneously. Smart tags execute custom AI prompts on transcripts for automatic categorization. The REST API with Swagger documentation supports signed webhooks integrating with n8n, Zapier, and Make. Auto-export pushes to Obsidian and Logseq, auto-processing watches directories, and the installable PWA provides mobile-first, offline-capable access with share-target support. 3,600+ stars since May 2025. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Deploy
Airi screenshot thumbnail

Airi

Project AIRI is the most popular open-source AI companion platform — a self-hosted recreation of Neuro-sama that brings AI-powered virtual characters into your world across web, desktop, and mobile. The system renders Live2D, Spine, and VRM 3D character models with auto-blink, eye tracking, and lip-sync driven by real-time voice synthesis, while the xsAI abstraction layer connects to 40+ LLM providers including OpenAI GPT-4, Anthropic Claude, Google Gemini, DeepSeek, and local models via Ollama and OpenRouter. Built from day one on WebGPU, WebAudio, Web Workers, WebAssembly, and WebSocket technologies, the browser version runs entirely client-side with PWA offline support while the server runtime enables persistent memory via PostgreSQL with pgvector embeddings and DuckDB WASM for client-side storage. The Minecraft agent plays autonomously using mineflayer with pathfinding, and a Factorio integration provides cooperative gameplay. Social integrations deploy your companion as a Discord bot joining voice channels, a Telegram bot, and a Twitter/X agent posting and replying autonomously. The desktop Stage Tamagotchi app provides an always-on-screen companion for Windows and macOS, while Stage Pocket brings the experience to mobile. Voice features include client-side speech recognition via VAD, multiple TTS providers including ElevenLabs, and screen vision capabilities. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
Chatterbox TTS screenshot thumbnail

Chatterbox TTS

With 26,000 GitHub stars and consistent victories over ElevenLabs in blind evaluations, Chatterbox delivers state-of-the-art text-to-speech with zero-shot voice cloning requiring only 5 seconds of reference audio. The model family spans three architectures: Chatterbox Multilingual V3 (500M parameters, 23+ languages including Arabic, Chinese, Japanese, Korean, Hindi, French, German, Spanish, and Portuguese), Chatterbox-Turbo (350M parameters optimized for voice agents with a single-step distilled decoder achieving ~200ms time-to-first-speech), and Chatterbox-Nano (110M parameters running 3x faster than realtime on 8 CPU cores for edge deployment). Unique among open-source TTS systems, Chatterbox introduces emotion exaggeration control — adjusting intensity from monotone to dramatically expressive via a single parameter — and native paralinguistic tagging where tokens like [laugh], [cough], [chuckle], and [gasp] inject natural vocal reactions inline without post-processing. The alignment-informed inference pipeline eliminates hallucinations and repetition artifacts common in autoregressive TTS. Built-in PerTh neural watermarking embeds imperceptible forensic identifiers in generated audio for provenance tracking. Trained on 500,000 hours of cleaned speech data across all supported languages. Voice conversion scripts enable transforming existing audio into any cloned voice. Deploy via pip install with PyTorch, serve through Gradio interfaces or custom FastAPI endpoints, and expose via HTTP streaming or WebSocket for sub-200ms conversational applications. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
Haven screenshot thumbnail

Haven

Haven gives your community a private chat server with voice calls, screen sharing, and end-to-end encrypted direct messages where friends join via invite link in their browser without installing apps or creating third-party accounts. Real-time messaging supports image uploads via paste and drag-drop, emoji reactions, replies, threads, typing indicators, @mentions with autocomplete, and inline GIF search through Tenor or GIPHY. Peer-to-peer WebRTC voice chat includes per-user volume sliders, mute and deafen controls, talking indicators, and screen sharing with picture-in-picture mode. Direct messages use ECDH P-256 key exchange with AES-256-GCM symmetric encryption where private keys never leave the browser, ensuring not even the server operator can read them. Twenty-plus visual themes with stackable effects including CRT scanlines, Matrix Rain, Cyberpunk Text Scramble, Snowfall, and Campfire Embers let users personalize the experience with configurable intensity sliders. Rich Presence integration shows what members are playing or listening to via Last.fm, Steam, and Spotify. A built-in bot API supports webhooks and custom slash commands, while Discord history import preserves channels, threads, forums, reactions, pins, and avatars. The Node.js server deploys via Docker Compose or a single batch file that auto-handles dependencies, SSL certificates, and configuration. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Deploy
Nodyx screenshot thumbnail

Nodyx

Nodyx puts forum threads, real-time chat channels, peer-to-peer voice calls, collaborative whiteboards, and a drag-and-drop homepage builder behind one installation script on hardware you control. Create categories and threaded discussions that search engines index, then switch to Socket.IO-powered chat rooms where replies, pins, link previews, and @mentions update instantly. Voice channels connect participants through WebRTC mesh networking, keeping audio between peers without routing through third-party servers; a Rust-based STUN/TURN relay handles restrictive NATs so home-server deployments work without port forwarding. The NodyxCanvas layer adds CRDT-synchronized collaborative drawing with 11 tools, voice-aware cursors, and board-scoped chat, all persisted to PostgreSQL snapshots. Build your community's landing page with a visual editor that places widgets across 11 layout zones, then extend functionality by dropping third-party widget archives into the admin panel or authoring your own with the plain-JavaScript Widget SDK. Private conversations stay end-to-end encrypted with ECDH P-256 key exchange and AES-256-GCM ciphers where the private key never leaves the browser, and the OctoGuard module automates moderation with five matcher types plus configurable actions from muting to banning. For streamers, a native Streamer Hub bridges Twitch chat bidirectionally, provides a drag-and-drop soundboard, generates OBS browser-source overlays, and offers a mobile Stream Deck. Instances discover each other through a gossip protocol, forming a decentralized network without central coordination. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Deploy
Lobe Chat screenshot thumbnail

Lobe Chat

A private ChatGPT built with Next.js: Lobe Chat is the open-source AI chat interface teams self-host instead. Its main advantage is provider breadth: one interface connects to 40+ model providers, including OpenAI, Anthropic Claude, Google Gemini, Mistral, Groq, AWS Bedrock, Azure, and local models served through Ollama, so you can switch models per conversation and compare outputs. It handles multi-modal work: image recognition, image generation, text-to-speech, and speech-to-text. A plugin system based on function calling and the Model Context Protocol (MCP) adds external tools like web search and code execution. Run it in standalone mode as a single container with settings in browser storage, or in database mode with PostgreSQL and S3-compatible storage for persistent history, multi-user auth, and RAG knowledge bases built from uploaded documents with pgvector retrieval. Because tools arrive through function calling and MCP rather than a proprietary plugin format, custom internal tools can be exposed to the assistant with a standard server over STDIO or HTTP. Hundreds of pre-configured assistant roles import from the community marketplace. For teams the cost model matters: provider API keys billed per token typically undercut a ChatGPT Plus seat per person, and self-hosting keeps API keys, uploaded files, embeddings, and conversation history entirely on your own server.

Deploy