Logo
Deploy Now

SaaS Alternative

Fireflies.ai Otter.ai Rev

Stars

3,669

Forks

302

Watchers

23

Developer links

Speakr

Speakr transforms audio recordings into organized, searchable, AI-enhanced notes with speaker recognition that identifies who said what across your entire recording library. The Python/Flask backend with Vue.js 3 and Tailwind CSS frontend deploys via Docker on port 8899, offering multiple transcription engines through auto-detected connectors: WhisperX for local processing with speaker diarization and voice embeddings, OpenAI Whisper and GPT-4o-transcribe, Mistral Voxtral for cloud diarization, AssemblyAI for multi-hour files, and any custom ASR webservice. Speaker voice profiles use embedding comparison to recognize individuals across different recordings automatically, while custom vocabulary biases the transcriber toward domain-specific jargon. The AI layer goes well beyond transcription: customizable summaries with per-recording, per-tag, and per-folder prompt templates; event extraction surfacing action items and calendar events; per-recording chat with streaming responses; and Inquire Mode for semantic search and natural-language queries across your entire library simultaneously. Smart tags execute custom AI prompts on transcripts for automatic categorization. The REST API with Swagger documentation supports signed webhooks integrating with n8n, Zapier, and Make. Auto-export pushes to Obsidian and Logseq, auto-processing watches directories, and the installable PWA provides mobile-first, offline-capable access with share-target support. 3,600+ stars since May 2025. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.

Speakr
Speakr
Speakr
Speakr
Speakr

Benefits

  • Privacy-First Self-Hosted Transcription
  • All audio processing and AI analysis runs entirely on your own infrastructure, ensuring sensitive meeting recordings and conversations never leave your network or reach third-party servers.
  • Multi-Engine Transcription Flexibility
  • Choose from WhisperX, OpenAI, Mistral Voxtral, AssemblyAI, or custom ASR endpoints with automatic connector detection and seamless switching between local and cloud processing.
  • AI-Powered Knowledge Extraction
  • Automatic summaries, event extraction, per-recording chat, and Inquire Mode semantic search transform raw transcripts into actionable knowledge across your entire recording library.
  • Speaker Recognition Across Recordings
  • Voice profile embeddings automatically identify speakers across different recordings, building a persistent speaker library with diarization labeling who said what in every transcript.

Features

  • Speaker Diarization and Voice Profiles
  • WhisperX-powered speaker identification labels who said what with voice embedding profiles that recognize individuals across all recordings automatically.
  • Inquire Mode Semantic Search
  • Natural-language queries search across your entire transcript library simultaneously using semantic embeddings and AI chat for cross-recording knowledge discovery.
  • REST API and Webhooks
  • Swagger-documented REST API with signed lifecycle webhooks integrates with n8n, Zapier, and Make for automated transcription workflow orchestration.
  • Smart Tags with AI Prompts
  • Custom AI prompt templates attached to tags transform transcripts automatically, enabling stackable categorization and content extraction per tag or folder.
  • Progressive Web App
  • Installable PWA with mobile-first design, offline capability, phone share-target, background sync, and light/dark themes across seven supported languages.