Kotaemon
Kotaemon is a document QA platform that combines advanced RAG techniques with a clean Gradio-based web interface for chatting with your documents. Built by Cinnamon, the Python backend supports any LLM provider including OpenAI, Azure OpenAI, Cohere, Groq, and local models via Ollama and llama-cpp-python, with a model management panel for configuring LLM and embedding providers from the UI. The default hybrid RAG pipeline combines full-text keyword retrieval with vector similarity search and applies re-ranking to ensure optimal result quality, while multi-modal document parsing extracts content from tables and figures alongside text. Advanced citations link every answer to specific source passages with relevance scores, viewable directly in the built-in PDF viewer with highlighted text spans. GraphRAG indexing via NanoGraphRAG, LightRAG, or Microsoft GraphRAG builds knowledge graphs from document collections for relationship-aware retrieval. Agent-based reasoning supports question decomposition for multi-hop queries using ReAct and ReWOO strategies. Multi-user authentication organizes documents into private and public collections with sharing and collaboration features. The platform supports Docker deployment in lite, full, and Ollama-bundled variants, runs on port 7860, and stores application data in a persistent volume. MCP tool integration enables external system connections for extended retrieval capabilities. On RepoCloud, deploy Kotaemon on a dedicated VPS with Docker, root SSH access, and complete control over your document AI infrastructure, all under the Apache 2.0 license.
Frappe Helpdesk
With over 3,200 GitHub stars, 900 forks, and backing from the team behind ERPNext, Frappe Helpdesk delivers a modern, streamlined alternative to Zendesk and Freshdesk with unlimited agents, no per-seat pricing, and full source code access under the AGPL-3.0 license. Built on the Frappe Framework with a Python backend and Vue 3 frontend using Frappe UI, the application collects customer inquiries from email, web forms, and the customer portal into a centralized ticketing queue with complete conversation history and threaded replies. Customizable SLA rules define response and resolution timelines by ticket type or team, triggering automatic alerts and escalations when deadlines approach or are missed. Assignment rules route incoming tickets to the appropriate agents based on priority, issue type, or workload balancing, while manual reassignment and transfer between teams remains available at any time. The customer self-service portal lets users submit tickets, track status, and search a knowledge base of published help articles that reduce repetitive support requests. Agents access saved reply templates for consistent, rapid responses to common queries. Custom fields, configurable workflows, and saved views adapt the interface to match each organization's support process. Real-time updates via WebSocket push ticket changes instantly to all connected agents. The PWA-compatible interface provides mobile access without a native app. Frappe Framework compatibility spans versions 15 and 16. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
Sage Wiki
Sage Wiki turns a pile of unstructured documents into a fully interlinked, searchable wiki by running them through a five-pass LLM compiler pipeline. Inspired by Andrej Karpathy's vision of LLM-compiled knowledge bases, the pipeline processes source files through diff detection, summarization, concept extraction, image captioning, and cross-reference discovery, with parallel LLM calls and checkpoint/resume for vaults scaling to 100,000+ documents. The typed ontology graph stores entities and relations with BFS traversal, configurable relation types, multilingual synonyms, and a promotion/demotion lifecycle backed by grounding verification and consensus scoring. Multi-format ingestion handles Markdown, PDF, Word, Excel, PowerPoint, EPUB, email, CSV, images, and code files without manual tagging. LLM provider support spans Anthropic, OpenAI, Gemini, Ollama, and any OpenAI-compatible API, with per-pass model routing enabling cost optimization by assigning cheaper models to simpler tasks. The built-in MCP server exposes 17 tools over SSE transport for integration with Claude, Cursor, and any MCP-compatible agent, while native Obsidian vault overlay ensures existing note workflows remain undisrupted. Team deployment supports Git-synced shared wikis, centralized server access, and hub federation across multiple projects. Ships as a single Go binary with Docker Compose multi-arch images serving the web UI on port 3333. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
LeafWiki
Engineering teams and technical writers maintain operational runbooks, system documentation, and knowledge bases using LeafWiki, an open-source self-hosted wiki that stores page contents directly as human-readable Markdown files on the local filesystem. Authors can compose documentation through a dual-pane editor equipped with live HTML preview, keyboard navigation, autocomplete for internal links, and native rendering for Mermaid diagrams, KaTeX mathematical formulas, and collapsible callout containers. The navigational tree organizes complex documentation into explicit directory hierarchies and custom drag-and-drop page sequences tracked in lightweight configuration files rather than arbitrary alphabetical lists. An integrated SQLite engine indexes page content and custom taxonomy tags for fast full-text searches, automatic incoming backlink tracking, and automated broken link detection. System operators can deploy the single-binary application without external database servers or runtime dependencies, backing up the entire repository by copying the data directory. Administrators can provision granular user roles, configure reverse proxy header authentication for single sign-on gateways, and manage read-only API access keys for automated documentation scripts and continuous integration pipelines. Public viewers can browse technical notes without authentication or switch between dark and light high-contrast themes. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
NoteDiscovery
Transforming fragmented research notes, daily technical logs, and project documentation into an interconnected knowledge base is what NoteDiscovery delivers for privacy-conscious teams and researchers. Writers compose rich documents using a dual-pane editor that renders MathJax equations, interactive task lists, and dynamic Mermaid sequence diagrams side by side with raw text. The interactive knowledge graph maps semantic relationships across notebooks, allowing researchers to explore backlinks, uncover hidden topical connections, and navigate complex idea webs visually. Team members sketch architecture diagrams and wireframes directly inside note canvases using the integrated drawing tool, saving revisions as embedded image layers without external editors. The platform organizes thoughts through flexible tag hierarchies, nested folder trees, and reusable document templates equipped with automatic date and variable substitutions. Autonomous coding assistants and language models query documentation, create structured meeting summaries, and update project indexes directly through the native Model Context Protocol server. Built-in export tools convert notebook collections into standalone HTML packages, printable documents, or portable Markdown archives for offline archiving. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Dialoqbase
Retrieval-augmented chatbots on your own knowledge base - that is the whole mission of Dialoqbase, an open-source bot-building platform. Feed it content through a broad set of data loaders - web pages and full crawls, sitemaps, PDFs, DOCX, CSV, plain text, GitHub repositories, YouTube videos, and MP3/MP4 audio - and it handles the whole RAG pipeline in one self-contained app: chunking, embedding, vector storage, and LLM querying. The distinguishing architecture choice is PostgreSQL with pgvector for embedding storage and similarity search, which removes the separate vector-database dependency, and Redis-backed Bull queues for ingesting large documents without blocking the API. Model choice is wide open: OpenAI, Anthropic Claude, Google Gemini, Cohere, Fireworks, Hugging Face, local models via Ollama, and any OpenAI-compatible endpoint, with an equally broad list of embedding providers. Finished bots embed on any website with customizable styling or deploy to Telegram, Discord, and WhatsApp, and an API creates and manages bots programmatically. Multi-user support adds registration limits and per-user bot quotas. MIT-licensed and free for commercial use.
CodeX Docs
Writing docs should feel like editing a modern document, not wrangling Markdown files - CodeX Docs delivers that on Editor.js, the block-styled editor its CodeX team builds and thousands of products use. Content is composed from clean blocks (headings, lists, code, images, embeds) with a UI that reads well on both desktop and mobile, and pages render statically with human-readable, SEO-friendly URLs. Structure is free-form: pages nest to any depth, so a flat FAQ and a deep product manual coexist in one instance, and the UI tunes to fit - collapse sections, hide the sidebar. The operational footprint is deliberately tiny. No database is required: the default driver persists to a local folder, with MongoDB available when you want it, and the whole app configures through one YAML file (overridable with APP_CONFIG_ environment variables) covering title, start page, auth password, and JWT secret. Editing mode sits behind password authentication. Thoughtful extras are wired in: readers can report misprints straight to your Telegram or Slack, Hawk error tracking catches frontend and backend exceptions, and Yandex Metrica analytics is a one-line config. A ready-made Helm chart covers Kubernetes. Written in TypeScript.
PenX
PenX delivers an open-source structured note-taking application that functions as a personal database disguised as an elegant editor — combining the outline workflow of Workflowy and Roam Research with the structured data capabilities of Tana through MetaTags that transform every note into a queryable database record. The local-first architecture stores all data on-device using PGLite, an in-process PostgreSQL-compatible engine, ensuring data ownership regardless of cloud connectivity. End-to-end encryption protects all synchronized data so that even the sync server cannot read your notes, tasks, ideas, or documents. GitHub-based version control provides out-of-the-box backup and history with full commit-level recovery. MetaTags are the core innovation — attaching structured tags to any note converts it into a database entry with typed fields, enabling table views, filters, and queries across your knowledge base without imposing rigid folder hierarchies. The daily notes workflow encourages free-form capture while MetaTags handle organization automatically, letting you record thoughts without deciding physical location upfront. AI-driven features assist with content generation, summarization, and intelligent search across your personal data hub. Real-time sync keeps web, desktop, and mobile in perfect alignment. Cross-platform availability includes web, desktop for Windows, macOS, and Linux, iOS, and Chrome extension. Deploy the web service via Next.js with pnpm using tRPC and Prisma. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
Nanote
100% portability is Nanote's one non-negotiable principle as a self-hosted note-taking app. There is no database - notebooks are plain folders and notes are plain Markdown files on your filesystem, so the same notes remain fully manageable from a terminal, Notepad, or any other editor, and walking away from Nanote costs nothing because your data was never in a proprietary format to begin with. Built with Nuxt and TypeScript around the Milkdown editor, it layers modern conveniences on that plain-file foundation: fast content search across all notes using OS-optimized tooling (ugrep), native Markdown rendering, image and file attachments, and a mobile-friendly layout for reading and editing on a phone. Clever remark directives make plain text interactive - typing ::file inserts an inline upload picker, while ::today, ::now, and ::tomorrow expand to live dates and times. A fully typed REST API with validation covers automation, and access is protected by a configurable secret key. Deployment is one container with three env vars: paths for notes, uploads, and config, all bind-mountable so your Markdown lives wherever you want it - including inside an existing sync setup. AGPL-licensed and actively daily-driven by its author.
Silicon Notes
"Somewhat lightweight, low-friction" is how Silicon Notes' author describes the personal knowledge base - written after DokuWiki's editor "drove me mad" and no existing wiki quite fit. The philosophy is that small frequent annoyances compound into cognitive load with no return, so everything here is optimized for frictionless daily use. Notes are written in plaintext Markdown and rendered as clean HTML with Pygments syntax highlighting for code blocks; pages get bi-directional relationships (backlinks), so the knowledge base becomes a connected web rather than a folder tree; and full-text plus title search retrieves anything fast. A table of contents lives in the left sidebar - "where it belongs" - editable while you read without scrolling away. Page history tracks revisions for auditing and rollback, JSON export/import keeps everything portable, and the mobile layout is genuinely usable. The stack is deliberately minimal: Python and Flask with Mistune for Markdown and SQLite for storage - no big frameworks, just a few small dependencies. One honest caveat: there is no built-in authentication, so deploy it behind a VPN, private network, or reverse-proxy auth layer. For a solo engineer's brain, it is exactly enough.