SurfSense
Positioned as the open-source NotebookLM alternative for AI agents, SurfSense delivers a live web research platform where your agents access structured data from Reddit, YouTube, Instagram, TikTok, Amazon, Walmart, Google Maps, Google Search, Indeed, and any page on the open web through one REST API or MCP server. Scheduled and event-triggered agents transform findings into briefs, alerts, podcasts, and presentations, while a built-in knowledge base keeps every discovery searchable with Perplexity-style cited answers using hybrid semantic and full-text search powered by PostgreSQL with pgvector. Upload PDFs, Office documents, images, and audio files, or sync Google Drive, OneDrive, and Dropbox — 50+ file formats supported with AI file sorting that auto-organizes documents by source, date, and topic. The MCP server exposes scrapers, knowledge base, and workspaces as native tools for Claude, Cursor, and any MCP-compatible agent. Cross-country proxy rotation handles Reddit, TikTok, and Google Search scraping with geo-aware sticky sessions and captcha-aware anti-bot handling. The platform features collaborative chats, multi-format document export, git-native knowledge base with Open Knowledge Format export, and a desktop quick-ask panel with global shortcut. Docker Compose deployment manages nine services including Caddy proxy, PostgreSQL, Redis, FastAPI backend, Celery workers, zero-cache real-time sync, and Next.js frontend with automatic Watchtower updates. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
Onyx
Formerly known as Danswer and now backed by over 31,000 GitHub stars with 253 releases, Onyx delivers a production-ready AI platform that turns any LLM into a context-aware enterprise assistant connected to your organization's actual knowledge. The agentic RAG pipeline combines BM-25 keyword search with prefix-aware embedding models in a hybrid index, then deploys AI agents to retrieve, verify, and synthesize answers with source citations from over 40 connected workplace tools including Google Drive, Confluence, Slack, Notion, Jira, SharePoint, GitHub, and Linear. Custom AI assistants with configurable prompts, backing knowledge sets, and document-level access control enable specialized agents for engineering, sales, support, and research workflows. The platform supports every major LLM provider — Anthropic Claude, OpenAI, Google Gemini, plus self-hosted options via Ollama, LiteLLM, and vLLM for fully air-gapped deployments. Beyond chat, Onyx provides web search with Serper, Google PSE, Brave, and SearXNG integration, an in-house web crawler, code execution, file creation, and multi-step deep research with report generation. Enterprise features include SSO via Google OAuth, OIDC, or SAML with SCIM provisioning, role-based access control, usage analytics by team and agent, query history auditing, PII removal through custom code hooks, and full whitelabeling. Deploy via Docker Compose on any infrastructure. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed (Community Edition).
ExcaliDash
ExcaliDash adds persistent storage, access control, version history, and real-time collaboration to Excalidraw, transforming ephemeral whiteboarding sessions into a managed drawing library your team can rely on. The Node.js/Express backend with React/TypeScript frontend deploys via Docker Compose on port 6767, using Prisma ORM on SQLite for all drawings, users, and metadata. WebSocket-powered collaboration lets multiple users edit the same canvas simultaneously with live cursor presence. Drawing snapshots preserve every revision with visual preview and one-click restore to any previous state. Three authentication modes handle different deployment needs: local email/password for personal use, hybrid mode mixing native credentials with OIDC for gradual enterprise adoption, and enforced OIDC-only for organizations requiring SSO through providers like Authentik or Keycloak. Scoped sharing controls determine whether drawings stay private, shared internally with team members, or accessible via external links without authentication. Collections organize drawings through drag-and-drop grouping, full-text search locates any diagram instantly, and exports use the non-proprietary .excalidraw format ensuring complete data portability. 1,350+ stars since November 2025. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
RAGFlow
RAGFlow has established itself as one of the most widely adopted open-source RAG engines available, powering production AI systems that demand traceable, hallucination-free answers from complex enterprise data. The platform processes PDF, DOCX, Excel, and PPT files through vision-based deep document understanding with layout analysis and OCR, extracting structured knowledge from tables, charts, and images that simpler parsers miss entirely. RAGFlow's hybrid retrieval pipeline combines vector search with BM25 keyword matching and multi-stage reranking across configurable document stores including Elasticsearch, InfiniFlow's Infinity engine, OpenSearch, and OceanBase. Developers connect any combination of LLM providers — OpenAI, DeepSeek, Anthropic Claude, Google Gemini, and locally-hosted models via Ollama — through a unified configuration layer. The visual agent workflow system enables multi-step reasoning chains with persistent memory, tool calling, and pre-built templates for common enterprise scenarios. RAGFlow synchronizes data from Confluence, S3, Notion, and Google Drive, and delivers answers through chat integrations with Feishu, Discord, Telegram, and Line. The Python SDK and RESTful API on port 9380 provide programmatic access to knowledge base management, document parsing, and conversational retrieval. The full stack deploys via Docker Compose with MySQL for metadata, Redis for task orchestration, and MinIO for object storage. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
TriliumNext
TriliumNext organizes notes in an infinitely deep tree where any single note can be cloned into multiple branches without duplication, building personal knowledge bases that mirror how ideas actually connect rather than forcing a single rigid folder hierarchy. Carrying forward the original Trilium project under active community stewardship with nearly 37,000 GitHub stars, the application is built on TypeScript with a CKEditor 5 WYSIWYG editor supporting rich text, tables, images, KaTeX math expressions, Mermaid diagrams, Excalidraw canvases, mind maps, spreadsheets with XLSX and CSV import/export, and code blocks with full syntax highlighting. A built-in JavaScript scripting engine runs on both frontend and backend, enabling custom widgets, automated workflows, scheduled tasks, and direct interaction with external REST services through a typed Script API. Full-text and fuzzy search with attribute-based queries locates any note instantly across databases tested at over 100,000 notes without performance degradation. The v0.104 release introduced importers for OneNote, Notion, Google Keep, Anytype, and Obsidian alongside 16 dedicated security fixes. Per-note AES encryption, OpenID Connect authentication, and TOTP two-factor protection safeguard sensitive content. The sync server keeps desktop clients, the progressive web app, and mobile devices in lockstep with zero third-party cloud dependency. Web Clipper captures content directly from browsers. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
OpenWiki
With over 15,900 GitHub stars and 40,000 weekly npm downloads in its first two months, OpenWiki from LangChain has rapidly become the standard for AI-generated codebase documentation. Built on the Deep Agents framework, it deploys a documentation agent that reads your repository's source code, tests, and configuration, then synthesizes a complete linked Markdown wiki with architecture overviews, integration guides, data-flow diagrams, and validated Mermaid visualizations. Two operating modes cover distinct workflows: code mode generates repository documentation in an openwiki/ folder with automatic AGENTS.md and CLAUDE.md integration for Codex, Claude Code, OpenCode, and Cursor, while personal mode builds a local knowledge base from nine connectors including Notion, Slack, Gmail, X/Twitter, Hacker News, LangSmith, Custom MCP, Web Search, and local git repositories. Thirteen model providers are supported out of the box — OpenAI, Anthropic, Gemini, AWS Bedrock, GitHub Copilot, OpenRouter, Nebius, Fireworks, Baseten, NVIDIA NIM, and any OpenAI-compatible endpoint like Ollama or LM Studio. Grounded Claims track every material assertion back to versioned source evidence, flagging stale propositions before they propagate. The interactive visualizer renders wiki pages as an explorable node graph with a side-by-side Markdown reader, exportable as a static site for GitHub Pages or MkDocs. Self-updating CI workflows via GitHub Actions, GitLab CI, or Bitbucket Pipelines open documentation PRs automatically when code changes. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Karakeep
Previously known as Hoarder and now holding 28,000+ GitHub stars, Karakeep is the most popular open-source bookmark-everything application — combining AI-powered automatic tagging with full-text search, page archival, and cross-platform access for digital content hoarders who refuse to let valuable links disappear. The Next.js frontend with tRPC communication delivers a responsive interface for saving links, notes, images, and PDFs, while Puppeteer crawls bookmarked pages to fetch titles, descriptions, and images automatically. LLM-based auto-tagging supports OpenAI, Anthropic, or local models via Ollama for privacy-first deployments that never send data to external services. Meilisearch powers full-text and semantic search across all stored content including OCR-extracted text from images. A rule-based automation engine triggers custom actions based on bookmark properties — automatically sorting, tagging, or archiving content matching defined conditions. Full page archival via Monolith preserves complete page snapshots against link rot, while yt-dlp integration archives videos from YouTube and other platforms. RSS feed ingestion automatically captures new articles from subscribed sources. Collaborative lists enable teams to build shared bookmark collections, with per-list permissions and real-time sync. Native iOS and Android apps, Chrome and Firefox extensions, and browser bookmark sync via Floccus ensure capture from any device. Importers migrate data from Chrome, Pocket, Linkwarden, Omnivore, and Tab Session Manager. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
SiYuan
Backed by over 45,000 GitHub stars and described as the tool that replaces Notion, Evernote, and Anki in a single Docker container, SiYuan is the privacy-first knowledge management system where every paragraph, heading, and list item is a uniquely addressable content block. The block-level architecture enables bidirectional links, transclusion embeds, and SQL query blocks that dynamically aggregate content across your entire workspace, while the knowledge graph visualization maps relationship networks between documents and blocks. Built-in databases support table views with relation and rollup columns, filter composition, sorting, and template-based calculations for structured data management alongside freeform notes. The FSRS spaced repetition engine turns any content block into a flashcard with scientifically calibrated review scheduling, eliminating the need for separate memorization tools. AI integration connects to OpenAI-compatible APIs for writing assistance, translation, summarization, and Q&A chat, with semantic search using embeddings and reranking for intelligent content retrieval. The Bazaar community marketplace delivers plugins, themes, templates, and widgets through a managed extension system with TypeScript plugin APIs. End-to-end encrypted synchronization works across S3-compatible storage, WebDAV servers, or SiYuan's own cloud service, while Tesseract OCR extracts searchable text from images and the web clipper captures pages from Chrome, Edge, and Firefox. Export targets include Markdown with assets, PDF, Word, and HTML. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.
Memos
Open the page, write a Markdown note, move on - Memos is a lightweight, self-hosted service built for quick capture. Instead of folders, notebooks, and titles, it presents a timeline: open the page, write a Markdown note, and move on. Notes support headings, code blocks with syntax highlighting, task lists, tables, and file attachments, with tags auto-extracted from #hashtags in the text. Each memo carries a visibility level, private, protected (logged-in users), or public, so one instance works as a personal log, a small team wiki, or a lightweight microblog. The backend is a single Go binary with a React frontend, around 50 MB of memory at runtime and a ~20 MB Docker image, so it fits comfortably on the smallest instance size with near-zero maintenance. SQLite is the default store, with MySQL and PostgreSQL supported for multi-user deployments needing more concurrency, and full REST and gRPC APIs - Connect RPC for browsers, gRPC-Gateway for external tools - make capture scriptable from CLIs, bots, and automation platforms. Fast full-text search spans all memos, pinned notes keep references handy, and a masonry view suits visual browsing. MIT-licensed with zero telemetry; content is stored as plain Markdown in a database you control, so notes remain readable, exportable, and free of proprietary formats.
XWiki
With over 140,000 code commits, 1,200+ GitHub stars, and continuous development since 2004 spanning more than two decades of active maintenance through version 18.6.0 released in July 2026, XWiki operates as a second-generation wiki platform that goes beyond static pages by enabling teams to build custom collaborative applications directly inside wiki pages using structured data forms and in-page scripting. The Java backend runs on Apache Tomcat with PostgreSQL, MySQL, or Oracle database storage, serving a responsive web interface with a WYSIWYG editor featuring real-time collaborative editing, link and macro editors, user mentions, inline comments, annotations, and complete version history with diff comparison. The App Within Minutes extension lets non-developers create custom data-driven applications using drag-and-drop form builders that generate filterable live tables for structured data browsing without writing code. Over 900 extensions from the built-in Extension Manager add functionality including blogs, task trackers, forums, diagram editors, and Confluence migration tools. Enterprise integration features include LDAP and Active Directory authentication, SAML and OIDC single sign-on, fine-grained per-page and per-space permissions with nested page hierarchies, and multi-wiki support for hosting multiple independent wikis from a single installation. The RESTful API provides programmatic access to pages, spaces, objects, attachments, and properties with XML and JSON representations. Office document import converts Word and Excel files directly into wiki pages while PDF export generates formatted documents from wiki content. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. LGPL-2.1 licensed.
ZenNotes
With over 2,200 GitHub stars and a philosophy that your notes should be files you own rather than rows in a database, ZenNotes is the keyboard-first Markdown editor that runs as a self-hosted web app backed by a Go server accessible from any browser on your network. Every note is a plain .md file in a vault directory you mount, with zero proprietary lock-in. Modal editing with real Vim motions, leader-key flows, and a command palette keeps your hands on the keyboard through edit, split, and preview modes. The rendering engine handles KaTeX math, Mermaid diagrams, TikZ graphics, and JSXGraph plots directly from Markdown syntax alongside wiki links and callout blocks. A first-party MCP server ships in the box with one-click integration for Claude Desktop and Cursor, letting AI assistants read and write the same Markdown files on disk without sync layers or duplicate copies. The bundled zen CLI provides note creation, search, tagging, task toggling, and piped capture with JSON output for shell scripting. Board views render plain CSV files as Kanban columns. Daily notes, quick capture, archive, and trash round out the vault workflow. The Go backend serves the browser frontend on port 7878 with token-based authentication, configurable browse roots, TLS proxy support, and file permission hardening at 0600/0700 defaults. Deploy via the multi-arch Docker image for linux/amd64 and linux/arm64. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Utopia
The first open-source substrate for enterprise knowledge engineering that learns passively and governs itself. The Rust-built backend paired with PostgreSQL and pgvector delivers a bitemporal knowledge graph where every fact carries two timelines: when it held in the real world and when the system came to believe it — enabling full audit trail replay of how understanding evolved. Document ingestion handles PDF, DOCX, PPTX, XLSX, CSV, Markdown, HTML, and plain text with legacy encoding detection, while scheduled syncing pulls from web pages, RSS feeds, GitHub, Jira, Notion, WebDAV, and S3-compatible buckets. Search fuses Tantivy full-text indexing with pgvector semantic vectors using Reciprocal Rank Fusion, streaming answers with inline citations that link directly to source passages. The built-in agent harness drives agentic RAG through conversation — searching documents, walking the knowledge graph at any historical date, and querying mounted databases via Ontology2SQL which achieves state-of-the-art results on BIRD Mini-Dev benchmarks. Five ontology packs ship inside the binary (schema.org, W3C Org, PROV-O, FOAF, IOF Core) with forward-chaining reasoning for transitivity, symmetry, inverses, and relation hierarchy. Entity resolution operates in three stages: exact name matching, embedding similarity, then model-based judgment with every merge reversible. Any OpenAI-compatible endpoint works including DeepSeek, Qwen, Ollama, and vLLM for fully air-gapped deployment. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache-2.0 licensed.
WeKnora
WeKnora turns scattered corporate documents into a searchable, reasoning-capable knowledge asset that your team can query in plain language and receive cited, sourced answers. Upload PDFs, Word files, web pages, Feishu wikis, Notion databases, Yuque docs, GitLab repositories, or RSS feeds into structured knowledge bases, and three distinct modes make the content actionable: RAG Quick Q&A retrieves relevant chunks and generates answers with source citations; the ReAct Agent autonomously orchestrates multi-step reasoning across knowledge retrieval, MCP tool calls, web search, and sandboxed code execution to produce comprehensive research reports; and Wiki Mode deploys LLM agents to distill raw documents into an interlinked markdown knowledge base with an interactive knowledge graph, revision history, and one-click rollback. Connect 20+ LLM providers including OpenAI, DeepSeek, Qwen, Claude, and local Ollama models without vendor lock-in, and choose from seven vector database backends (Qdrant, Milvus, Weaviate, and more) for embedding storage. Enterprise features include four-tier RBAC with per-resource ownership and per-workspace audit logs, AES-256-GCM credential encryption, scoped API keys, Langfuse observability tracing for every agent loop and tool call, and a runtime task-queue dashboard for worker-pool governance. Cross-session long-term memory preserves conversational context across interactions. The Agent Skills catalog lets teams install and share sandboxed scripts executed in Docker or E2B containers. A Chrome Extension captures web content directly into knowledge bases. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
Joplin
Notes on Windows, macOS, Linux, Android, iOS, and the terminal, synced through your own server: Joplin pairs its open-source clients with Joplin Server, the official self-hosted backend that replaces Dropbox, OneDrive, or Nextcloud as the synchronization target. Notes are Markdown with inline attachments (images, PDFs, audio), organized into hierarchical notebooks and sub-notebooks with cross-cutting tags, alongside to-do lists with reminders and alarms. End-to-end encryption is the headline feature: enabled in the clients, it encrypts sync payloads on-device before upload, so the server stores blobs it cannot read - genuine protection even if the host is compromised. The desktop app offers both Rich Text and Markdown editors, extended by a plugin ecosystem, custom themes, and an Extension API for writing your own scripts; a Web Clipper for Chrome and Firefox captures full pages or screenshots straight into notebooks. Joplin Server ships as a Docker image with SQLite for evaluation and PostgreSQL for production, offers a filesystem storage driver for large content, and includes multi-user support and note sharing - all free under AGPL-3.0 when self-hosted. Notes stay in an open format, so the exit path always exists.
Khoj
A self-hosted "second brain": Khoj indexes your own files and answers questions from them, parsing Markdown (whole Obsidian vaults included), org-mode, PDF, Word, plain text, Notion pages, GitHub repositories, and images described by a vision model, then embedding everything with sentence-transformers into a vector index for semantic search and RAG with cited sources. Any LLM backend works: local models like Llama, Qwen, or Mistral via Ollama, or cloud models like GPT, Claude, and Gemini. You can build custom agents, each with its own persona, scoped knowledge base, chat model, and tools such as web search and code execution. Scheduled automations run recurring research and deliver newsletters or notifications to your inbox, and research mode performs multi-hop web searches with inline citations. Access it from a browser, the Obsidian plugin, Emacs, desktop, or WhatsApp - all clients connect to the same self-hosted instance, making Khoj one of the few AI assistants Emacs users can point at decades of org files. Semantic search means recall works without exact keywords: "that paper about forecasting with transformers" surfaces the right PDF even when you cannot remember its title. Switching LLM backends never requires re-indexing your documents, and with a local model via Ollama, even inference stays on hardware you control - journals, research, and private notes are never sent anywhere. Python/FastAPI stack, AGPL-licensed, with PostgreSQL storage.
Trilium Notes
For people whose notes number in the tens of thousands, Trilium Notes is the hierarchical note-taking application built specifically for large personal knowledge bases - actively maintained as TriliumNext. Notes arrange into arbitrarily deep trees where every note is both content and container, and cloning lets a single note live in multiple places at once - bash notes belong under both Linux and Scripting, and Trilium refuses to make you choose. A WYSIWYG editor handles rich text, tables, math, and syntax-highlighted code blocks with Markdown-style shortcuts, while dedicated note types cover Excalidraw sketches, mind maps, geo maps with GPX tracks, relation maps that visualize connections between notes, and tables with typed columns. The attribute system is the power layer: labels attach queryable metadata (#year=1999, #author), relations create named links between notes, and both inherit down the tree - feeding full-text search, saved queries, and scripting. Scripting is Trilium's deepest differentiator: JavaScript code notes run on events like note changes or hourly schedules, build custom widgets, and add server-side logic, turning the knowledge base into a programmable platform. Protected notes encrypt sensitive content, note hoisting focuses on subtrees, and the self-hosted server syncs desktop clients across devices.
Open Notebook
The most feature-complete open-source alternative to Google's NotebookLM — a self-hosted research platform where you upload PDFs, videos, audio files, and web pages into organized notebooks, then chat with your content, generate multi-speaker podcasts, and run semantic search across everything without sending a single byte to Google's servers. The podcast engine supports 1-4 fully customizable speakers with backstories, personalities, and expertise profiles, generating professional audio dialogue through OpenAI, ElevenLabs, Google TTS, or completely local text-to-speech via Kokoro for maximum privacy. Content processing uses token-based chunking with RAG-powered retrieval grounded in your uploaded sources, while both full-text keyword search and semantic vector search via SurrealDB enable conceptual discovery across all notebooks. The 18+ supported AI providers include OpenAI, Anthropic, Google Gemini, Groq, Ollama, LM Studio, and more — configurable per task so you can route cheap models to summarization and powerful models to analysis. Content transformations extract insights, generate summaries, create study guides, and produce structured outputs from any source material. The MCP integration connects Open Notebook to Claude Desktop, VS Code, and other MCP clients for seamless workflow integration. A full REST API on port 5055 enables complete automation of notebook management, source upload, and podcast generation. Deploy via Docker Compose with the application container, SurrealDB v2 on RocksDB, and optional TTS containers. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.
DeepTutor
With 34,000+ GitHub stars and a v1.5 release driven by 36 merged community pull requests, DeepTutor from Hong Kong University's Data Science Lab delivers a full agent-native learning workspace that goes far beyond chatbot wrappers. Eight integrated surfaces — Chat, Deep Solve, Quiz Generation, Deep Research, Math Animator, Co-Writer, Book generation, and Mastery Practice — share a unified context so the objective follows the learner, not the tool. The platform's three-layer memory architecture (L1 working, L2 session, L3 long-term) makes personalization inspectable rather than opaque, letting users see exactly what the system remembers and why. Knowledge retrieval operates across five pluggable engines — LlamaIndex with FAISS vectors, PageIndex for page-level citations, GraphRAG for knowledge-graph traversal, LightRAG for local or server-offloaded retrieval, and linked Obsidian vaults — with document parsing via MinerU, Docling, markitdown, or PyMuPDF4LLM. Partners extend the tutoring brain to 15+ messaging platforms including Slack, Discord, Telegram, Matrix with E2EE, and Mattermost, each carrying private memory with branch, resume, and replay capabilities. Subagent integration brings Claude Code, Codex, Gemini, and Kimi directly into learning sessions. The system supports 30+ LLM providers from OpenAI and Anthropic to Ollama for fully local operation, with multi-user isolation, admin controls, and a full CLI interface. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.