Hoarder
Hoarder (now Karakeep) is a bookmark manager that actually fights link rot: every page you save gets archived at capture time using Monolith, so the content survives even when the original URL dies. Beyond archival, an AI layer powered by OpenAI or local Ollama models auto-tags everything by analyzing page content. Prefer full privacy? Ollama keeps all inference on your server with zero external API calls. Full-text search through Meilisearch indexes the actual scraped content of every bookmark, not just titles and tags, so you find articles by what they say rather than labels you half-remember. Save links with automatic metadata extraction, plain text notes, uploaded images, and PDF documents, all organized into shareable lists with collaborative access. Browser extensions for Chrome and Firefox make saving a one-click operation from any page. Migrating is painless with importers for Chrome, Pocket, Linkwarden, Omnivore, and Tab Session Manager. LLM summarization condenses saved pages into brief overviews for quick scanning. The AI layer is entirely optional: Hoarder works perfectly as a manual bookmark manager, with intelligence adding convenience rather than imposing a requirement. SSO integration and responsive dark mode round out the package.
Mayan EDMS
Mayan EDMS stores, classifies, and retrieves millions of documents with automatic OCR, workflow automation, and audit-ready access controls that organizations have relied on for over a decade. Tesseract integration extracts searchable text from scanned PDFs and images in over 100 languages, transforming paper archives into instantly queryable digital collections without manual data entry. The workflow engine routes documents through approval chains using configurable state machines that trigger notifications, enforce retention policies, and maintain complete audit trails for regulatory compliance. Version tracking preserves every revision with full diff capabilities, while GnuPG digital signatures provide cryptographic proof of authenticity and tamper detection for sensitive records. Role-based permissions combined with object-level ACLs and LDAP integration ensure documents remain visible only to authorized users, down to individual file granularity. Full-text search powered by Whoosh or ElasticSearch handles advanced queries across massive document stores with faceted filtering and relevance ranking. The Django REST Framework API enables programmatic upload, metadata extraction, and workflow triggering from external systems. Beyond simple folder hierarchies, metadata schemas, document types, tags, and cabinet structures provide multi-dimensional classification tailored to how your organization actually works. Background processing through Celery handles OCR, conversion, and preview generation asynchronously, keeping the web interface responsive under heavy ingest loads.