Docspell
Convert paper archives, email attachments, and PDF records into an organized digital repository with Docspell, an open-source document management system that automates optical character recognition, metadata extraction, and full-text search. Users can ingest receipts and invoices through network scanners, monitored IMAP email mailboxes, mobile uploads, or drag-and-drop browser tools. The background processing pipeline executes optical character recognition via Tesseract, enhances image contrast with unpaper, and extracts text from Word documents and PDFs. Machine learning algorithms analyze document syntax to predict correspondents, suggest relevant organization tags, and identify due dates automatically. Team members can search their entire filing cabinet using complex Boolean queries and full-text indexing powered by Apache Solr or PostgreSQL. Administrators can configure multi-user collectives with isolated permissions, define custom metadata attributes, set up webhooks, and share time-limited document download links. Users can also merge multi-page scans, track processing job queues, and export curated document collections for tax filings or legal audits. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. GNU AGPL v3.0 licensed.
Deploy