2 apps Collibra
DataHub screenshot thumbnail

DataHub

DataHub maps your entire data ecosystem into a searchable, governed catalog where every table, pipeline, dashboard, and metric is discoverable and traceable from source to consumer. Originally built at LinkedIn to manage metadata at hyperscale and proven to handle 10 million+ assets and billions of relationships in production, the platform is now trusted by 3,000+ organizations including Netflix, Visa, Slack, and Pinterest. The Spring Java backend (GMS) exposes both GraphQL and OpenAPI REST endpoints, while the React frontend delivers an intuitive interface for searching, browsing, and governing data assets. The Python-based ingestion framework provides 80+ production-grade connectors extracting deep metadata from Snowflake, BigQuery, Redshift, Databricks, dbt, Airflow, Spark, Kafka, Looker, Tableau, Power BI, Superset, PostgreSQL, MySQL, Hive, Glue, S3, Iceberg, and Unity Catalog through pull-based scheduled crawls and push-based emission via Python and Java SDKs. Automatic table-level and column-level lineage detection uses SQL parsing with 97-99% accuracy, tracing data flows from ingestion pipelines through warehouses to BI dashboards. Real-time metadata streaming via Kafka keeps the catalog continuously synchronized as schemas evolve and pipelines execute. The governance layer provides business glossary management, tag propagation along lineage graphs, domain-based organization, and fine-grained access control policies. DataHub Actions triggers automated responses to metadata changes, enabling notifications, quality checks, and downstream workflows. Elasticsearch powers full-text search with faceted filtering across entities. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
OpenMetadata screenshot thumbnail

OpenMetadata

OpenMetadata builds a unified knowledge graph connecting schemas, tables, columns, dashboards, pipelines, ML models, and data products into one searchable catalog accessible at port 8585. The ingestion framework ships 130+ connectors covering Snowflake, BigQuery, Redshift, Databricks, PostgreSQL, MySQL, Kafka, Airflow, dbt, Tableau, Looker, Power BI, Metabase, and Superset, automatically extracting metadata on configurable schedules. Column-level lineage traces data flow across transformations, joins, and aggregations, while built-in data quality testing executes profiling and validation rules as data contracts with automated alerting on failures. Governance features include role-based access control, PII auto-detection, glossary term propagation, and domain-based ownership assignment. The native MCP server and AI SDK expose semantic search, lineage queries, and governance metadata as tools any LLM agent can call, enabling AI systems to discover and reason about enterprise data with full trust context. The architecture requires only PostgreSQL or MySQL plus Elasticsearch, no Kafka, no graph database, and deploys via a single Docker Compose file. Created by the founders of Apache Hadoop, Apache Atlas, and Uber's Databook, the platform has earned over 14,700 GitHub stars and adoption by 3,000+ organizations. Apache 2.0 licensed.

Deploy