Categories
Analytics Observability Data Catalog Metadata Management Data Governance Data Lineage Data DiscoveryStars
Forks
Watchers
Developer links
DataHub
DataHub maps your entire data ecosystem into a searchable, governed catalog where every table, pipeline, dashboard, and metric is discoverable and traceable from source to consumer. Originally built at LinkedIn to manage metadata at hyperscale and proven to handle 10 million+ assets and billions of relationships in production, the platform is now trusted by 3,000+ organizations including Netflix, Visa, Slack, and Pinterest. The Spring Java backend (GMS) exposes both GraphQL and OpenAPI REST endpoints, while the React frontend delivers an intuitive interface for searching, browsing, and governing data assets. The Python-based ingestion framework provides 80+ production-grade connectors extracting deep metadata from Snowflake, BigQuery, Redshift, Databricks, dbt, Airflow, Spark, Kafka, Looker, Tableau, Power BI, Superset, PostgreSQL, MySQL, Hive, Glue, S3, Iceberg, and Unity Catalog through pull-based scheduled crawls and push-based emission via Python and Java SDKs. Automatic table-level and column-level lineage detection uses SQL parsing with 97-99% accuracy, tracing data flows from ingestion pipelines through warehouses to BI dashboards. Real-time metadata streaming via Kafka keeps the catalog continuously synchronized as schemas evolve and pipelines execute. The governance layer provides business glossary management, tag propagation along lineage graphs, domain-based organization, and fine-grained access control policies. DataHub Actions triggers automated responses to metadata changes, enabling notifications, quality checks, and downstream workflows. Elasticsearch powers full-text search with faceted filtering across entities. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
Benefits
- 80+ Production-Grade Data Connectors
- Extracts metadata from Snowflake, BigQuery, Redshift, dbt, Airflow, Looker, Tableau, Power BI, Kafka, Spark, and dozens more through pull-based scheduled crawls and push-based SDK emission.
- Automatic Column-Level Lineage
- SQL parser with 97-99% accuracy traces data flows from ingestion pipelines through warehouses to BI dashboards, detecting table-level and column-level dependencies across the entire data stack.
- Real-Time Metadata Streaming
- Kafka-powered event stream keeps the metadata graph continuously synchronized as schemas evolve, pipelines execute, and data assets change, enabling immediate propagation of governance decisions.
- Proven LinkedIn-Scale Architecture
- Originally built at LinkedIn and proven to handle 10 million+ data assets and billions of relationships in production, now trusted by 3,000+ organizations including Netflix, Visa, and Pinterest.
Features
- GraphQL & REST APIs
- Strongly-typed GraphQL API and OpenAPI REST endpoints enable programmatic access to the metadata graph, with Python and Java SDKs for building custom integrations and automations.
- Data Governance Framework
- Business glossary management, automated tag propagation along lineage graphs, domain-based organization, ownership assignment, and fine-grained access control policies for data stewardship.
- DataHub Actions Automation
- Event-driven framework triggers automated responses to metadata changes in real-time, enabling notifications, quality checks, and downstream workflow execution on asset modifications.
- AI-Native MCP Integration
- Model Context Protocol server enables AI agents to query the metadata graph, grounding LLM responses in verified data context for intelligent data discovery and governance automation.
- Elasticsearch Search Engine
- Full-text search with faceted filtering across datasets, dashboards, pipelines, and data products, supporting keyword, entity-type, platform, owner, tag, and domain-based discovery queries.