Logo
Deploy Now

Stars

14,953

Forks

2,330

Watchers

59

Developer links

OpenMetadata

OpenMetadata builds a unified knowledge graph connecting schemas, tables, columns, dashboards, pipelines, ML models, and data products into one searchable catalog accessible at port 8585. The ingestion framework ships 130+ connectors covering Snowflake, BigQuery, Redshift, Databricks, PostgreSQL, MySQL, Kafka, Airflow, dbt, Tableau, Looker, Power BI, Metabase, and Superset, automatically extracting metadata on configurable schedules. Column-level lineage traces data flow across transformations, joins, and aggregations, while built-in data quality testing executes profiling and validation rules as data contracts with automated alerting on failures. Governance features include role-based access control, PII auto-detection, glossary term propagation, and domain-based ownership assignment. The native MCP server and AI SDK expose semantic search, lineage queries, and governance metadata as tools any LLM agent can call, enabling AI systems to discover and reason about enterprise data with full trust context. The architecture requires only PostgreSQL or MySQL plus Elasticsearch, no Kafka, no graph database, and deploys via a single Docker Compose file. Created by the founders of Apache Hadoop, Apache Atlas, and Uber's Databook, the platform has earned over 14,700 GitHub stars and adoption by 3,000+ organizations. Apache 2.0 licensed.

OpenMetadata
OpenMetadata
OpenMetadata
OpenMetadata
OpenMetadata

Benefits

  • 130+ Data Source Connectors
  • Ingest metadata from Snowflake, BigQuery, Redshift, Databricks, PostgreSQL, MySQL, Kafka, Airflow, dbt, Tableau, Looker, Power BI, and dozens more on configurable schedules.
  • Built-In Data Quality Testing
  • Execute profiling and validation rules as data contracts with automated alerting, column-level statistics, freshness checks, and custom SQL test suites without external tools.
  • AI-Native MCP Server and SDK
  • Expose metadata, lineage, glossaries, and governance context as tools any LLM agent can call via the native Model Context Protocol server and Python AI SDK.
  • No Kafka or Graph Database Required
  • Deploy with only PostgreSQL or MySQL plus Elasticsearch — a deliberately simple architecture that runs via a single Docker Compose file in minutes.

Features

  • Column-Level Lineage
  • Trace data flow across transformations, joins, and aggregations at the column level, visualizing end-to-end lineage paths from source to dashboard.
  • Governance and Classification
  • Role-based access control, automatic PII classification, glossary term propagation, domain-based ownership, and tag-based data policies for compliance.
  • Semantic Data Discovery
  • Full-text and semantic search across all metadata assets with faceted filtering by service, owner, tag, tier, domain, and data quality status.
  • Data Contracts
  • Define schema expectations, freshness SLAs, and quality rules as enforceable contracts with automated validation and alerting on violations.
  • Collaboration and Activity Feeds
  • Threaded conversations on any data asset, task assignments, announcements, @mentions, and activity feeds tracking all metadata changes across the organization.