Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Top DATA-ENGINEERING GitHub Repositories & Tools (2026)
Discover the most starred and trending open source tools tagged with #data-engineering.
An orchestration platform for the development, production, and observation of data assets.
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
Apache Superset is a Data Visualization and Data Exploration Platform
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Event Driven Orchestration & Scheduling Platform for Mission Critical Applications
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
Graph-Native Infrastructure for Context and Accountable AI Systems
Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.
Quilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.
Egeria core
Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.
The data-validation toolkit for enhanced dbt (data build tool) PR review
dbt adapter for SQL Server and Azure SQL
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
Workflow Engine for Kubernetes
A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts, per-model cost. Adapters: Databricks, Snowflake, BigQuery, DuckDB. Single static Rust binary. Apache 2.0.
Business intelligence as code: build fast, interactive data visualizations in SQL and markdown