Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
-
Updated
Sep 25, 2026 - Python
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
An orchestration platform for the development, production, and observation of data assets.
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.
🧙 Build, run, and manage data pipelines for integrating and transforming data.
[SIGMOD'27] Easy Data Preparation with latest LLMs-based Operators and Pipelines.
Preswald is a WASM packager for Python-based interactive data apps: bundle full complex data workflows, particularly visualizations, into single files, runnable completely in-browser, using Pyodide, DuckDB, Pandas, and Plotly, Matplotlib, etc. Build dashboards, reports, and notebooks that run offline, load fast, and share like a document.
A system for agentic LLM-powered data processing and ETL
Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
Concurrent Python made simple
This dbt package captures metadata, artifacts, and test results so you can detect anomalies, monitor data quality, and build metadata tables. It powers Elementary OSS and feeds the wider context layer used by Elementary Cloud’s full Data & AI Control Plane.
⛔ [ARCHIVING ON 2026-11-01] One framework to develop, deploy and operate data workflows with Python and SQL.
AI agent tooling for data engineering workflows.
Work with your web service, database, and streaming schemas in a single format.
Main repo including core data model, data marts, data quality tests, and terminology sets.
The world's simplest distributed computing framework.
A System for Optimized Semantic Computation
Runtime for building and managing AI agents and Workflows. Easy to learn, fast to build, High Performance, Reliable by design, Intuitive UI, Production Ready.
Relational Workflows: where database schemas define executable data pipelines.
Cloud-native, data onboarding architecture for Google Cloud Datasets
To associate your repository with the data-pipelines topic, visit your repo's landing page and select "manage topics."