Skip to content
#

incremental-loading

Here are 24 public repositories matching this topic...

A Databricks data engineering project simulating a full medallion architecture (Bronze → Silver → Gold) for e-commerce sales data, featuring incremental loading, a Kimball-style star schema, and SCD Type 1 & Type 2 implementations using PySpark and Delta Lake MERGE INTO.

  • Updated Sep 19, 2026
  • Python

Built a metadata-driven Azure lakehouse that incrementally ingests Spotify-style SQL data with ADF, processes new Parquet files with Databricks Auto Loader, and publishes SCD-managed Delta facts and dimensions.

  • Updated Aug 14, 2026
  • Python

Real-time AWS ecommerce data pipeline using MSK Serverless, MSK Connect, S3, AWS Glue, Delta Lake, SCD Type 2, and Redshift Serverless.

  • Updated Aug 18, 2026
  • Python

An end-to-end retail data ETL pipeline that extracts CSV data, transforms and validates it with Pandas, and incrementally loads it into PostgreSQL. Includes automated testing, structured logging, idempotent processing, CPU-optimized loading, and HashiCorp Vault for secure credential management.

  • Updated Aug 13, 2026
  • Python

Add this topic to your repo

To associate your repository with the incremental-loading topic, visit your repo's landing page and select "manage topics."

Learn more