Menu

Beginner project · Project 1 of 8

CSV to Data Warehouse Pipeline

A small online shop exports orders as daily CSV files. Build a pipeline that loads them into a star schema so that sales can be reported reliably, even when files are resent or contain bad rows.

  • Beginner
  • Python · SQLite or PostgreSQL · SQL · pytest
  • 2 min read
  • Updated Oct 2026

Requirements

  • Load daily order CSV files into a local database
  • Model the data as a fact table and at least two dimensions
  • Reruns must not create duplicates
  • Bad rows are logged and counted, not silently dropped
  • Automated tests prove idempotency

Technology stack

Python, SQLite or PostgreSQL, SQL, pytest

Dataset

Generate your own CSV files with a small script, or use any public sample retail dataset whose licence allows reuse.

Business context

Reports built from hand-edited spreadsheets are slow and error-prone. A small, reliable pipeline with a clear model gives the shop consistent numbers every morning and is a good first project because every Data Engineering concept appears in miniature: modelling, idempotency, quality and testing.

Architecture

  1. Daily CSV files land in a data/ folder.
  2. A Python loader writes them to a staging table with upserts and an audit record.
  3. SQL builds dim_customer, dim_product and fact_order_line.
  4. Quality checks run; failures stop the report refresh.
A single-machine batch pipeline with staging, modelling and checks.

Start from the idempotent CSV loader tutorial, then add the modelling layer described in the star schema guide.

By DataDank Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type