Course · Data platforms
Delta Lake
Delta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.
- Lessons
- 6
- Interview questions
- 3
- Projects & case studies
- 8
- Reading time
- ~1 h
Your progress
Saved in this browser onlyCourse structure
- 1 lessonStart hereThe complete overview of the course in one read.
- 1 lessonBeginnerCore concepts you will use every day.
- 3 lessonsIntermediatePatterns used in production pipelines.
- 1 lessonAdvancedPerformance, internals and edge cases.
Practise
- InterviewDelta Lake interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetDatabricks Cheat SheetA quick Databricks reference: Unity Catalog names and grants, Delta table operations, medallion layers, job design and the compute choices that keep costs down.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
Intermediate
Patterns used in production pipelines.
- Delta Lake: Transactions, Schema Evolution and Time TravelSee how Delta Lake's transaction log gives ACID guarantees on data lake files, how schema enforcement and evolution work, and how time travel and VACUUM interact.
- Delta Lake vs Traditional Data Lake TablesWhat changes when a folder of Parquet files becomes a Delta table: atomic writes, updates and deletes, schema enforcement, time travel and faster metadata.
- Parquet vs Avro vs ORC for Data EngineeringCompare Parquet, ORC and Avro by layout, compression, schema evolution and use case, with a measured size comparison and how column pruning works in Spark.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- AdvancedChange Data Capture PipelineReplicate an operational PostgreSQL table into a lakehouse table within minutes, including updates and deletes, so analysts query current data without touching the production database.
- AdvancedFraud Detection Data PipelineBuild the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.
- AdvancedKafka → Spark → Delta Lake Streaming PipelineAn application emits user events to Kafka. Build a streaming pipeline that lands them in Delta Lake within a minute, deduplicates replays, and produces per-minute aggregates that tolerate late events.
- IntermediateLarge-Scale Batch Processing PipelineProcess a large public dataset (several gigabytes or more) with PySpark into partitioned, query-ready tables, and document how you found and fixed the main performance bottleneck.
System design case studies
- AdvancedDesign a GDPR- and PII-Compliant Data PipelineA European e-commerce company streams customer events and database changes into a lakehouse and warehouse for analytics and ML. Design the pipeline so that personal data is identified, minimised and protected at every stage, processing respects consent and purpose, and data-subject requests (access and erasure) are completed across raw data, curated tables, streams, backups and downstream tools within the legal deadline.
- AdvancedDesign a Geospatial Analytics PipelineDesign a pipeline for a delivery company that ingests GPS pings from couriers and order locations from customers, cleans and enriches them with geographic zones (cities, delivery zones, postcodes, store catchments), and produces analytics such as delivery times by zone, demand heatmaps and route efficiency, at both daily and near-real-time granularity.
- AdvancedDesign an Idempotent Reprocessing SystemA bug in the revenue logic went unnoticed for three weeks; a source re-sent a month of corrected data; a new metric needs two years of history. Today each of these takes a week of manual work and risks duplicates. Design a reprocessing system that lets engineers re-run any pipeline over any range of history, safely and repeatably, while daily runs continue, and that propagates corrections downstream.
- AdvancedDesign a Lakehouse with Bronze, Silver and Gold LayersDesign a company-wide lakehouse in which many teams ingest batch files, database changes and event streams; data is refined through bronze, silver and gold layers; analysts query gold tables with SQL and data scientists train models from silver and gold, all on one governed copy of the data.
Resources
Cheat sheets
Related courses
- Apache SparkHow Spark works inside: driver and executors, RDDs, jobs and stages, shuffles and skew, caching, Catalyst, AQE, memory tuning, Kubernetes and the History Server.
- PySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Data modelingData modeling and warehousing: star schemas and grain, fact and dimension design, SCDs, Data Vault and other methods, dbt, semantic layers and incremental models.

