Course · Languages & query
Python
Python glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.
- Lessons
- 5
- Interview questions
- 4
- Projects & case studies
- 5
- Reading time
- ~1 h
Your progress
Saved in this browser onlyCourse structure
- 1 lessonStart hereThe complete overview of the course in one read.
- 2 lessonsBeginnerCore concepts you will use every day.
- 2 lessonsIntermediatePatterns used in production pipelines.
Practise
- InterviewPython interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetPython for Data Engineers Cheat SheetA quick Python reference for pipelines: files and CSV, JSON, dates, collections, generators and batching, error handling, logging and testing with pytest.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
- Python Functions, Modules and Reusable Data Pipeline CodeStructure Python pipeline code as small pure functions, clear modules and a thin entry point so it is testable, reusable and safe to run from any scheduler.
- Python Data Structures for Data Engineering InterviewsChoose between list, tuple, set, dict, deque and Counter by their operations and costs, with the patterns that come up in data engineering coding interviews.
Intermediate
Patterns used in production pipelines.
- Python Iterators, Generators and Memory-Efficient ProcessingUse Python iterators and generators to process files and API pages lazily, in batches, with flat memory use, and avoid the mistakes that silently exhaust them.
- Build an Idempotent CSV Loader in PythonBuild a small Python loader that reads CSV files, skips and counts bad rows, and loads into a database so running it twice gives the same result.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- BeginnerCSV to Data Warehouse PipelineA small online shop exports orders as daily CSV files. Build a pipeline that loads them into a star schema so that sales can be reported reliably, even when files are resent or contain bad rows.
- IntermediateE-commerce Analytics Data PlatformAn online store has orders, customers, products and web sessions in separate systems. Build an ELT platform that models them into trusted marts for revenue, retention and product performance.
- AdvancedFraud Detection Data PipelineBuild the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.
System design case studies
- AdvancedDesign an LLM RAG Data Ingestion PipelineDesign the ingestion pipeline behind an internal assistant that answers employees' questions using company documents (wiki pages, shared drives, PDFs, support tickets, policies), so that retrieved passages are relevant, current, and never shown to someone who could not open the original document.
- AdvancedDesign a Vector Embeddings PipelineDesign a platform that computes, stores and serves vector embeddings for several use cases (product search, recommendations, document retrieval) over hundreds of millions of items: keep embeddings in sync with changing source data, let teams upgrade embedding models safely, and serve low-latency nearest-neighbour queries with measured recall.
Resources
Cheat sheets
Related courses
- SQLSQL is the core language of data work: querying, transforming and modelling data in warehouses, lakehouses and Spark. Start here before any other tool.
- PySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.

