Menu

Course · Languages & query

Python

Python glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.

Lessons
5
Interview questions
4
Projects & case studies
5
Reading time
~1 h

Your progress

Saved in this browser only

Course structure

Practise

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Python for Data EngineeringThe Python a Data Engineer actually uses: structuring pipeline code, data structures, generators and batching, error handling, idempotent loads, testing and orchestration.Beginner2 min

Beginner

Core concepts you will use every day.

  1. Python Functions, Modules and Reusable Data Pipeline CodeStructure Python pipeline code as small pure functions, clear modules and a thin entry point so it is testable, reusable and safe to run from any scheduler.Beginner3 min
  2. Python Data Structures for Data Engineering InterviewsChoose between list, tuple, set, dict, deque and Counter by their operations and costs, with the patterns that come up in data engineering coding interviews.Beginner3 min

Intermediate

Patterns used in production pipelines.

  1. Python Iterators, Generators and Memory-Efficient ProcessingUse Python iterators and generators to process files and API pages lazily, in batches, with flat memory use, and avoid the mistakes that silently exhaust them.Intermediate4 min
  2. Build an Idempotent CSV Loader in PythonBuild a small Python loader that reads CSV files, skips and counts bad rows, and loads into a database so running it twice gives the same result.IntermediateTutorial6 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

System design case studies

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type