Course · Distributed processing
PySpark
PySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Lessons
- 10
- Interview questions
- 7
- Projects & case studies
- 2
- Reading time
- ~3 h
Your progress
Saved in this browser onlyCourse structure
- 1 lessonStart hereThe complete overview of the course in one read.
- 4 lessonsBeginnerCore concepts you will use every day.
- 5 lessonsIntermediatePatterns used in production pipelines.
Practise
- InterviewPySpark interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
- PySpark DataFrames, Columns and SchemasBuild PySpark DataFrames with explicit StructType schemas, transform them with select, withColumn and filter, handle nulls, and cast types safely under Spark 4 ANSI mode.
- PySpark Transformations vs ActionsUnderstand lazy evaluation in PySpark: transformations build a plan, actions run it, and narrow versus wide transformations decide where Spark shuffles data.
- PySpark Aggregations, Pivot and Temporary ViewsAggregate PySpark DataFrames with groupBy and agg, build subtotals with rollup and cube, pivot and unpivot data, and share DataFrames with SQL through temp views.
- Reading and Writing Data in PySpark: Parquet, CSV, JSON, ORC, Avro and JDBCRead and write Parquet, CSV, JSON, ORC and Avro in PySpark, capture corrupt records, load JDBC tables in parallel and use partition discovery to prune folders.
Intermediate
Patterns used in production pipelines.
- PySpark Joins and Join StrategyWrite correct PySpark joins, then control how Spark executes them: broadcast hash, sort-merge and shuffle hash joins, hints, AQE, bucketing and bucketed joins.
- PySpark Window Functions: Ranking, Lag and Running TotalsUse PySpark window functions for top-N per group, deduplication, lag and lead, running and rolling totals, gap filling and sessions, and avoid the single-partition trap.
- PySpark UDFs, pandas UDFs and the Pandas API on SparkWhen to write Python UDFs in PySpark, how Arrow and pandas UDFs reduce their cost, how applyInPandas and UDTFs work, and when the pandas API on Spark fits.
- Nested and Complex Types in PySpark: Structs, Arrays, Maps and explodeWork with nested JSON in PySpark: query structs, arrays and maps, use higher-order functions, flatten with explode and posexplode, and rebuild nested output.
- Writing Efficient Output: File Sizes, Small Files and Partitioned WritesControl how many files PySpark writes and how big they are: repartition vs coalesce before a write, partitioned writes, the small files problem, fewer shuffles and custom partitioners.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- IntermediateLarge-Scale Batch Processing PipelineProcess a large public dataset (several gigabytes or more) with PySpark into partitioned, query-ready tables, and document how you found and fixed the main performance bottleneck.
- IntermediateS3 → PySpark → Snowflake Data PipelineRaw event files arrive in object storage every day. Clean and aggregate them with PySpark and load curated tables into a cloud warehouse for analysts, with each day reprocessable on demand.
Resources
Cheat sheets
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- Cheat sheetApache Spark Interview Cheat SheetThe Spark concepts interviewers ask about most, in one page: lazy evaluation, stages and shuffles, joins, partitions, skew, AQE, caching and Spark 4 defaults.
Related courses
- Apache SparkHow Spark works inside: driver and executors, RDDs, jobs and stages, shuffles and skew, caching, Catalyst, AQE, memory tuning, Kubernetes and the History Server.
- SQLSQL is the core language of data work: querying, transforming and modelling data in warehouses, lakehouses and Spark. Start here before any other tool.
- PythonPython glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.
- Delta LakeDelta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.

