Menu

Course · Distributed processing

PySpark

PySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.

Lessons
10
Interview questions
7
Projects & case studies
2
Reading time
~3 h

Your progress

Saved in this browser only

Course structure

Practise

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. PySpark Fundamentals for Data EngineersThe PySpark fundamentals in one guide: DataFrames and schemas, lazy evaluation, joins, window functions, UDF alternatives, partitions, shuffles and how to debug performance.Beginner2 min

Beginner

Core concepts you will use every day.

  1. PySpark DataFrames, Columns and SchemasBuild PySpark DataFrames with explicit StructType schemas, transform them with select, withColumn and filter, handle nulls, and cast types safely under Spark 4 ANSI mode.Beginner24 min
  2. PySpark Transformations vs ActionsUnderstand lazy evaluation in PySpark: transformations build a plan, actions run it, and narrow versus wide transformations decide where Spark shuffles data.Beginner2 min
  3. PySpark Aggregations, Pivot and Temporary ViewsAggregate PySpark DataFrames with groupBy and agg, build subtotals with rollup and cube, pivot and unpivot data, and share DataFrames with SQL through temp views.Beginner16 min
  4. Reading and Writing Data in PySpark: Parquet, CSV, JSON, ORC, Avro and JDBCRead and write Parquet, CSV, JSON, ORC and Avro in PySpark, capture corrupt records, load JDBC tables in parallel and use partition discovery to prune folders.Beginner25 min

Intermediate

Patterns used in production pipelines.

  1. PySpark Joins and Join StrategyWrite correct PySpark joins, then control how Spark executes them: broadcast hash, sort-merge and shuffle hash joins, hints, AQE, bucketing and bucketed joins.Intermediate26 min
  2. PySpark Window Functions: Ranking, Lag and Running TotalsUse PySpark window functions for top-N per group, deduplication, lag and lead, running and rolling totals, gap filling and sessions, and avoid the single-partition trap.Intermediate18 min
  3. PySpark UDFs, pandas UDFs and the Pandas API on SparkWhen to write Python UDFs in PySpark, how Arrow and pandas UDFs reduce their cost, how applyInPandas and UDTFs work, and when the pandas API on Spark fits.Intermediate18 min
  4. Nested and Complex Types in PySpark: Structs, Arrays, Maps and explodeWork with nested JSON in PySpark: query structs, arrays and maps, use higher-order functions, flatten with explode and posexplode, and rebuild nested output.Intermediate16 min
  5. Writing Efficient Output: File Sizes, Small Files and Partitioned WritesControl how many files PySpark writes and how big they are: repartition vs coalesce before a write, partitioned writes, the small files problem, fewer shuffles and custom partitioners.Intermediate21 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type