Menu

Course · Distributed processing

Apache Spark

How Spark works inside: driver and executors, RDDs, jobs and stages, shuffles and skew, caching, Catalyst, AQE, memory tuning, Kubernetes and the History Server.

Lessons
9
Interview questions
9
Projects & case studies
25
Reading time
~3 h

Your progress

Saved in this browser only

Course structure

Practise

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Apache Spark Architecture: Driver, Executors and Cluster ManagersHow a Spark application is built: the driver and executors, standalone, YARN and Kubernetes cluster managers, deploy modes, and SparkSession versus SparkContext.Beginner14 min

Beginner

Core concepts you will use every day.

  1. Spark RDD Fundamentals: Transformations, Actions and Key-Value OperationsLearn Spark's RDD layer: lineage and partitions, transformations versus actions, map versus flatMap, reduceByKey versus groupByKey, mapPartitions and closures.Beginner17 min

Intermediate

Patterns used in production pipelines.

  1. Spark Execution Model: Lazy Evaluation, Jobs, Stages and TasksHow Spark turns lazy DataFrame code into a DAG of jobs, stages and tasks, why shuffles split stages, and what to look for in each tab of the Spark UI.Intermediate15 min
  2. Spark Partitions, Shuffles and Data SkewHow Spark partitions data, what a shuffle does on disk and over the network, how to size shuffle partitions, and how to detect and fix skew with AQE and salting.Intermediate28 min
  3. Spark Caching, Checkpointing, Broadcast Variables and AccumulatorsWhen to cache or persist in Spark and at which storage level, how checkpointing cuts lineage, and how broadcast variables and accumulators share data safely.Intermediate19 min

Advanced

Performance, internals and edge cases.

  1. Catalyst, Tungsten and Code Generation: How Spark SQL Optimises QueriesHow Spark SQL optimises a query: Catalyst's plan phases, Tungsten memory and whole-stage code generation, cost-based optimisation, pushdown, pruning and UDF costs.Advanced17 min
  2. Adaptive Query Execution and Dynamic Partition PruningHow Spark re-plans queries at run time with Adaptive Query Execution, how to read AQE plans, and how dynamic partition pruning skips fact-table partitions.Advanced16 min
  3. Spark Memory, Executor Sizing and Cluster TuningHow Spark's unified memory works, how to size executors with a worked example, and how to tune GC, diagnose spill, and use dynamic allocation and speculation.Advanced20 min
  4. Deploying and Monitoring Spark: Kubernetes and the History ServerRun Spark on Kubernetes: images, service accounts, pod resources and dependencies, then keep every job inspectable with event logs and the History Server.Advanced14 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

System design case studies

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type