Course · Distributed processing
Apache Spark
How Spark works inside: driver and executors, RDDs, jobs and stages, shuffles and skew, caching, Catalyst, AQE, memory tuning, Kubernetes and the History Server.
- Lessons
- 9
- Interview questions
- 9
- Projects & case studies
- 25
- Reading time
- ~3 h
Your progress
Saved in this browser onlyCourse structure
- 1 lessonStart hereThe complete overview of the course in one read.
- 1 lessonBeginnerCore concepts you will use every day.
- 3 lessonsIntermediatePatterns used in production pipelines.
- 4 lessonsAdvancedPerformance, internals and edge cases.
Practise
- InterviewApache Spark interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
Intermediate
Patterns used in production pipelines.
- Spark Execution Model: Lazy Evaluation, Jobs, Stages and TasksHow Spark turns lazy DataFrame code into a DAG of jobs, stages and tasks, why shuffles split stages, and what to look for in each tab of the Spark UI.
- Spark Partitions, Shuffles and Data SkewHow Spark partitions data, what a shuffle does on disk and over the network, how to size shuffle partitions, and how to detect and fix skew with AQE and salting.
- Spark Caching, Checkpointing, Broadcast Variables and AccumulatorsWhen to cache or persist in Spark and at which storage level, how checkpointing cuts lineage, and how broadcast variables and accumulators share data safely.
Advanced
Performance, internals and edge cases.
- Catalyst, Tungsten and Code Generation: How Spark SQL Optimises QueriesHow Spark SQL optimises a query: Catalyst's plan phases, Tungsten memory and whole-stage code generation, cost-based optimisation, pushdown, pruning and UDF costs.
- Adaptive Query Execution and Dynamic Partition PruningHow Spark re-plans queries at run time with Adaptive Query Execution, how to read AQE plans, and how dynamic partition pruning skips fact-table partitions.
- Spark Memory, Executor Sizing and Cluster TuningHow Spark's unified memory works, how to size executors with a worked example, and how to tune GC, diagnose spill, and use dynamic allocation and speculation.
- Deploying and Monitoring Spark: Kubernetes and the History ServerRun Spark on Kubernetes: images, service accounts, pod resources and dependencies, then keep every job inspectable with event logs and the History Server.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- AdvancedChange Data Capture PipelineReplicate an operational PostgreSQL table into a lakehouse table within minutes, including updates and deletes, so analysts query current data without touching the production database.
- AdvancedFraud Detection Data PipelineBuild the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.
- AdvancedKafka → Spark → Delta Lake Streaming PipelineAn application emits user events to Kafka. Build a streaming pipeline that lands them in Delta Lake within a minute, deduplicates replays, and produces per-minute aggregates that tolerate late events.
- IntermediateLarge-Scale Batch Processing PipelineProcess a large public dataset (several gigabytes or more) with PySpark into partitioned, query-ready tables, and document how you found and fixed the main performance bottleneck.
- AdvancedReal-Time Analytics PipelineBuild a pipeline that turns a stream of order events into per-minute revenue and order counts by category, visible on a dashboard within a minute, and correct even when events arrive late.
System design case studies
- AdvancedDesign an A/B Testing Data PipelineDesign the data pipeline behind a company's experimentation platform: record which users saw which variant, join that to behavioural and business events, and produce daily, statistically sound results for hundreds of concurrent experiments.
- AdvancedDesign an Ad-Bidding Analytics PipelineDesign the analytics pipeline for a demand-side platform that bids in real-time ad auctions: log bid requests, bids, wins, impressions, clicks and conversions; give the bidder near-real-time spend for budget pacing; give advertisers campaign reports; and produce billing-grade spend and training data for bid models.
- AdvancedDesign a Churn Prediction Data PipelineDesign the data pipeline behind a churn prediction model for a subscription business: define churn labels, build leak-free features from product usage, billing and support data, produce reproducible training datasets, score every active customer daily, and deliver the scores to the customer-success team's tools.
- AdvancedDesign a Clickstream Analytics PipelineDesign a pipeline that collects every page view, click and app interaction from a website and mobile apps, and turns it into reliable product analytics: real-time traffic monitoring, sessions, funnels, retention and attribution, while respecting consent and keeping cost under control.
- AdvancedDesign a Customer 360 PlatformA retailer holds customer data in a CRM, an e-commerce platform, a support desk, a loyalty app, marketing tools and web analytics, each with its own ids. Design a Customer 360 platform that resolves these into one customer, builds a trusted profile with history and consent, serves it to analysts and to real-time applications, and respects privacy law.
- AdvancedDesign a Feature StoreTwenty ML teams each build their own feature pipelines, compute the same customer features differently, and regularly ship models whose online features do not match what they were trained on. Design a shared feature store that lets teams define features once, generate point-in-time-correct training data, and serve the same features online with low latency.
- AdvancedDesign a Real-Time Fraud Detection PipelineA payments company must decide whether to approve, review or decline each card payment while the customer waits, using the payment details, the customer's recent behaviour and machine-learning models, and must keep learning as fraud patterns change and chargeback labels arrive weeks later.
- AdvancedDesign a Geospatial Analytics PipelineDesign a pipeline for a delivery company that ingests GPS pings from couriers and order locations from customers, cleans and enriches them with geographic zones (cities, delivery zones, postcodes, store catchments), and produces analytics such as delivery times by zone, demand heatmaps and route efficiency, at both daily and near-real-time granularity.
- AdvancedDesign an IoT Sensor Data PipelineAn industrial company has 500,000 sensors on pumps, compressors and cooling systems across 2,000 sites, many on unreliable cellular links. Design a pipeline that ingests their telemetry securely, raises alerts on dangerous conditions within seconds, stores readings for dashboards and long-term analysis, and supports predictive-maintenance models, despite late, duplicated and out-of-order data.
- AdvancedDesign a Real-Time Streaming PlatformDesign a shared real-time streaming platform where hundreds of services publish domain events, platform users build stream-processing jobs on them, and the results reach the lakehouse, search, caches and alerting within seconds, reliably and with clear ownership.
- AdvancedDesign a Near-Real-Time Dashboard BackendDesign the backend for live business dashboards (orders per minute, revenue, conversion and delivery times by region and category) that show data no more than a minute old, stay fast for hundreds of concurrent viewers, and reconcile with the daily finance numbers.
- AdvancedDesign a Recommendation Data PipelineAn online marketplace wants personalised product recommendations on the home page, product pages and in emails. Design the data pipelines that collect user interactions, build training data and features, produce candidate and ranked recommendations, serve them with low latency, and measure whether they work.
- AdvancedDesign a Ride-Hailing Surge Pricing PipelineDesign the data pipeline that computes a price multiplier for each small area of a city every few seconds from live supply (available drivers) and demand (ride requests and app opens), serves it to the pricing service with low latency, and keeps a complete record of every multiplier for audit, analysis and model training.
- IntermediateDesign a Batch Ingestion FrameworkA data team writes a new pipeline by hand for every source, and now runs 150 slightly different jobs pulling from databases, SFTP drops, object storage and REST APIs. Design a reusable, metadata-driven batch ingestion framework that onboards a new source through configuration, lands data reliably and idempotently in the lakehouse, and is easy to operate, backfill and monitor.
- AdvancedDesign a Lakehouse with Bronze, Silver and Gold LayersDesign a company-wide lakehouse in which many teams ingest batch files, database changes and event streams; data is refined through bronze, silver and gold layers; analysts query gold tables with SQL and data scientists train models from silver and gold, all on one governed copy of the data.
- AdvancedDesign a Social Media Feed Analytics SystemDesign the analytics system for a social app's feed: collect impressions and engagements (likes, comments, shares, watch time) on posts, and give creators near-real-time post statistics, give product teams daily engagement and ranking-quality metrics, and give the ranking team clean training data.
- AdvancedDesign a Streaming ETL Pipeline with Kafka and SparkDesign a streaming ETL pipeline that reads application events from Kafka, cleans, enriches and deduplicates them with Spark Structured Streaming, and lands them in lakehouse tables that analysts can query within a few minutes, without losing or double-counting events.
- AdvancedDesign a Unified Batch and Streaming Platform: Lambda vs KappaA company runs nightly batch pipelines for accurate reporting and separate streaming jobs for real-time dashboards, and the two disagree. Design a platform that serves both fresh and accurate results from one set of business logic, can reprocess history when logic changes, and is cheaper to run and maintain.
- AdvancedDesign a Vector Embeddings PipelineDesign a platform that computes, stores and serves vector embeddings for several use cases (product search, recommendations, document retrieval) over hundreds of millions of items: keep embeddings in sync with changing source data, let teams upgrade embedding models safely, and serve low-latency nearest-neighbour queries with measured recall.
- AdvancedDesign a Video Streaming Analytics PipelineDesign the analytics pipeline for a video streaming service: collect player telemetry from apps, TVs and browsers, measure viewing (watch time, completion) and quality of experience (start-up time, rebuffering, bitrate) in near real time for operations and live events, and produce trusted daily content and royalty reporting.
Resources
Cheat sheets
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- Cheat sheetApache Spark Interview Cheat SheetThe Spark concepts interviewers ask about most, in one page: lazy evaluation, stages and shuffles, joins, partitions, skew, AQE, caching and Spark 4 defaults.
Related courses
- PySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Delta LakeDelta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.
- KafkaKafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.

