Interview preparation
Apache Spark interview questions
Every question starts with a direct answer you could say out loud, then explains why, when and what goes wrong. Tick questions off as you practise.
All Apache Spark questions
| # | Question | Difficulty | Type | Done |
|---|---|---|---|---|
| 1 | What is the difference between a transformation and an action in Spark? | Easy | conceptual | |
| 2 | Explain Spark jobs, stages and tasks. | Medium | conceptual | |
| 3 | How does partition count affect Spark performance? | Medium | optimization, conceptual | |
| 4 | What causes a shuffle in Spark? | Medium | conceptual | |
| 5 | What is data skew and how can you mitigate it? | Hard | optimization, debugging |
Ticks are saved in this browser only. No account needed.
No questions match these filters
Clear a filter to see more questions.
Concept review
- LessonAdvancedAdaptive Query Execution and Dynamic Partition Pruning
- LessonBeginnerApache Spark Architecture: Driver, Executors and Cluster Managers
- LessonIntermediateSpark Caching, Checkpointing, Broadcast Variables and Accumulators
- LessonAdvancedCatalyst, Tungsten and Code Generation: How Spark SQL Optimises Queries
- LessonAdvancedDeploying and Monitoring Spark: Kubernetes and the History Server
- LessonIntermediateSpark Execution Model: Lazy Evaluation, Jobs, Stages and Tasks
- LessonAdvancedSpark Memory, Executor Sizing and Cluster Tuning
- LessonIntermediateSpark Partitions, Shuffles and Data Skew
- LessonBeginnerSpark RDD Fundamentals: Transformations, Actions and Key-Value Operations
Continue preparing
- CourseApache Spark courseOrdered lessons, projects and a cheat sheet.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
- System designCase studiesPractise full data platform designs.
- CompaniesCompany preparationGuides with attributed sources and practice questions.

