Course · Data platforms
Databricks
Databricks is a lakehouse platform built around Apache Spark and Delta Lake. Learn its workspace, compute, jobs and Unity Catalog governance model.
- Lessons
- 3
- Interview questions
- 1
- Projects & case studies
- 2
- Reading time
- ~1 h
Your progress
Saved in this browser onlyCourse structure
- 1 lessonStart hereThe complete overview of the course in one read.
- 1 lessonBeginnerCore concepts you will use every day.
- 1 lessonIntermediatePatterns used in production pipelines.
Practise
- InterviewDatabricks interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetDatabricks Cheat SheetA quick Databricks reference: Unity Catalog names and grants, Delta table operations, medallion layers, job design and the compute choices that keep costs down.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
System design case studies
- AdvancedDesign a Data Catalog and Lineage SystemAnalysts at a large company spend days finding the right table, nobody knows who owns half the datasets, and engineers cannot tell what will break if they change a column. Design a data catalog and lineage system that automatically harvests metadata from warehouses, lakehouses, pipelines and BI tools, makes data discoverable and trustworthy, shows table- and column-level lineage, and supports governance workflows such as ownership, classification and access requests.
- AdvancedDesign a Data Mesh ArchitectureA large retailer's central data team of 25 engineers is a bottleneck for 15 business domains: requests wait months, the team does not understand every domain's data, and quality problems are found far from where they start. Leadership wants to move to a data mesh. Design the architecture and operating model: how domains own and publish data products, what the shared platform provides, how governance works across domains, and how to migrate without breaking existing reporting.
Resources
Cheat sheets
Related courses
- Apache SparkHow Spark works inside: driver and executors, RDDs, jobs and stages, shuffles and skew, caching, Catalyst, AQE, memory tuning, Kubernetes and the History Server.
- PySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Delta LakeDelta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.

