Course · Streaming & orchestration
Kafka
Kafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.
- Lessons
- 10
- Interview questions
- 3
- Projects & case studies
- 28
- Reading time
- ~3 h
Your progress
Saved in this browser onlyCourse structure
- 1 lessonStart hereThe complete overview of the course in one read.
- 1 lessonBeginnerCore concepts you will use every day.
- 6 lessonsIntermediatePatterns used in production pipelines.
- 2 lessonsAdvancedPerformance, internals and edge cases.
Practise
- InterviewKafka interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetKafka Cheat SheetA quick Kafka reference: topics, partitions, keys, replication, consumer groups, offsets, delivery semantics and the CLI commands for inspecting a cluster.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
Intermediate
Patterns used in production pipelines.
- Kafka vs Queue-Based Messaging for Data PipelinesWhen to use a distributed log like Kafka and when a traditional message queue fits better: retention, replay, ordering, fan-out, per-message acknowledgement and operations.
- Kafka Connect and Debezium CDCHow Kafka Connect moves data in and out of Kafka: source and sink connectors, workers, converters, SMTs, offsets and dead-letter queues, plus Debezium change data capture from Postgres.
- Kafka Consumer Groups and RebalancingHow the group coordinator manages Kafka consumer groups, eager versus cooperative versus KIP-848 rebalancing, the cooperative sticky assignor, static membership and partition moves.
- Kafka Consumers: Poll Loop, Offsets, Commits and LagHow Kafka consumers read data: the poll loop, committed offsets, auto versus manual commits, consumer lag, seeking and replay, deserialisation errors and fetch sizing.
- Kafka Log Storage: Segments, Retention, Compaction and Topic ConfigurationHow Kafka stores partitions as segment files, how time and size retention delete data, how compaction and tombstones keep the latest value per key, and tiered storage.
- Kafka Producers: acks, Batching, Compression, Idempotence and OrderingHow the Kafka producer sends records: acks and durability, idempotence, batching with linger.ms, compression codecs, partitioners, retries, in-flight requests and buffering.
Advanced
Performance, internals and edge cases.
- Kafka Delivery Semantics: At-Most-Once, At-Least-Once and Exactly-OnceWhat at-most-once, at-least-once and exactly-once mean in Kafka, how transactions and read_committed work, and how to get exactly-once effects in external sinks.
- Kafka Replication: Leaders, ISR, min.insync.replicas and KRaftHow Kafka replicates partitions: leaders and followers, the in-sync replica set, min.insync.replicas, unclean leader election, the high watermark, rack awareness and KRaft.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- AdvancedChange Data Capture PipelineReplicate an operational PostgreSQL table into a lakehouse table within minutes, including updates and deletes, so analysts query current data without touching the production database.
- AdvancedFraud Detection Data PipelineBuild the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.
- AdvancedKafka → Spark → Delta Lake Streaming PipelineAn application emits user events to Kafka. Build a streaming pipeline that lands them in Delta Lake within a minute, deduplicates replays, and produces per-minute aggregates that tolerate late events.
- AdvancedReal-Time Analytics PipelineBuild a pipeline that turns a stream of order events into per-minute revenue and order counts by category, visible on a dashboard within a minute, and correct even when events arrive late.
System design case studies
- AdvancedDesign an Ad-Bidding Analytics PipelineDesign the analytics pipeline for a demand-side platform that bids in real-time ad auctions: log bid requests, bids, wins, impressions, clicks and conversions; give the bidder near-real-time spend for budget pacing; give advertisers campaign reports; and produce billing-grade spend and training data for bid models.
- AdvancedDesign a CDC Pipeline from an OLTP Database to the WarehouseReplicate inserts, updates and deletes from a production PostgreSQL (or MySQL) database into the analytics warehouse within minutes, keeping both a current-state copy and a change history, without adding query load to the source or losing a single change.
- AdvancedDesign a Clickstream Analytics PipelineDesign a pipeline that collects every page view, click and app interaction from a website and mobile apps, and turns it into reliable product analytics: real-time traffic monitoring, sessions, funnels, retention and attribution, while respecting consent and keeping cost under control.
- AdvancedDesign an E-Commerce Inventory Sync SystemDesign a system that keeps product availability consistent across a retailer's website, mobile app, physical stores and third-party marketplaces, using stock data from several warehouse management systems, so that customers rarely see items that cannot be fulfilled and the business does not hide stock it could sell.
- AdvancedDesign an Event-Driven ArchitectureAn e-commerce company's services call each other synchronously, so one slow service stalls checkout and analytics depends on nightly database dumps. Design an event-driven architecture in which services publish business events reliably, other services react to them independently, and the same events feed analytics, without losing or double-applying anything.
- AdvancedDesign a Feature StoreTwenty ML teams each build their own feature pipelines, compute the same customer features differently, and regularly ship models whose online features do not match what they were trained on. Design a shared feature store that lets teams define features once, generate point-in-time-correct training data, and serve the same features online with low latency.
- AdvancedDesign a Real-Time Fraud Detection PipelineA payments company must decide whether to approve, review or decline each card payment while the customer waits, using the payment details, the customer's recent behaviour and machine-learning models, and must keep learning as fraud patterns change and chargeback labels arrive weeks later.
- AdvancedDesign a GDPR- and PII-Compliant Data PipelineA European e-commerce company streams customer events and database changes into a lakehouse and warehouse for analytics and ML. Design the pipeline so that personal data is identified, minimised and protected at every stage, processing respects consent and purpose, and data-subject requests (access and erasure) are completed across raw data, curated tables, streams, backups and downstream tools within the legal deadline.
- AdvancedDesign an IoT Sensor Data PipelineAn industrial company has 500,000 sensors on pumps, compressors and cooling systems across 2,000 sites, many on unreliable cellular links. Design a pipeline that ingests their telemetry securely, raises alerts on dangerous conditions within seconds, stores readings for dashboards and long-term analysis, and supports predictive-maintenance models, despite late, duplicated and out-of-order data.
- AdvancedDesign a Real-Time Streaming PlatformDesign a shared real-time streaming platform where hundreds of services publish domain events, platform users build stream-processing jobs on them, and the results reach the lakehouse, search, caches and alerting within seconds, reliably and with clear ownership.
- AdvancedDesign a Log Ingestion and Search PlatformA company runs 3,000 services on Kubernetes and VMs across two regions, producing about 20 TB of logs a day. Engineers need to search recent logs within seconds during incidents, security needs a year of audit logs, and the current logging bill grows faster than traffic. Design a log ingestion and search platform that collects, parses, redacts, indexes and retains logs reliably and affordably.
- IntermediateDesign a Notification and Alerting PipelineDesign a pipeline that turns business events (order shipped, price dropped, payment failed, threshold breached) into notifications delivered by push, email, SMS or chat, respecting each user's preferences, never spamming, and tracking whether each message was delivered.
- AdvancedDesign an Order-Events Processing SystemDesign the event backbone for an e-commerce order lifecycle (created, paid, packed, shipped, delivered, cancelled, refunded) so that downstream services and analytics see every state change reliably and in order, stuck orders are detected, and a correct current state and full history of each order are always available.
- AdvancedDesign a Payment Events Pipeline with Exactly-Once ProcessingDesign the pipeline that carries payment events (authorised, captured, refunded, charged back) from the payments service and payment provider webhooks to a double-entry ledger, merchant balances and analytics, so that every payment is recorded exactly once, no money is created or lost by retries, and every number can be reconciled with the provider.
- AdvancedDesign a Near-Real-Time Dashboard BackendDesign the backend for live business dashboards (orders per minute, revenue, conversion and delivery times by region and category) that show data no more than a minute old, stay fast for hundreds of concurrent viewers, and reconcile with the daily finance numbers.
- AdvancedDesign a Real-Time LeaderboardDesign a leaderboard for a mobile game with tens of millions of players: show the global top 100, each player's own rank and neighbours, friends and regional boards, and daily, weekly and all-time boards, all updated within a couple of seconds of a score change, and resilient to duplicates, cheating and component failures.
- AdvancedDesign a Recommendation Data PipelineAn online marketplace wants personalised product recommendations on the home page, product pages and in emails. Design the data pipelines that collect user interactions, build training data and features, produce candidate and ranked recommendations, serve them with low latency, and measure whether they work.
- AdvancedDesign a Ride-Hailing Surge Pricing PipelineDesign the data pipeline that computes a price multiplier for each small area of a city every few seconds from live supply (available drivers) and demand (ride requests and app opens), serves it to the pricing service with low latency, and keeps a complete record of every multiplier for audit, analysis and model training.
- AdvancedDesign a Schema Registry and Data Contract SystemDesign a schema registry and data contract system for a company where 60 product teams publish events to Kafka and expose database tables to the data platform, so that producers can evolve their data without breaking hundreds of downstream pipelines, dashboards and ML models, and breaking changes are caught before deployment rather than in production.
- AdvancedDesign a Search Indexing PipelineDesign the pipeline that keeps a product search index (for an online marketplace) in sync with the catalogue, pricing and inventory databases, so that changes are searchable within seconds and the whole index can be rebuilt with a new mapping without downtime.
- AdvancedDesign a Social Media Feed Analytics SystemDesign the analytics system for a social app's feed: collect impressions and engagements (likes, comments, shares, watch time) on posts, and give creators near-real-time post statistics, give product teams daily engagement and ranking-quality metrics, and give the ranking team clean training data.
- AdvancedDesign a Streaming ETL Pipeline with Kafka and SparkDesign a streaming ETL pipeline that reads application events from Kafka, cleans, enriches and deduplicates them with Spark Structured Streaming, and lands them in lakehouse tables that analysts can query within a few minutes, without losing or double-counting events.
- AdvancedDesign a Unified Batch and Streaming Platform: Lambda vs KappaA company runs nightly batch pipelines for accurate reporting and separate streaming jobs for real-time dashboards, and the two disagree. Design a platform that serves both fresh and accurate results from one set of business logic, can reprocess history when logic changes, and is cheaper to run and maintain.
- AdvancedDesign a Video Streaming Analytics PipelineDesign the analytics pipeline for a video streaming service: collect player telemetry from apps, TVs and browsers, measure viewing (watch time, completion) and quality of experience (start-up time, rebuffering, bitrate) in near real time for operations and live events, and produce trusted daily content and royalty reporting.
Resources
Cheat sheets
Related courses
- Apache SparkHow Spark works inside: driver and executors, RDDs, jobs and stages, shuffles and skew, caching, Catalyst, AQE, memory tuning, Kubernetes and the History Server.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.
- Data modelingData modeling and warehousing: star schemas and grain, fact and dimension design, SCDs, Data Vault and other methods, dbt, semantic layers and incremental models.

