Menu

Course · Streaming & orchestration

Kafka

Kafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.

Lessons
10
Interview questions
3
Projects & case studies
28
Reading time
~3 h

Your progress

Saved in this browser only

Course structure

Practise

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Kafka and Real-Time Data EngineeringReal-time data engineering with Kafka: logs, partitions and consumer groups, delivery semantics, schemas, stream processing, CDC and operating streaming pipelines.Intermediate2 min

Beginner

Core concepts you will use every day.

  1. Kafka Topics, Partitions and Consumer GroupsHow Kafka topics split into partitions, how keys decide ordering, what replication factor protects, and how consumer groups divide partitions between consumers.Beginner22 min

Intermediate

Patterns used in production pipelines.

  1. Kafka vs Queue-Based Messaging for Data PipelinesWhen to use a distributed log like Kafka and when a traditional message queue fits better: retention, replay, ordering, fan-out, per-message acknowledgement and operations.Intermediate2 min
  2. Kafka Connect and Debezium CDCHow Kafka Connect moves data in and out of Kafka: source and sink connectors, workers, converters, SMTs, offsets and dead-letter queues, plus Debezium change data capture from Postgres.Intermediate20 min
  3. Kafka Consumer Groups and RebalancingHow the group coordinator manages Kafka consumer groups, eager versus cooperative versus KIP-848 rebalancing, the cooperative sticky assignor, static membership and partition moves.Intermediate16 min
  4. Kafka Consumers: Poll Loop, Offsets, Commits and LagHow Kafka consumers read data: the poll loop, committed offsets, auto versus manual commits, consumer lag, seeking and replay, deserialisation errors and fetch sizing.Intermediate23 min
  5. Kafka Log Storage: Segments, Retention, Compaction and Topic ConfigurationHow Kafka stores partitions as segment files, how time and size retention delete data, how compaction and tombstones keep the latest value per key, and tiered storage.Intermediate22 min
  6. Kafka Producers: acks, Batching, Compression, Idempotence and OrderingHow the Kafka producer sends records: acks and durability, idempotence, batching with linger.ms, compression codecs, partitioners, retries, in-flight requests and buffering.Intermediate17 min

Advanced

Performance, internals and edge cases.

  1. Kafka Delivery Semantics: At-Most-Once, At-Least-Once and Exactly-OnceWhat at-most-once, at-least-once and exactly-once mean in Kafka, how transactions and read_committed work, and how to get exactly-once effects in external sinks.Advanced19 min
  2. Kafka Replication: Leaders, ISR, min.insync.replicas and KRaftHow Kafka replicates partitions: leaders and followers, the in-sync replica set, min.insync.replicas, unclean leader election, the high watermark, rack awareness and KRaft.Advanced20 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

System design case studies

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type