System design
Data Engineering system design
Each case study follows the same structure, so you practise a repeatable approach rather than memorising one answer.
- IntermediateDesign a Batch Ingestion FrameworkA data team writes a new pipeline by hand for every source, and now runs 150 slightly different jobs pulling from databases, SFTP drops, object storage and REST APIs. Design a reusable, metadata-driven batch ingestion framework that onboards a new source through configuration, lands data reliably and idempotently in the lakehouse, and is easy to operate, backfill and monitor.
- IntermediateDesign a Notification and Alerting PipelineDesign a pipeline that turns business events (order shipped, price dropped, payment failed, threshold breached) into notifications delivered by push, email, SMS or chat, respecting each user's preferences, never spamming, and tracking whether each message was delivered.
- IntermediateDesign an ELT Pipeline with dbtA growing company loads data from its product database, Stripe, Salesforce and web events into Snowflake with hand-written scripts and runs hundreds of untested SQL views. Design an ELT pipeline built on dbt that loads raw data reliably, transforms it into tested, documented models, deploys changes safely, and finishes the daily build before the business day starts.
- AdvancedDesign a Backfill and Late-Data Handling SystemDesign the platform capabilities that let a data team handle late-arriving data automatically in daily and hourly pipelines, and run large backfills (after bug fixes, new columns or new pipelines) across months of history and dozens of dependent tables, safely, cheaply and without disturbing production loads or consumers.
- AdvancedDesign a CDC Pipeline from an OLTP Database to the WarehouseReplicate inserts, updates and deletes from a production PostgreSQL (or MySQL) database into the analytics warehouse within minutes, keeping both a current-state copy and a change history, without adding query load to the source or losing a single change.
- AdvancedDesign a Churn Prediction Data PipelineDesign the data pipeline behind a churn prediction model for a subscription business: define churn labels, build leak-free features from product usage, billing and support data, produce reproducible training datasets, score every active customer daily, and deliver the scores to the customer-success team's tools.
- AdvancedDesign a Clickstream Analytics PipelineDesign a pipeline that collects every page view, click and app interaction from a website and mobile apps, and turns it into reliable product analytics: real-time traffic monitoring, sessions, funnels, retention and attribution, while respecting consent and keeping cost under control.
- AdvancedDesign a Cloud Data Warehouse for a SaaS ProductDesign a cloud data warehouse for a B2B SaaS company that consolidates its multi-tenant product database, product usage events, billing system and CRM, so the company can run trusted revenue and usage reporting (MRR, churn, activation) and offer usage analytics to its own customers without one customer ever seeing another's data.
- AdvancedDesign a Cost-Optimised Warehouse StrategyA company's cloud warehouse bill has tripled in a year to well over the budget, while data volume only doubled. Nobody can say which teams, pipelines or dashboards drive the spend. Design a strategy that makes cost visible and attributable, cuts waste without hurting SLAs, chooses the right pricing model, and keeps cost growing slower than usage from now on.
- AdvancedDesign a Customer 360 PlatformA retailer holds customer data in a CRM, an e-commerce platform, a support desk, a loyalty app, marketing tools and web analytics, each with its own ids. Design a Customer 360 platform that resolves these into one customer, builds a trusted profile with history and consent, serves it to analysts and to real-time applications, and respects privacy law.
- AdvancedDesign a Data Catalog and Lineage SystemAnalysts at a large company spend days finding the right table, nobody knows who owns half the datasets, and engineers cannot tell what will break if they change a column. Design a data catalog and lineage system that automatically harvests metadata from warehouses, lakehouses, pipelines and BI tools, makes data discoverable and trustworthy, shows table- and column-level lineage, and supports governance workflows such as ownership, classification and access requests.
- AdvancedDesign a Data Lake on Cloud Object StorageDesign a central data lake on cloud object storage that receives files, database extracts and event streams from dozens of sources, keeps raw data cheaply for years, and lets analysts, Spark jobs and data scientists query curated data with SQL without the lake turning into an ungoverned swamp.
- AdvancedDesign a Data Mesh ArchitectureA large retailer's central data team of 25 engineers is a bottleneck for 15 business domains: requests wait months, the team does not understand every domain's data, and quality problems are found far from where they start. Leadership wants to move to a data mesh. Design the architecture and operating model: how domains own and publish data products, what the shared platform provides, how governance works across domains, and how to migrate without breaking existing reporting.
- AdvancedDesign a Data Observability SystemA data platform runs 3,000 tables across a warehouse and a lakehouse, fed by Airflow, dbt, Spark and streaming jobs. Problems are usually found by business users hours later. Design a data observability system that monitors pipelines and data automatically, detects freshness, volume, schema and distribution anomalies, finds the likely root cause through lineage, and drives incidents to resolution against defined SLAs.
- AdvancedDesign a Data Quality FrameworkBad data keeps reaching dashboards and models: duplicated orders after a retry, a source that silently sent half its rows, negative prices after an upstream change. Design a company-wide data quality framework that lets teams declare expectations on their datasets, enforces them at the right points in batch and streaming pipelines, stops bad data from being published, and routes problems to the people who can fix them.
- AdvancedDesign a Data SLA and Freshness Monitoring SystemDesign a system that tells a company, for each of its important datasets, whether the data is fresh, complete and correct enough to use right now; alerts the right owner before consumers notice a problem; shows which downstream dashboards and models are affected; and reports SLA attainment over time.
- AdvancedDesign a Feature StoreTwenty ML teams each build their own feature pipelines, compute the same customer features differently, and regularly ship models whose online features do not match what they were trained on. Design a shared feature store that lets teams define features once, generate point-in-time-correct training data, and serve the same features online with low latency.
- AdvancedDesign a Financial Reconciliation PipelineDesign a daily batch pipeline that reconciles the company's internal payment ledger with settlement files from payment service providers (PSPs) and statements from banks, so finance can prove every transaction was received, settled and paid out, and can investigate every difference.
- AdvancedDesign a GDPR- and PII-Compliant Data PipelineA European e-commerce company streams customer events and database changes into a lakehouse and warehouse for analytics and ML. Design the pipeline so that personal data is identified, minimised and protected at every stage, processing respects consent and purpose, and data-subject requests (access and erasure) are completed across raw data, curated tables, streams, backups and downstream tools within the legal deadline.
- AdvancedDesign a Geospatial Analytics PipelineDesign a pipeline for a delivery company that ingests GPS pings from couriers and order locations from customers, cleans and enriches them with geographic zones (cities, delivery zones, postcodes, store catchments), and produces analytics such as delivery times by zone, demand heatmaps and route efficiency, at both daily and near-real-time granularity.
- AdvancedDesign a Lakehouse with Bronze, Silver and Gold LayersDesign a company-wide lakehouse in which many teams ingest batch files, database changes and event streams; data is refined through bronze, silver and gold layers; analysts query gold tables with SQL and data scientists train models from silver and gold, all on one governed copy of the data.
- AdvancedDesign a Log Ingestion and Search PlatformA company runs 3,000 services on Kubernetes and VMs across two regions, producing about 20 TB of logs a day. Engineers need to search recent logs within seconds during incidents, security needs a year of audit logs, and the current logging bill grows faster than traffic. Design a log ingestion and search platform that collects, parses, redacts, indexes and retains logs reliably and affordably.
- AdvancedDesign a Marketing Attribution PipelineDesign a pipeline that credits conversions (sign-ups, purchases) to the marketing touchpoints that preceded them, joins that to ad spend from each advertising platform, and gives the marketing team daily return-on-spend by channel and campaign.
- AdvancedDesign a Metrics and KPI PlatformLeadership tracks about 80 company KPIs (revenue, active users, conversion, retention, delivery time) in dashboards, spreadsheets, board decks and product experiments, and the numbers rarely agree. Design a metrics platform where every KPI is defined once, computed consistently at any grain and dimension, versioned, monitored for anomalies, and served to every tool through one interface.
- AdvancedDesign a Multi-Tenant Data PlatformA B2B SaaS company with 4,000 customer organisations wants to offer in-product analytics, scheduled exports and a data-sharing feature, built on a shared data platform. Design a multi-tenant data platform that ingests each tenant's data, keeps tenants strictly isolated, gives every tenant fast and fair query performance, supports enterprise tenants with stronger isolation and residency needs, and attributes cost per tenant.
- AdvancedDesign a Near-Real-Time Dashboard BackendDesign the backend for live business dashboards (orders per minute, revenue, conversion and delivery times by region and category) that show data no more than a minute old, stay fast for hundreds of concurrent viewers, and reconcile with the daily finance numbers.
- AdvancedDesign a Near-Zero Downtime Data Platform MigrationDesign the migration of a live on-premises data warehouse, the ETL jobs that load it and the dashboards that read it to a cloud warehouse or lakehouse, so that consumers see no more than a few minutes of disruption and every number can be proved to match before the old system is switched off.
- AdvancedDesign a Payment Events Pipeline with Exactly-Once ProcessingDesign the pipeline that carries payment events (authorised, captured, refunded, charged back) from the payments service and payment provider webhooks to a double-entry ledger, merchant balances and analytics, so that every payment is recorded exactly once, no money is created or lost by retries, and every number can be reconciled with the provider.
- AdvancedDesign a Real-Time Fraud Detection PipelineA payments company must decide whether to approve, review or decline each card payment while the customer waits, using the payment details, the customer's recent behaviour and machine-learning models, and must keep learning as fraud patterns change and chargeback labels arrive weeks later.
- AdvancedDesign a Real-Time LeaderboardDesign a leaderboard for a mobile game with tens of millions of players: show the global top 100, each player's own rank and neighbours, friends and regional boards, and daily, weekly and all-time boards, all updated within a couple of seconds of a score change, and resilient to duplicates, cheating and component failures.
- AdvancedDesign a Real-Time Streaming PlatformDesign a shared real-time streaming platform where hundreds of services publish domain events, platform users build stream-processing jobs on them, and the results reach the lakehouse, search, caches and alerting within seconds, reliably and with clear ownership.
- AdvancedDesign a Recommendation Data PipelineAn online marketplace wants personalised product recommendations on the home page, product pages and in emails. Design the data pipelines that collect user interactions, build training data and features, produce candidate and ranked recommendations, serve them with low latency, and measure whether they work.
- AdvancedDesign a Ride-Hailing Surge Pricing PipelineDesign the data pipeline that computes a price multiplier for each small area of a city every few seconds from live supply (available drivers) and demand (ride requests and app opens), serves it to the pricing service with low latency, and keeps a complete record of every multiplier for audit, analysis and model training.
- AdvancedDesign a Schema Registry and Data Contract SystemDesign a schema registry and data contract system for a company where 60 product teams publish events to Kafka and expose database tables to the data platform, so that producers can evolve their data without breaking hundreds of downstream pipelines, dashboards and ML models, and breaking changes are caught before deployment rather than in production.
- AdvancedDesign a Search Indexing PipelineDesign the pipeline that keeps a product search index (for an online marketplace) in sync with the catalogue, pricing and inventory databases, so that changes are searchable within seconds and the whole index can be rebuilt with a new mapping without downtime.
- AdvancedDesign a Self-Serve Analytics PlatformDesign a platform that lets 1,000 employees across product, marketing, finance and operations find trustworthy data, answer their own questions with SQL or a BI tool, and build their own dashboards and models, without a central data team writing every query, and without losing control of data quality, access to sensitive data or cost.
- AdvancedDesign a Slowly Changing Dimension FrameworkA warehouse has 60 dimensions (customers, products, stores, sales territories, employees) whose attributes change over time. Each team handles history differently, some overwrite values and lose history, others produce overlapping or duplicate versions. Design a reusable slowly changing dimension framework that applies the right history rules per attribute, loads changes idempotently from batch and CDC sources, handles late and out-of-order changes, and lets facts join to the version that was valid at the time.
- AdvancedDesign a Social Media Feed Analytics SystemDesign the analytics system for a social app's feed: collect impressions and engagements (likes, comments, shares, watch time) on posts, and give creators near-real-time post statistics, give product teams daily engagement and ranking-quality metrics, and give the ranking team clean training data.
- AdvancedDesign a Streaming ETL Pipeline with Kafka and SparkDesign a streaming ETL pipeline that reads application events from Kafka, cleans, enriches and deduplicates them with Spark Structured Streaming, and lands them in lakehouse tables that analysts can query within a few minutes, without losing or double-counting events.
- AdvancedDesign a Time-Series Metrics StoreDesign a store for operational and business metrics (request latency, error counts, CPU, orders per minute) that ingests millions of samples per second from thousands of services, answers dashboard and alerting queries in under a second, and keeps a year of history affordably.
- AdvancedDesign a Unified Batch and Streaming Platform: Lambda vs KappaA company runs nightly batch pipelines for accurate reporting and separate streaming jobs for real-time dashboards, and the two disagree. Design a platform that serves both fresh and accurate results from one set of business logic, can reprocess history when logic changes, and is cheaper to run and maintain.
- AdvancedDesign a Vector Embeddings PipelineDesign a platform that computes, stores and serves vector embeddings for several use cases (product search, recommendations, document retrieval) over hundreds of millions of items: keep embeddings in sync with changing source data, let teams upgrade embedding models safely, and serve low-latency nearest-neighbour queries with measured recall.
- AdvancedDesign a Video Streaming Analytics PipelineDesign the analytics pipeline for a video streaming service: collect player telemetry from apps, TVs and browsers, measure viewing (watch time, completion) and quality of experience (start-up time, rebuffering, bitrate) in near real time for operations and live events, and produce trusted daily content and royalty reporting.
- AdvancedDesign an A/B Testing Data PipelineDesign the data pipeline behind a company's experimentation platform: record which users saw which variant, join that to behavioural and business events, and produce daily, statistically sound results for hundreds of concurrent experiments.
- AdvancedDesign an Ad-Bidding Analytics PipelineDesign the analytics pipeline for a demand-side platform that bids in real-time ad auctions: log bid requests, bids, wins, impressions, clicks and conversions; give the bidder near-real-time spend for budget pacing; give advertisers campaign reports; and produce billing-grade spend and training data for bid models.
- AdvancedDesign an Analytics and BI PlatformDesign the analytics and BI platform for a mid-sized company where executives, finance, marketing and operations all need dashboards and self-service analysis, data comes from about 40 sources, and today different teams report different numbers for the same metric.
- AdvancedDesign an E-Commerce Inventory Sync SystemDesign a system that keeps product availability consistent across a retailer's website, mobile app, physical stores and third-party marketplaces, using stock data from several warehouse management systems, so that customers rarely see items that cannot be fulfilled and the business does not hide stock it could sell.
- AdvancedDesign an Event-Driven ArchitectureAn e-commerce company's services call each other synchronously, so one slow service stalls checkout and analytics depends on nightly database dumps. Design an event-driven architecture in which services publish business events reliably, other services react to them independently, and the same events feed analytics, without losing or double-applying anything.
- AdvancedDesign an Idempotent Reprocessing SystemA bug in the revenue logic went unnoticed for three weeks; a source re-sent a month of corrected data; a new metric needs two years of history. Today each of these takes a week of manual work and risks duplicates. Design a reprocessing system that lets engineers re-run any pipeline over any range of history, safely and repeatably, while daily runs continue, and that propagates corrections downstream.
- AdvancedDesign an IoT Sensor Data PipelineAn industrial company has 500,000 sensors on pumps, compressors and cooling systems across 2,000 sites, many on unreliable cellular links. Design a pipeline that ingests their telemetry securely, raises alerts on dangerous conditions within seconds, stores readings for dashboards and long-term analysis, and supports predictive-maintenance models, despite late, duplicated and out-of-order data.
- AdvancedDesign an LLM RAG Data Ingestion PipelineDesign the ingestion pipeline behind an internal assistant that answers employees' questions using company documents (wiki pages, shared drives, PDFs, support tickets, policies), so that retrieved passages are relevant, current, and never shown to someone who could not open the original document.
- AdvancedDesign an Order-Events Processing SystemDesign the event backbone for an e-commerce order lifecycle (created, paid, packed, shipped, delivered, cancelled, refunded) so that downstream services and analytics see every state change reliably and in order, stuck orders are detected, and a correct current state and full history of each order are always available.

