Course · Cloud
AWS
AWS for Data Engineers: S3, IAM, Lambda, Glue, Athena, Redshift, EMR, Kinesis, Step Functions, Lake Formation and how they fit into one data platform.
- Lessons
- 14
- Interview questions
- 0
- Projects & case studies
- 2
- Reading time
- ~4 h
Your progress
Saved in this browser onlyCourse structure
- 1 lessonStart hereThe complete overview of the course in one read.
- 2 lessonsBeginnerCore concepts you will use every day.
- 8 lessonsIntermediatePatterns used in production pipelines.
- 3 lessonsAdvancedPerformance, internals and edge cases.
Practise
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
- Amazon S3 for Data EngineersS3 for data lakes: buckets and keys, storage classes, lifecycle rules, versioning, partitioned layouts, events, encryption, access points and multipart upload.
- AWS IAM for Data Engineers: Roles, Policies and Cross-Account AccessIAM for data pipelines: users, groups and roles, identity and resource policies, least privilege, STS, cross-account access, boundaries and policy evaluation.
Intermediate
Patterns used in production pipelines.
- AWS Lambda for Data PipelinesUse AWS Lambda in data pipelines: triggers, S3 and Kinesis processing, concurrency, cold starts, layers, memory and timeout tuning, and error handling with DLQs.
- AWS Glue: Data Catalog, Crawlers and Spark ETL JobsAWS Glue for Data Engineers: the Data Catalog, crawlers, partitions, Spark ETL jobs, DynamicFrames, Glue Studio, job bookmarks, triggers, workflows and cost tuning.
- Amazon Athena: Serverless SQL on S3Query S3 with Amazon Athena: external tables, CTAS, partition projection, performance tuning, workgroups, federated queries and what drives the cost per query.
- Amazon Redshift for Data EngineersRedshift for Data Engineers: architecture, RA3, distribution and sort keys, COPY and UNLOAD, vacuum, WLM, concurrency scaling, Spectrum and Serverless.
- Amazon Kinesis: Data Streams, Firehose and Stream ProcessingAmazon Kinesis for Data Engineers: Data Streams, shard throughput, KCL and KPL, enhanced fan-out, Amazon Data Firehose, Managed Service for Apache Flink, and Kafka.
- AWS Step Functions for Data Pipeline OrchestrationOrchestrate AWS data pipelines with Step Functions: state machines, Standard vs Express, Map and Parallel, retries and catches, Glue and Lambda integrations.
- Monitoring Data Pipelines on AWS: CloudWatch and EventBridgeMonitor AWS data pipelines with CloudWatch metrics, alarms, custom metrics, Logs Insights and dashboards, plus EventBridge rules that react to job failures.
- AWS Data Services Landscape: Networking, Databases, Messaging, CDC and CostThe AWS services around a data platform: VPC networking, RDS and Aurora, DynamoDB, SQS and SNS, DMS for CDC, Amazon MSK, Cost Explorer and Budgets, Well-Architected.
Advanced
Performance, internals and edge cases.
- Amazon EMR: Spark Clusters, EMR on EKS and EMR ServerlessRun Spark on Amazon EMR: cluster node types, Spark tuning, instance fleets and Spot, bootstrap actions, EMR on EKS, EMR Serverless and how to keep EMR costs down.
- AWS Lake Formation: Data Lake Governance and Fine-Grained AccessGovern an S3 data lake with AWS Lake Formation: permissions over the Glue Data Catalog, column, row and cell filters, LF-Tags, blueprints and cross-account sharing.
- Modern AWS Data Services: Iceberg, Data Quality, Serverless and IaCModern AWS data services: S3 Tables and Iceberg, Glue Data Quality, DataBrew, AppFlow, Batch, OpenSearch, QuickSight, Secrets Manager, CDK and cost allocation tags.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
System design case studies
Resources
Related courses
- Apache SparkHow Spark works inside: driver and executors, RDDs, jobs and stages, shuffles and skew, caching, Catalyst, AQE, memory tuning, Kubernetes and the History Server.
- KafkaKafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.
- SnowflakeSnowflake is a cloud data warehouse that separates storage from compute. Learn virtual warehouses, micro-partitions, pruning, caching and cost control.

