AWS Data Lake ETL Pipeline — AWS architecture diagram

About This Architecture

AWS Data Lake ETL Pipeline ingests IoT events via Kinesis, Kafka topics via MSK, and batch extracts from on-premises databases, routing them through Glue Streaming and Batch ETL jobs into S3 zones (Raw, Curated, Aggregated). EMR Spark transforms curated data while Redshift serves analytics queries, with Athena enabling ad-hoc exploration and QuickSight powering dashboards. Lake Formation governs access across the data lake, CloudWatch monitors pipeline health, and Glue Data Catalog maintains metadata for all assets. This architecture demonstrates a modern medallion pattern combining real-time and batch ingestion with centralized governance and multi-query serving. Fork this diagram on Diagrams.so to customize ingestion sources, add additional transformation stages, or adjust storage zones for your organization's data maturity model.

People also ask

How do I design a scalable AWS data lake ETL pipeline that handles both real-time and batch data with proper governance?

This diagram shows a medallion-pattern data lake using Kinesis and MSK for real-time ingestion, Glue Streaming and Batch ETL for transformation into Raw/Curated/Aggregated S3 zones, EMR Spark for complex transforms, and Redshift plus Athena for analytics serving. Lake Formation enforces access control, Glue Data Catalog maintains metadata, and CloudWatch monitors pipeline health.

AWS Data Lake ETL Pipeline

AWSadvanceddata-engineeringETLdata-lakeGlueRedshift
Domain: Data EngineeringAudience: Data engineers building scalable ETL pipelines on AWS
1 views0 favoritesPublic

Created by

August 12, 2026

Updated

August 14, 2026 at 10:12 PM

Type

data pipeline

Need a custom architecture diagram?

Describe your architecture in plain English and get a production-ready Draw.io diagram in seconds. Works for AWS, Azure, GCP, Kubernetes, and more.

Generate with AI

AI-generated. Verify before production use. Learn more

Report this diagram