About This Architecture

BKPM Investment Data Pipeline implements a three-layer medallion architecture ingesting investment data via web scraping into a Bronze raw layer, then flowing through Data Cleaning, Standardization, and Validation stages into a Silver processed layer, and finally materializing into a Gold structured dataset across Parquet Files, Database, and Object Storage. The pipeline extracts from BKPM Data Source through a Web Scraping Job, persists raw data in a Raw Data Store, and applies progressive data quality transformations before serving curated datasets. This architecture separates concerns between ingestion, transformation, and consumption, enabling data quality gates and independent scaling of each layer. Fork this diagram on Diagrams.so to customize data sources, add monitoring, or adapt the validation rules for your investment analytics use case.

People also ask

How do you build a scalable ETL pipeline for investment data using the medallion architecture pattern?

This BKPM Investment Data Pipeline demonstrates the medallion pattern: Bronze layer ingests raw data via web scraping, Silver layer applies data cleaning, standardization, and validation transformations, and Gold layer serves curated datasets across Parquet Files, Database, and Object Storage for analytics consumption.

BKPM Investment Data Pipeline

Autointermediatedata-engineeringmedallion-architectureETL-pipelinedata-qualityweb-scrapingdata-transformation
Domain: Data EngineeringAudience: Data engineers building medallion architecture ETL pipelines
1 views0 favoritesPublic

Created by

August 14, 2026

Updated

August 15, 2026 at 2:35 PM

Type

data pipeline

Need a custom architecture diagram?

Describe your architecture in plain English and get a production-ready Draw.io diagram in seconds. Works for AWS, Azure, GCP, Kubernetes, and more.

Generate with AI

AI-generated. Verify before production use. Learn more

Report this diagram