Databases · DBC-52

Modern Data Engineering and Lakehouse Architecture

Modern Data Engineering and Lakehouse Architecture is a hands-on course in building an open source lakehouse, from Bronze, Silver and Gold layers on Apache Iceberg and Trino to Airflow ingestion, dbt transformation, data contracts, time travel and lineage. It suits data engineers, DBAs and teams building an on-premise data platform, and you leave with a working reference pipeline.

Updated
From 8,010 THB / person 8,900 −10% excl. VAT 7% · group rates available
PDFDownload the course outline
  • Duration18 hours · 3 days
  • FormatOnsite / live online
  • Next roundOn request
  • CertificateIncluded

Course overview

Many organisations still keep a data warehouse for reporting apart from a data lake for raw files, and pull data from their core systems with nightly batch ETL. The result is data stored twice, numbers that do not match, slow answers to questions that need recent data, and, when an AI project or an audit arrives, no way to show where a figure came from, because lineage and data quality gates were never built in. The lakehouse, built on an open table format such as Apache Iceberg, brings both worlds together in a single layer.

This course walks learners through building a lakehouse with open source tools as one continuous project across all three days. It starts by designing Bronze, Silver and Gold layers on S3-compatible object storage with Apache Iceberg and Trino, then ingests simulated deposit transactions incrementally with Apache Airflow and covers CDC concepts. Data is transformed with dbt, guarded by data quality tests and data contracts, before learners use Iceberg time travel, schema evolution and compaction, set up observability and lineage, and close with a capstone that explains the quality control point at every layer. Every lab runs on the learner's own machine with Docker and uses synthetic data only. (3 days, 6 hours per day, 18 hours in total, Intermediate level.)

What you’ll gain

  • Explain how a data warehouse, a data lake and a lakehouse differ, and choose the right architecture for the organisation
  • Design a medallion architecture with Bronze, Silver and Gold layers and clear rules for using each layer
  • Run a lakehouse stack with Docker Compose, including object storage, an Iceberg catalog, Trino and Airflow
  • Choose between full load, incremental load and CDC, and build ingestion pipelines that are safe to re-run
  • Use dbt with Trino to transform data in Iceberg tables, with tests, SCD Type 2 and generated documentation
  • Define data contracts with model contracts and source freshness so the pipeline stops when data is wrong
  • Use time travel, schema evolution and Iceberg table maintenance such as compaction and snapshot expiry
  • Put data observability and lineage in place from the source system to the final report

Who this course is for

  • Data engineers and ETL developers moving from traditional ETL to a lakehouse architecture
  • DBAs and system analysts who look after source databases and feed data to reporting or AI teams
  • IT infrastructure teams that need to run an on-premise data platform inside the organisation
  • Data architects and tech leads who design data layers and quality controls across the organisation
  • Anyone who already knows Airflow and dbt and wants to move on to open table formats and the lakehouse

Prerequisites

  • SQL at the level of JOINs, GROUP BY and subqueries
  • Basic Linux command-line and Git skills, and the ability to read simple Python code
  • No prior Spark, Kafka or cloud experience is needed; some Airflow or dbt experience helps you move faster
  • A laptop or VM that can run Docker, ideally with 16 GB RAM or more, 4 CPU cores and 40 GB of free disk, with permission to install software

Curriculum

Course Details

This is a lakehouse architecture course that builds on the Data Pipeline with Apache Airflow and dbt course. It runs for 3 days, 6 hours per day (18 hours in total, 09:00-16:00), as lectures with labs built on one continuous project across all three days. Intermediate level. The course focuses on medallion design on S3-compatible object storage with Apache Iceberg and Trino, ingestion patterns and CDC concepts, data contracts, time travel and schema evolution, observability and lineage, and a capstone project. Airflow and

dbt are used as working tools, with their basics kept short; learners who want them in depth can take the Data Pipeline with Apache Airflow and dbt course. Every lab runs on the learner's own laptop or VM with Docker (16 GB RAM or more recommended) and uses a simulated deposit system only, so the approach can be applied inside an organisation without moving real data out, in line with PDPA principles. Learners take home the workshop guide, DAG code, the dbt project and the Docker Compose files.

Day 1 Lakehouse Architecture and Data Ingestion

Section 1: From Data Warehouse and Data Lake to Lakehouse

  • The cost of keeping a warehouse and a lake apart: duplicated data, mismatched numbers and slow nightly ETL
  • What a lakehouse is: object storage, an open table format, a catalog and a query engine
  • Separating storage from compute so several engines can read the same data
  • The three-day project at a glance: from a simulated deposit system to Gold tables for dashboards
  • Synthetic data and on-premise labs, in line with PDPA principles

Section 2: Designing the Medallion Architecture and Data Model

  • The role of the Bronze, Silver and Gold layers and the rules for what each layer may be used for
  • Consistent naming for schemas, tables and columns, and a retention policy for each layer
  • A star schema recap: facts and dimensions for transaction data
  • Slowly changing dimensions Type 2 for keeping customer and account history
  • Workshop: design the data layers and data model for the project

Section 3: Lab: Starting the Lakehouse Stack with Docker Compose

  • The stack: a PostgreSQL source, object storage, an Iceberg catalog, Trino and Airflow
  • Configure Trino catalogs for Iceberg and PostgreSQL
  • Create the bucket and the Bronze, Silver and Gold schemas
  • Query across systems with the Trino CLI or DBeaver
  • Check machine resources and fix common container problems

Section 4: Ingestion Patterns and CDC Concepts

  • How full load, watermark-based incremental load and change data capture differ
  • Choosing by data volume, required freshness and load on the source system
  • Log-based versus query-based CDC, and handling deleted and late-arriving records
  • Idempotency: designing loads that give the same result when re-run
  • Instructor demo: CDC with Debezium and Apache Kafka

Section 5: Lab: An Airflow DAG for Bronze Ingestion

  • Airflow 3 essentials for this project: DAGs, tasks and the TaskFlow API
  • Write a task that loads PostgreSQL data incrementally into a Bronze Iceberg table
  • Store the watermark and metadata columns such as load time and run ID
  • Set the schedule, retries and alerts for pipeline failures
  • Workshop: new transactions flow into the Bronze layer automatically
Day 2 Transformation with dbt, Data Quality and Data Contracts

Section 6: dbt on the Lakehouse

  • dbt essentials for this project: models, source, ref and materializations
  • Connect dbt to Trino with dbt-trino and build models as Iceberg tables
  • Organise the project folders to match the Bronze, Silver and Gold layers
  • Lab: set up the dbt project and declare sources from the Bronze layer

Section 7: Lab: From Bronze to Silver

  • Clean data, cast data types and remove duplicates
  • Incremental models with the merge strategy on Iceberg tables
  • Handle late-arriving data and cancelled or corrected transactions
  • Lab: build reliable Silver tables for transactions and accounts

Section 8: Lab: From Silver to Gold

  • Build facts and dimensions from the star schema designed earlier
  • Keep SCD Type 2 history with dbt snapshots
  • Summary tables for dashboards, such as balances and daily transactions
  • Lab: build the Gold layer and query it through Trino

Section 9: Pipeline-Level Data Quality Tests

  • Generic tests: not_null, unique, accepted_values and relationships
  • Singular tests for business rules, such as balances matching the sum of transactions
  • Severity levels warn and error, and storing failing records
  • Lab: add tests at every layer and make them catch bad data

Section 10: Lab: Data Contracts and the End-to-End Pipeline

  • What a data contract is, and what producers and consumers agree on
  • Model contracts: enforcing column names and data types, and the limits on Trino
  • Source freshness: setting freshness thresholds for source data
  • Have Airflow run dbt after ingestion and stop as soon as a contract is broken
  • Lab: change the source schema and watch the pipeline stop before bad data reaches Gold
Day 3 Apache Iceberg, Observability and Capstone

Section 11: Apache Iceberg in Depth

  • The problems an open table format solves for Parquet files sitting in a data lake
  • Inside a table: metadata files, manifests, snapshots and the catalog
  • ACID transactions and concurrent writes from several engines
  • A brief comparison of Iceberg and Delta Lake

Section 12: Lab: Time Travel for Audit and Review

  • Inspect snapshots and table history through metadata tables
  • Query data as of a point in time with FOR TIMESTAMP AS OF and FOR VERSION AS OF
  • Roll a table back to an earlier snapshot after a bad load
  • Lab: answer what the balances were at the end of the previous day

Section 13: Schema Evolution and Partitioning

  • Add, rename or widen columns without rewriting data
  • Hidden partitioning and choosing partitions to match query patterns
  • Partition evolution as data volumes change
  • Lab: add a transaction channel column and change the partitioning of the transactions table

Section 14: Lab: Table Maintenance and Performance

  • The small files problem caused by frequent loads
  • Compaction with optimize and removing old snapshots with expire_snapshots
  • Remove unreferenced files and schedule table maintenance with Airflow
  • Lab: measure query time before and after compaction

Section 15: Data Observability and Data Lineage

  • Monitor freshness, volume and schema drift at each layer
  • Show lineage from the source system to the report with dbt docs
  • Lineage expectations in financial services, for example the BCBS 239 principles
  • Lab: alert when data is stale or row counts look abnormal

Section 16: Workshop: Capstone Project

  • Run the whole pipeline from new source transactions to the Gold tables
  • Present the architecture and the data quality control point at each layer
  • Demonstrate time travel and lineage to answer audit-style questions
  • Peer review, and next steps such as full CDC and CI/CD for data work

Schedule & training options

For individuals — public rounds

No public rounds are open right now. Join the waiting list and we will contact you first when the next round opens, or ask us on LINE. Or call 02-570-8449 or 088-807-9770

For organisations — in-house / private

  • Tailor the content to your team’s tools and projects
  • Your dates, at your office or live online
  • Quotation with tax ID for procurement
Corporate training quote

Instructors

Frequently asked questions

How does this course differ from Data Pipeline with Apache Airflow and dbt?

Data Pipeline with Apache Airflow and dbt teaches those two tools in depth, from installing Airflow, writing DAGs and scheduling to a full dbt project with tests on a warehouse. This course is about lakehouse architecture. It uses Airflow and dbt as working tools with their basics kept short, and spends most of its time on medallion design on object storage with Apache Iceberg and Trino, ingestion patterns and CDC concepts, data contracts, time travel, schema evolution, Iceberg table maintenance, observability and lineage, ending with a capstone. If you have not used Airflow or dbt before, we suggest taking the Data Pipeline course first or preparing the basics in advance.

What machine do I need for the labs?

All labs run with Docker Compose on your own laptop or VM: PostgreSQL, object storage, an Iceberg catalog, Trino and Airflow. Airflow alone needs at least 4 GB of memory for Docker, and Trino also needs a fair amount, so we recommend 16 GB RAM or more (12 GB at minimum with other programs closed), 4 CPU cores and about 40 GB of free disk, with permission to install Docker. Debezium and Kafka are an instructor demo, so they do not run on your machine.

Why does the course talk about S3-compatible object storage rather than MinIO alone?

In 2026 the MinIO community edition moved to source-only distribution and its main repository is no longer maintained. The course therefore teaches concepts that apply to any S3-compatible object storage and picks a suitable option for the labs on the training day, such as MinIO built from source or SeaweedFS. Iceberg and Trino connect through the same S3 API either way. For production use, choose object storage that is actively maintained or covered by a support agreement that fits your organisation's policy.

Do I need to bring real company data?

No. Every lab uses a simulated deposit system generated by scripts, and everything runs on your own machine with no data sent to outside services. This lets you keep experimenting inside your organisation without real personal data, in line with PDPA principles. If you want the labs adapted to your own systems, the course can be run in-house.

Which versions of Airflow and dbt does the course use?

The course uses Apache Airflow 3 with DAGs written using the TaskFlow API. Transformation uses dbt Core 1.x with the latest compatible dbt-trino adapter on the training day. dbt v2, released in 2026, uses a new adapter model, so the course gives an overview of v2 while the labs use the combination that already supports Trino. These tools release often, and the content follows the latest compatible versions on the training day.

Which course should I take if my organisation uses Microsoft Fabric?

This course uses open source tools that can be installed inside your organisation. Concepts such as the medallion architecture, data quality, data contracts and lineage carry over to Fabric, but if you want hands-on work in Fabric itself, with OneLake, Lakehouse, pipelines and notebooks, we recommend the institute's Microsoft Fabric Essentials course.