Section 11: Apache Iceberg in Depth
- The problems an open table format solves for Parquet files sitting in a data lake
- Inside a table: metadata files, manifests, snapshots and the catalog
- ACID transactions and concurrent writes from several engines
- A brief comparison of Iceberg and Delta Lake
Section 12: Lab: Time Travel for Audit and Review
- Inspect snapshots and table history through metadata tables
- Query data as of a point in time with FOR TIMESTAMP AS OF and FOR VERSION AS OF
- Roll a table back to an earlier snapshot after a bad load
- Lab: answer what the balances were at the end of the previous day
Section 13: Schema Evolution and Partitioning
- Add, rename or widen columns without rewriting data
- Hidden partitioning and choosing partitions to match query patterns
- Partition evolution as data volumes change
- Lab: add a transaction channel column and change the partitioning of the transactions table
Section 14: Lab: Table Maintenance and Performance
- The small files problem caused by frequent loads
- Compaction with optimize and removing old snapshots with expire_snapshots
- Remove unreferenced files and schedule table maintenance with Airflow
- Lab: measure query time before and after compaction
Section 15: Data Observability and Data Lineage
- Monitor freshness, volume and schema drift at each layer
- Show lineage from the source system to the report with dbt docs
- Lineage expectations in financial services, for example the BCBS 239 principles
- Lab: alert when data is stale or row counts look abnormal
Section 16: Workshop: Capstone Project
- Run the whole pipeline from new source transactions to the Gold tables
- Present the architecture and the data quality control point at each layer
- Demonstrate time travel and lineage to answer audit-style questions
- Peer review, and next steps such as full CDC and CI/CD for data work