Databases · DBC-44

Data Pipeline with Apache Airflow and dbt

18 hours 3 days
Last updated
Data Pipeline ด้วย Apache Airflow และ dbt

Data Pipeline with Apache Airflow and dbt is a 18-hour training course by IT Genius Institute. Many organizations already own BI tools such as Power BI or Tableau, yet still struggle with data preparation: pulling data from several systems by hand,…

Training schedule

No public rounds are open right now — register your interest and we will contact you when the next round opens, or request an in-house session for your team.

Corporate training quote

Many organizations already own BI tools such as Power BI or Tableau, yet still struggle with data preparation: pulling data from several systems by hand, cleaning it in Excel, and constantly finding that numbers do not match between reports. This course fills the data engineering layer that sits between source databases and reports, using two industry-standard tools: Apache Airflow for scheduling and orchestrating jobs, and dbt for transforming data with SQL plus automatic testing and

documentation. Learners build a real pipeline end to end: designing an ELT architecture, writing Airflow DAGs, ingesting data from multiple sources into a warehouse, transforming it into staging and mart layers with dbt, writing tests that catch bad data before users see it, building incremental models that process only new data, and finally connecting Airflow with dbt and taking the pipeline to production with monitoring. (3 days, 6 hours per day, 18 hours in total, Intermediate level.)

Objectives

  • Understand modern data architecture and the difference between ETL and ELT
  • Install and operate Apache Airflow and understand its core components
  • Write DAGs and choose the right operator for each kind of task
  • Schedule jobs, define dependencies and handle failures and backfills
  • Ingest data from databases, files and APIs into a warehouse systematically
  • Transform data with dbt into maintainable staging and mart layers
  • Write data quality tests and generate documentation and lineage automatically
  • Build incremental models that process only changed data
  • Connect Airflow with dbt and run the pipeline in production with monitoring

Who this course is for

  • Data engineers and anyone moving into data engineering
  • Data analysts and BI developers tired of preparing data by hand
  • Database administrators responsible for delivering data to analytics teams
  • Developers who must build scheduled data processing jobs
  • Data teams who want to fix mismatched report numbers systematically

Prerequisites

  • Working SQL skills (SELECT, JOIN, GROUP BY)
  • Understanding of relational databases and tables
  • Ability to read and edit basic Python (developer-level skill not required)
  • Prior Docker and Git experience helps but is not required

Curriculum

Course Details

A 3-day course, 6 hours per day (18 hours in total), delivered as lectures with hands-on labs for every topic. Intermediate level. Learners take home a lab guide, complete DAG code and dbt project samples, practice datasets and a data quality checklist. All tools are open source and can be used in the organization afterwards.

Day 1: Data Architecture and Airflow

Section 1: Modern Data Architecture

  • Common problems: manual extraction, mismatched numbers and unknown data provenance
  • ETL versus ELT and why ELT dominates with modern warehouses
  • Data platform components: sources, ingestion, warehouse, transformation and BI
  • Designing staging, intermediate and mart layers

Section 2: Installing and Using Apache Airflow

  • Airflow components: scheduler, webserver, worker, metadata database and executor
  • Installing Airflow with Docker Compose and using the web interface
  • DAG, task and task instance concepts and reading job status on screen
  • Lab: install Airflow and run your first DAG

Section 3: Writing DAGs and Operators

  • Writing DAGs in Python and using the modern TaskFlow API
  • Common operators: Python, Bash, SQL and sensors for waiting on conditions
  • Passing values between tasks with XCom and using variables and connections securely
  • Lab: write a DAG that pulls data from a file and an API and stores it in a database

Section 4: Scheduling and Handling Failures

  • Setting schedules and understanding data intervals and catchup
  • Defining task dependencies and running tasks in parallel
  • Automatic retries, failure notifications and backfilling historical runs
  • Lab: schedule a daily job and test failure and backfill scenarios

Day 2: Ingestion and dbt

Section 5: Ingesting Data into the Warehouse

  • Extracting from relational databases, CSV files and APIs
  • Full load versus incremental load and when each applies
  • Handling mid-run failures and making jobs idempotent
  • Lab: build a DAG that ingests several sources into the raw layer

Section 6: Introduction to dbt and Building Models

  • What dbt is and how it turns SQL transformation into disciplined engineering
  • Setting up a dbt project, profiles and the warehouse connection
  • Writing models, using ref and source, and view versus table materialization
  • Lab: build staging models from the raw data you ingested

Section 7: Data Testing and Documentation

  • Built-in tests: not_null, unique, accepted_values and relationships
  • Writing custom SQL tests for business rules
  • Generating documentation automatically and reading lineage graphs
  • Lab: add tests and documentation to your models and watch a test catch bad data

Section 8: Incremental Models and Project Structure

  • Building incremental models to process only new data and cut runtime
  • Structuring the project into staging, intermediate and mart layers
  • Using macros and Jinja to avoid repeating SQL, and seeds for reference tables
  • Lab: convert a model to incremental and measure the time saved

Day 3: Integration and Production

Section 9: Connecting Airflow with dbt

  • Patterns for invoking dbt from Airflow and separating run and test steps
  • Designing an end-to-end DAG: ingest, transform, test and notify
  • Managing dependencies between ingestion and transformation jobs
  • Lab: build a full pipeline where Airflow drives dbt automatically

Section 10: Data Quality and Monitoring

  • Placing quality gates at each layer and deciding when to stop the pipeline
  • Alerting on data anomalies or pipeline failure via email or team channels
  • Tracking data freshness and communicating with report users
  • Lab: add quality checks and alerting to your own pipeline

Section 11: Going to Production

  • Separating dev and production environments and handling credentials safely
  • Keeping pipeline code in Git and reviewing through pull requests
  • CI/CD approaches for data work and testing before production
  • Planning long-term pipeline maintenance and handover to the team

Section 12: Capstone and Next Steps

  • Capstone: build a complete pipeline from raw data to BI-ready tables
  • Present your architecture and the reasoning behind your decisions
  • Moving on to cloud data warehouses and other data engineering tooling
  • Workshop: present the pipeline you built and review it with the instructor

Frequently asked questions

Who is Data Pipeline with Apache Airflow and dbt for, and what background is needed?

Built for Data engineers and anyone moving into data engineering · Data analysts and BI developers tired of preparing data by hand · Database administrators responsible for delivering data to analytics teams Background you should have: Working SQL skills (SELECT, JOIN, GROUP BY) · Understanding of relational databases and tables Not sure the fit is right? Talk to our team on LINE @itgenius or call 02-570-8449.

How much does Data Pipeline with Apache Airflow and dbt cost and how long does it run?

THB 12,500 (currently THB 11,250 on promotion). The course runs 18 hours. The fee covers course materials, lunch and refreshments throughout. Pay by bank transfer to the company account, confirm it on our payment page, and we can issue the receipt or tax invoice in your company's name.

Do I get a certificate?

Yes. Everyone who completes the course receives a Certificate of Completion from IT Genius Institute. Each certificate carries its own number, and anyone holding that number can verify it online on our certificate page, so you can add it to your portfolio or pass it to HR as evidence of training.

Where does the training take place, and is there an online option?

You can attend onsite at IT Genius Institute or arrange to join online, and we also run it as a private in-house session for your team. Ask about dates and venues on LINE @itgenius or call 02-570-8449.

What if I fall behind or miss a session — can I retake it?

Yes. You may retake the same course free of charge in a later round, under the institute's conditions. Tell our team which course and round you attended, and we will check it and offer you the rounds that still have seats. Ask us on LINE @itgenius or call 02-570-8449.

How do I enrol, or request a quotation for my company?

Enrol online with the registration form on this page. You can register several attendees at once and enter your tax ID and billing address for the tax invoice. Or request a company quotation straight from the quote button. For anything else call 02-570-8449 or reach us on LINE @itgenius.

Instructors

Run this course for your whole team

We run this course in-house, tailored to your stack.

Corporate training quote