Data Pipeline with Apache Airflow and dbt is a 18-hour training course by IT Genius Institute. Many organizations already own BI tools such as Power BI or Tableau, yet still struggle with data preparation: pulling data from several systems by hand,…
Training schedule
No public rounds are open right now — register your interest and we will contact you when the next round opens, or request an in-house session for your team.
Many organizations already own BI tools such as Power BI or Tableau, yet still struggle with data preparation: pulling data from several systems by hand, cleaning it in Excel, and constantly finding that numbers do not match between reports. This course fills the data engineering layer that sits between source databases and reports, using two industry-standard tools: Apache Airflow for scheduling and orchestrating jobs, and dbt for transforming data with SQL plus automatic testing and
documentation. Learners build a real pipeline end to end: designing an ELT architecture, writing Airflow DAGs, ingesting data from multiple sources into a warehouse, transforming it into staging and mart layers with dbt, writing tests that catch bad data before users see it, building incremental models that process only new data, and finally connecting Airflow with dbt and taking the pipeline to production with monitoring. (3 days, 6 hours per day, 18 hours in total, Intermediate level.)
Objectives
Understand modern data architecture and the difference between ETL and ELT
Install and operate Apache Airflow and understand its core components
Write DAGs and choose the right operator for each kind of task
Schedule jobs, define dependencies and handle failures and backfills
Ingest data from databases, files and APIs into a warehouse systematically
Transform data with dbt into maintainable staging and mart layers
Write data quality tests and generate documentation and lineage automatically
Build incremental models that process only changed data
Connect Airflow with dbt and run the pipeline in production with monitoring
Who this course is for
Data engineers and anyone moving into data engineering
Data analysts and BI developers tired of preparing data by hand
Database administrators responsible for delivering data to analytics teams
Developers who must build scheduled data processing jobs
Data teams who want to fix mismatched report numbers systematically
Prerequisites
Working SQL skills (SELECT, JOIN, GROUP BY)
Understanding of relational databases and tables
Ability to read and edit basic Python (developer-level skill not required)
Prior Docker and Git experience helps but is not required
Curriculum
Course Details
A 3-day course, 6 hours per day (18 hours in total), delivered as lectures with hands-on labs for every topic. Intermediate level. Learners take home a lab guide, complete DAG code and dbt project samples, practice datasets and a data quality checklist. All tools are open source and can be used in the organization afterwards.
Day 1: Data Architecture and Airflow
Section 1: Modern Data Architecture
Common problems: manual extraction, mismatched numbers and unknown data provenance
ETL versus ELT and why ELT dominates with modern warehouses
Data platform components: sources, ingestion, warehouse, transformation and BI
Designing staging, intermediate and mart layers
Section 2: Installing and Using Apache Airflow
Airflow components: scheduler, webserver, worker, metadata database and executor
Installing Airflow with Docker Compose and using the web interface
DAG, task and task instance concepts and reading job status on screen
Lab: install Airflow and run your first DAG
Section 3: Writing DAGs and Operators
Writing DAGs in Python and using the modern TaskFlow API
Common operators: Python, Bash, SQL and sensors for waiting on conditions
Passing values between tasks with XCom and using variables and connections securely
Lab: write a DAG that pulls data from a file and an API and stores it in a database
Section 4: Scheduling and Handling Failures
Setting schedules and understanding data intervals and catchup
Defining task dependencies and running tasks in parallel
Automatic retries, failure notifications and backfilling historical runs
Lab: schedule a daily job and test failure and backfill scenarios
Day 2: Ingestion and dbt
Section 5: Ingesting Data into the Warehouse
Extracting from relational databases, CSV files and APIs
Full load versus incremental load and when each applies
Handling mid-run failures and making jobs idempotent
Lab: build a DAG that ingests several sources into the raw layer
Section 6: Introduction to dbt and Building Models
What dbt is and how it turns SQL transformation into disciplined engineering
Setting up a dbt project, profiles and the warehouse connection
Writing models, using ref and source, and view versus table materialization
Lab: build staging models from the raw data you ingested
Section 7: Data Testing and Documentation
Built-in tests: not_null, unique, accepted_values and relationships
Writing custom SQL tests for business rules
Generating documentation automatically and reading lineage graphs
Lab: add tests and documentation to your models and watch a test catch bad data
Section 8: Incremental Models and Project Structure
Building incremental models to process only new data and cut runtime
Structuring the project into staging, intermediate and mart layers
Using macros and Jinja to avoid repeating SQL, and seeds for reference tables
Lab: convert a model to incremental and measure the time saved
Day 3: Integration and Production
Section 9: Connecting Airflow with dbt
Patterns for invoking dbt from Airflow and separating run and test steps
Designing an end-to-end DAG: ingest, transform, test and notify
Managing dependencies between ingestion and transformation jobs
Lab: build a full pipeline where Airflow drives dbt automatically
Section 10: Data Quality and Monitoring
Placing quality gates at each layer and deciding when to stop the pipeline
Alerting on data anomalies or pipeline failure via email or team channels
Tracking data freshness and communicating with report users
Lab: add quality checks and alerting to your own pipeline
Section 11: Going to Production
Separating dev and production environments and handling credentials safely
Keeping pipeline code in Git and reviewing through pull requests
CI/CD approaches for data work and testing before production
Planning long-term pipeline maintenance and handover to the team
Section 12: Capstone and Next Steps
Capstone: build a complete pipeline from raw data to BI-ready tables
Present your architecture and the reasoning behind your decisions
Moving on to cloud data warehouses and other data engineering tooling
Workshop: present the pipeline you built and review it with the instructor
Frequently asked questions
Who is Data Pipeline with Apache Airflow and dbt for, and what background is needed?
Built for Data engineers and anyone moving into data engineering · Data analysts and BI developers tired of preparing data by hand · Database administrators responsible for delivering data to analytics teams Background you should have: Working SQL skills (SELECT, JOIN, GROUP BY) · Understanding of relational databases and tables Not sure the fit is right? Talk to our team on LINE @itgenius or call 02-570-8449.
How much does Data Pipeline with Apache Airflow and dbt cost and how long does it run?
THB 12,500 (currently THB 11,250 on promotion). The course runs 18 hours. The fee covers course materials, lunch and refreshments throughout. Pay by bank transfer to the company account, confirm it on our payment page, and we can issue the receipt or tax invoice in your company's name.
Do I get a certificate?
Yes. Everyone who completes the course receives a Certificate of Completion from IT Genius Institute. Each certificate carries its own number, and anyone holding that number can verify it online on our certificate page, so you can add it to your portfolio or pass it to HR as evidence of training.
Where does the training take place, and is there an online option?
You can attend onsite at IT Genius Institute or arrange to join online, and we also run it as a private in-house session for your team. Ask about dates and venues on LINE @itgenius or call 02-570-8449.
What if I fall behind or miss a session — can I retake it?
Yes. You may retake the same course free of charge in a later round, under the institute's conditions. Tell our team which course and round you attended, and we will check it and offer you the rounds that still have seats. Ask us on LINE @itgenius or call 02-570-8449.
How do I enrol, or request a quotation for my company?
Enrol online with the registration form on this page. You can register several attendees at once and enter your tax ID and billing address for the tax invoice. Or request a company quotation straight from the quote button. For anything else call 02-570-8449 or reach us on LINE @itgenius.
Producing video content every day is one of the most time-consuming tasks for business owners and marketing teams today. Editing a single clip with click-and-drag tools (CapCut, Premiere Pro) takes 1-3 hours, including cutting silences, adding subtitles,…
Systems wired together with direct point-to-point calls turn into a tangle as services multiply, and when a downstream system goes down the data goes with it. Apache Kafka solves this by acting as a durable central hub for events, so producers and consumers…
18 hours3 days
11,250THB12,500 THB
View details
Run this course for your whole team
We run this course in-house, tailored to your stack.