DevOps · DOP-02

Observability with Prometheus and Grafana

18 hours 3 days
Last updated
สร้างระบบ Observability ด้วย Prometheus และ Grafana

Observability with Prometheus and Grafana is a 18-hour training course by IT Genius Institute. When a system fails or slows down, the first question is how you find out and how fast. Many organizations still learn about problems from customer calls…

Training schedule

No public rounds are open right now — register your interest and we will contact you when the next round opens, or request an in-house session for your team.

Corporate training quote

When a system fails or slows down, the first question is how you find out and how fast. Many organizations still learn about problems from customer calls rather than from their own monitoring. Prometheus and Grafana are the open-source standard that DevOps and SRE teams worldwide use to collect metrics, visualize them on dashboards and alert before users are affected. This course builds an observability stack

from scratch: installing Prometheus, collecting metrics from servers, containers, databases and applications through exporters, writing PromQL to query and compute values, designing readable Grafana dashboards, configuring alert rules that do not create noise with Alertmanager, and finally long-term storage plus an introduction to logs and traces. Every topic is hands-on in a lab environment. (3 days, 6 hours per day, 18 hours in total, Intermediate level.)

Objectives

  • Understand observability concepts and the difference between metrics, logs and traces
  • Install and configure Prometheus and understand its pull-based architecture
  • Collect metrics from servers, containers and databases using exporters
  • Configure service discovery so new targets are monitored automatically
  • Write PromQL to query, compute rates and aggregate metric data
  • Design Grafana dashboards that are readable and useful to the team
  • Instrument your own application and monitor it using the four golden signals
  • Write alert rules and manage notifications effectively with Alertmanager
  • Plan long-term storage and extend into log and trace collection

Who this course is for

  • DevOps engineers and site reliability engineers responsible for system stability
  • System administrators and infrastructure teams who need monitoring that actually works
  • Developers who want to understand application behavior and performance in production
  • Teams running Docker and Kubernetes who need container monitoring
  • Technical leads who need to establish alerting and availability metrics

Prerequisites

  • Basic Linux command line skills (the Linux Administration course is recommended first)
  • Basic networking concepts such as ports, HTTP and DNS
  • Prior Docker experience helps but is not required
  • Programming is not required, but you should be able to read YAML configuration

Curriculum

Course Details

A 3-day course, 6 hours per day (18 hours in total), delivered as lectures with hands-on labs for every topic. Intermediate level. Learners take home a lab guide, Prometheus and Alertmanager configuration files, sample Grafana dashboards and a set of commonly used PromQL queries. All tools are open source and can be deployed in the organization at no licensing cost.

Day 1: Prometheus Fundamentals and Metric Collection

Section 1: Observability and Monitoring Concepts

  • The difference between monitoring and observability and why organizations need both
  • The three pillars: metrics, logs and traces, and choosing the right one for each problem
  • Key indicators: the four golden signals and SLI, SLO and error budget concepts
  • How Prometheus, Grafana and Alertmanager fit together

Section 2: Installing and Configuring Prometheus

  • The pull-based architecture and the Prometheus time series database
  • Installing Prometheus with Docker Compose and the structure of prometheus.yml
  • Metrics, labels and time series concepts and good metric naming
  • Lab: install Prometheus and monitor itself as the first target

Section 3: Collecting Metrics with Exporters

  • Collecting server metrics with Node Exporter (CPU, memory, disk, network)
  • Monitoring containers with cAdvisor and collecting database and web server metrics
  • Metric types: counter, gauge, histogram and summary, and when to use each
  • Lab: install several exporters and verify Prometheus is collecting everything

Section 4: Service Discovery and Target Management

  • Static target configuration and its limits as the system grows
  • Service discovery from files, DNS and cloud providers, and using it with Kubernetes
  • Using relabeling to organize labels and filter targets
  • Lab: configure service discovery so new targets are monitored automatically

Day 2: PromQL and Grafana Dashboards

Section 5: PromQL Fundamentals

  • Selecting data with selectors and filtering with label matchers
  • Common functions: rate, increase, sum, avg, max and using by and without
  • Calculating percentiles from histograms and spotting anomalies
  • Lab: write PromQL to answer real questions such as error rate and resource usage

Section 6: Grafana and Building Dashboards

  • Installing Grafana and connecting Prometheus as a data source
  • Panel types and choosing the right visualization for each kind of data
  • Using variables to build dashboards that can switch server or service
  • Lab: build a dashboard monitoring your own servers and containers

Section 7: Designing Dashboards That Work

  • Design principles so a dashboard immediately shows whether the system is healthy
  • Tailoring dashboards by audience: executives, operations teams and developers
  • Importing community dashboards and adapting them to the organization
  • Lab: refine a dashboard so it communicates system status on a single screen

Section 8: Application Monitoring

  • Instrumenting an application with a client library and exposing a metrics endpoint
  • Measuring the four golden signals: latency, traffic, errors and saturation
  • External monitoring with Blackbox Exporter to simulate the user perspective
  • Lab: add metrics to a sample application and build a dashboard to track it

Day 3: Alerting, Storage and Capstone

Section 9: Writing Alerting Rules

  • Writing alerting rules in PromQL and using the for duration to avoid noisy alerts
  • Setting severity levels and including the context the team needs to fix the problem
  • Symptom-based alerting: alert on what users feel, not on every internal metric
  • Lab: write alert rules for the servers and services you are monitoring

Section 10: Alertmanager and Notifications

  • Installing and configuring Alertmanager and connecting it to Prometheus
  • Grouping, silencing and inhibition to reduce alert noise
  • Sending notifications to email, LINE, Slack or your organization channels
  • Lab: send alerts to a real channel and test an outage scenario

Section 11: Long-term Storage and Performance

  • Setting retention periods and estimating the storage required
  • Long-term storage approaches and federation as the system scales
  • Performance tuning: scrape interval, label count and high cardinality problems
  • Backing up and maintaining the monitoring stack itself

Section 12: Logs, Traces and Capstone

  • Collecting logs with Loki and viewing logs alongside metrics in Grafana
  • Distributed tracing concepts and OpenTelemetry for following requests across services
  • Capstone: build a complete observability setup for one service from metrics to alerts
  • Workshop: present your dashboards and alert rules and receive feedback

Frequently asked questions

Who is Observability with Prometheus and Grafana for, and what background is needed?

Built for DevOps engineers and site reliability engineers responsible for system stability · System administrators and infrastructure teams who need monitoring that actually works · Developers who want to understand application behavior and performance in production Background you should have: Basic Linux command line skills (the Linux Administration course is recommended first) · Basic networking concepts such as ports, HTTP and DNS Not sure the fit is right? Talk to our team on LINE @itgenius or call 02-570-8449.

How much does Observability with Prometheus and Grafana cost and how long does it run?

THB 10,500 (currently THB 9,450 on promotion). The course runs 18 hours. The fee covers course materials, lunch and refreshments throughout. Pay by bank transfer to the company account, confirm it on our payment page, and we can issue the receipt or tax invoice in your company's name.

Do I get a certificate?

Yes. Everyone who completes the course receives a Certificate of Completion from IT Genius Institute. Each certificate carries its own number, and anyone holding that number can verify it online on our certificate page, so you can add it to your portfolio or pass it to HR as evidence of training.

Where does the training take place, and is there an online option?

You can attend onsite at IT Genius Institute or arrange to join online, and we also run it as a private in-house session for your team. Ask about dates and venues on LINE @itgenius or call 02-570-8449.

What if I fall behind or miss a session — can I retake it?

Yes. You may retake the same course free of charge in a later round, under the institute's conditions. Tell our team which course and round you attended, and we will check it and offer you the rounds that still have seats. Ask us on LINE @itgenius or call 02-570-8449.

How do I enrol, or request a quotation for my company?

Enrol online with the registration form on this page. You can register several attendees at once and enter your tax ID and billing address for the tax invoice. Or request a company quotation straight from the quote button. For anything else call 02-570-8449 or reach us on LINE @itgenius.

Instructors

Run this course for your whole team

We run this course in-house, tailored to your stack.

Corporate training quote