DevOps · DOP-06

LLMOps: Evaluation, Observability and CI/CD

LLMOps: Evaluation, Observability and CI/CD is a course in taking RAG and chatbot applications to production with discipline, covering tracing with Langfuse, prompt versioning, evaluation with Ragas and promptfoo, quality gates in CI/CD, monitoring and cost control. It suits AI engineers, DevOps engineers and tech leads who look after LLM applications in production.

Updated
From 8,010 THB / person 8,900 −10% excl. VAT 7% · group rates available
PDFDownload the course outline
  • Duration12 hours · 2 days
  • FormatOnsite / live online
  • Next roundOn request
  • CertificateIncluded

Course overview

Getting a chatbot or RAG system to answer well on a developer's laptop does not take long. Keeping it answering well, consistently, in production is another matter. A small prompt edit quietly makes some answers worse, a model upgrade multiplies the bill, and when users report a wrong answer the team cannot tell which step caused it. Because LLM applications are not deterministic, the usual ways of testing and deploying software are not enough.

This course builds on the RAG and AI chatbot courses and shows learners how to take a RAG application to production in a disciplined way. Learners install Langfuse to trace every step of the app, manage prompts with versions, build evaluation datasets from real data, measure RAG quality with Ragas and test behaviour with promptfoo. They then turn those evaluations into a gate in GitHub Actions, deploy with Docker behind an LLM gateway, and monitor quality, latency and cost after release. The course closes with a capstone that builds a complete pipeline, from a prompt change to production monitoring. (2 days, 6 hours per day, 12 hours in total, Intermediate level.)

What you’ll gain

  • Explain the LLMOps lifecycle and how it differs from MLOps and traditional DevOps
  • Install Langfuse and trace a RAG application end to end, from retrieval to answer
  • Manage prompts with versions, separate production from staging and roll back when needed
  • Build evaluation datasets from real traces and synthetic data
  • Evaluate RAG quality with Ragas and test application behaviour with promptfoo
  • Set up an evaluation gate in CI/CD so changes that lower quality never reach production
  • Deploy LLM applications with Docker and an LLM gateway, with fallbacks and rate limits
  • Monitor quality, latency and cost after release, and put guardrails in place

Who this course is for

  • AI engineers and developers who have built a chatbot or RAG system and need to take it to production
  • DevOps and platform engineers who run LLM applications alongside the development team
  • ML engineers who want to extend their MLOps practice to LLM work
  • Tech leads who set testing and deployment standards for the team's AI applications
  • Graduates of RAG & Knowledge Base or one of the AI chatbot courses

Prerequisites

  • Python, and experience building an LLM application or RAG system, or completion of RAG & Knowledge Base
  • Basic Git, GitHub and Docker skills
  • A basic understanding of CI/CD
  • An LLM API key with credit and a GitHub account of your own, plus a laptop that can run Docker Desktop

Curriculum

Course Details

A 2-day course, 6 hours per day (12 hours in total, 09:00-16:00), alternating lectures with labs on a single sample RAG application throughout. Intermediate level. The course builds on RAG & Knowledge Base and the AI chatbot courses. It does not teach how to build RAG from scratch, and instead focuses on taking an application to production in a disciplined way, covering tracing, prompt versioning, evaluation,

CI/CD, monitoring and cost control, so it pairs well with the DevOps track. Every lab runs on the learner's own machine with Docker. Learners bring their own or company LLM API key with credit and a GitHub account; API usage and GitHub Actions minutes are billed by actual use. Learners take home a lab guide, the sample RAG application, docker-compose files for Langfuse and LiteLLM, evaluation suites and GitHub Actions workflows.

Day 1 Tracing, Prompt Versioning and Evaluation

Section 1: How LLMOps Differs from MLOps and DevOps

  • Problems specific to LLM apps: non-deterministic output, hard-to-measure quality and usage-driven cost
  • The LLMOps lifecycle: develop, evaluate, deploy, monitor and feed real data back into improvement
  • What to measure: answer quality, groundedness, latency, cost and safety
  • Lab: set up the sample RAG app and try to find which step caused a wrong answer without any tooling

Section 2: Lab: Tracing with Langfuse

  • Self-host Langfuse with Docker Compose
  • Instrument the RAG app with the Python SDK to see retrieval, prompts and model calls
  • Traces, spans, sessions, tokens and cost per request, following the OpenTelemetry approach for GenAI
  • Lab: follow the trace of a wrong answer until the problem turns out to be retrieval

Section 3: Lab: Prompt Management and Versioning

  • Keeping prompts in Git compared with Langfuse prompt management
  • Versions and labels such as production and staging, together with model config
  • Link traces to prompt versions to compare results before and after a change
  • Lab: release a new prompt version, then roll it back when a problem appears

Section 4: Lab: Building Evaluation Datasets

  • What a golden dataset is and which kinds of questions it should contain
  • Pull real traces with problems or poor feedback into test cases
  • Generate synthetic questions from documents with Ragas, then have people review them
  • Lab: build a 50-item dataset from the sample app's documents and traces

Section 5: Lab: Evaluating RAG with Ragas

  • Core RAG metrics: faithfulness, answer relevance, context precision and context recall
  • How LLM-as-a-judge works and the weaknesses to watch for
  • Use the individual scores to tell retrieval problems from generation problems
  • Lab: evaluate the RAG app before and after a chunking change and decide whether it really improved

Section 6: Lab: Behaviour Testing with promptfoo

  • Write assertion-based test cases: required keywords, valid JSON and staying on topic
  • Compare several prompts and models in one matrix
  • Use LLM-graded assertions for criteria that rules cannot check
  • Lab: choose the best-value model on quality, speed and cost
Day 2 CI/CD, Production and Cost Control

Section 7: Lab: An Evaluation Gate in GitHub Actions

  • Run the eval suite automatically on every pull request that changes a prompt or code
  • Set pass thresholds and handle score noise from non-deterministic models
  • Post a comparison with the previous run as a pull request comment
  • Lab: a pipeline that blocks the merge when faithfulness falls below the threshold

Section 8: Lab: Deploying with Docker and an LLM Gateway

  • Containerise the RAG app and handle config and secrets safely
  • Use LiteLLM as a central gateway to switch models without code changes
  • Configure fallbacks, retries, timeouts and rate limits per team or per app
  • Lab: deploy the app behind the gateway and simulate a primary model outage to test fallback

Section 9: Lab: Online Evaluation and Monitoring

  • Sample production traces and score them continuously with LLM-as-a-judge
  • Dashboards for quality, latency, error rate and cost
  • Alerts when scores drop or latency and cost spike
  • Lab: close the loop from poor production feedback to a new test case in CI

Section 10: Lab: Controlling Cost and Latency

  • Track tokens and cost by user, feature and prompt version
  • Provider prompt caching, and a semantic cache for repeated questions
  • Model routing: easy tasks to small models, hard tasks to large ones
  • Lab: cut the cost per question, then prove with evals that quality has not dropped

Section 11: Guardrails, Safety and Governance

  • Mask personal data in traces and set retention periods in line with PDPA
  • Input and output guardrails and their impact on latency
  • An incident plan for when an LLM app gives wrong answers or leaks data
  • Workshop: write a release checklist for the team's LLM applications

Section 12: Workshop: Capstone LLMOps Pipeline

  • Take on a RAG app with quality problems and high costs
  • Analyse it with traces, build a dataset and set evaluation criteria
  • Fix the prompt or retrieval through a pull request that must pass the evaluation gate
  • Deploy through the gateway, switch on monitoring and prove the result with numbers

Schedule & training options

For individuals — public rounds

No public rounds are open right now. Join the waiting list and we will contact you first when the next round opens, or ask us on LINE. Or call 02-570-8449 or 088-807-9770

For organisations — in-house / private

  • Tailor the content to your team’s tools and projects
  • Your dates, at your office or live online
  • Quotation with tax ID for procurement
Corporate training quote

Instructors

Frequently asked questions

Who is LLMOps: Evaluation, Observability and CI/CD for, and what background is needed?

Built for AI engineers and developers who have built a chatbot or RAG system and need to take it to production · DevOps and platform engineers who run LLM applications alongside the development team · ML engineers who want to extend their MLOps practice to LLM work Background you should have: Python, and experience building an LLM application or RAG system, or completion of RAG & Knowledge Base · Basic Git, GitHub and Docker skills Not sure the fit is right? Talk to our team on LINE @itgenius or call 02-570-8449.

How much does LLMOps: Evaluation, Observability and CI/CD cost and how long does it run?

THB 8,900 (currently THB 8,010 on promotion). The course runs 12 hours. The price excludes 7% VAT (for payment in a company's name). Pay by bank transfer to the company account, confirm it on our payment page, and we can issue the receipt or tax invoice in your company's name.

Do I get a certificate?

Yes. Everyone who completes the course receives a Certificate of Completion from IT Genius Institute. Each certificate carries its own number, and anyone holding that number can verify it online on our certificate page, so you can add it to your portfolio or pass it to HR as evidence of training.

Where does the training take place, and is there an online option?

You can attend onsite at IT Genius Institute or arrange to join online, and we also run it as a private in-house session for your team. Ask about dates and venues on LINE @itgenius or call 02-570-8449.

What if I fall behind or miss a session — can I retake it?

Yes. You may retake the same course free of charge in a later round, under the institute's conditions. Tell our team which course and round you attended, and we will check it and offer you the rounds that still have seats. Ask us on LINE @itgenius or call 02-570-8449.

How do I enrol, or request a quotation for my company?

Enrol online with the registration form on this page. You can register several attendees at once and enter your tax ID and billing address for the tax invoice. Or request a company quotation straight from the quote button. For anything else call 02-570-8449 or reach us on LINE @itgenius.