Section 7: Lab: An Evaluation Gate in GitHub Actions
- Run the eval suite automatically on every pull request that changes a prompt or code
- Set pass thresholds and handle score noise from non-deterministic models
- Post a comparison with the previous run as a pull request comment
- Lab: a pipeline that blocks the merge when faithfulness falls below the threshold
Section 8: Lab: Deploying with Docker and an LLM Gateway
- Containerise the RAG app and handle config and secrets safely
- Use LiteLLM as a central gateway to switch models without code changes
- Configure fallbacks, retries, timeouts and rate limits per team or per app
- Lab: deploy the app behind the gateway and simulate a primary model outage to test fallback
Section 9: Lab: Online Evaluation and Monitoring
- Sample production traces and score them continuously with LLM-as-a-judge
- Dashboards for quality, latency, error rate and cost
- Alerts when scores drop or latency and cost spike
- Lab: close the loop from poor production feedback to a new test case in CI
Section 10: Lab: Controlling Cost and Latency
- Track tokens and cost by user, feature and prompt version
- Provider prompt caching, and a semantic cache for repeated questions
- Model routing: easy tasks to small models, hard tasks to large ones
- Lab: cut the cost per question, then prove with evals that quality has not dropped
Section 11: Guardrails, Safety and Governance
- Mask personal data in traces and set retention periods in line with PDPA
- Input and output guardrails and their impact on latency
- An incident plan for when an LLM app gives wrong answers or leaks data
- Workshop: write a release checklist for the team's LLM applications
Section 12: Workshop: Capstone LLMOps Pipeline
- Take on a RAG app with quality problems and high costs
- Analyse it with traces, build a dataset and set evaluation criteria
- Fix the prompt or retrieval through a pull request that must pass the evaluation gate
- Deploy through the gateway, switch on monitoring and prove the result with numbers