Section 7: Excessive Agency and the Agentic Top 10
- Excessive agency: more permissions, functions and autonomy than the agent needs
- The OWASP Top 10 for Agentic Applications: goal hijack, tool misuse and identity abuse
- Memory poisoning, cascading failures, rogue agents and over-trusting agents
- Workshop: threat model a sample agent that reads email and calls APIs
Section 8: Lab: Testing AI Agents, Tools and MCP
- Agent goal hijack through data the agent reads while it works
- Tool misuse: chaining permitted tools into a harmful outcome
- MCP server risks such as tampered tool descriptions and overly broad permissions
- Lab: lead the demo agent into sending data out of the system, then close the gap with least privilege
Section 9: Lab: Automated Red Teaming with promptfoo
- Configure promptfoo to test the target application through an HTTP or Python provider
- Choose plugins and strategies that match the application's OWASP risks
- Read the report, separate false positives and prioritise what to fix
- Lab: scan the demo application before and after the fixes and compare the results
Section 10: Lab: Guardrails and Defence in Depth
- Layers of defence: input checks, output checks, tool permission separation and human-in-the-loop
- Separating untrusted data from instructions and limiting the blast radius when the model is fooled
- Human confirmation before irreversible actions, and audit logs for agents
- Lab: add guardrails to the demo application and measure them with the same test suite
Section 11: Continuous Testing and Production Monitoring
- Run red team test suites in CI/CD to catch regressions when prompts or models change
- Logging and monitoring for LLM applications to detect real attacks
- PDPA considerations when prompts and logs contain personal data
- Lab: set up GitHub Actions to run promptfoo whenever a prompt changes
Section 12: Workshop: Capstone Red Team Engagement
- Receive a new AI agent application with its scope and rules of engagement
- Threat model it, test it by hand and scan it with promptfoo
- Rate the severity of each finding and map it to the OWASP lists
- Write a findings report with recommended fixes, then present and review it together