Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Agentic Systems in Production
- Agentic architectures: loops, tools, memory, and orchestration layers.
- Agent lifecycle: development, deployment, and continuous operation.
- Challenges associated with managing agents at production scale.
Infrastructure and Deployment Models
- Deploying agents within containerised and cloud environments.
- Scaling patterns: comparing horizontal versus vertical scaling, concurrency, and throttling.
- Orchestration of multi-agent systems and workload balancing.
Monitoring and Observability
- Key metrics: latency, success rate, memory usage, and agent call depth.
- Tracing agent activity and visualising call graphs.
- Instrumenting observability using Prometheus, OpenTelemetry, and Grafana.
Logging, Auditing, and Compliance
- Centralised logging and structured event collection.
- Ensuring compliance and auditability within agentic workflows.
- Designing audit trails and replay mechanisms for effective debugging.
Performance Tuning and Resource Optimisation
- Reducing inference overhead and optimising agent orchestration cycles.
- Utilising model caching and lightweight embeddings for faster retrieval.
- Conducting load testing and stress scenarios for AI pipelines.
Cost Control and Governance
- Identifying agent cost drivers: API calls, memory, compute, and external integrations.
- Tracking agent-level costs and implementing chargeback models.
- Enacting automation policies to prevent agent sprawl and idle resource consumption.
CI/CD and Rollout Strategies for Agents
- Integrating agent pipelines into CI/CD systems.
- Strategies for testing, versioning, and rolling back iterative agent updates.
- Implementing progressive rollouts and safe deployment mechanisms.
Failure Recovery and Reliability Engineering
- Designing for fault tolerance and graceful degradation.
- Applying retry, timeout, and circuit breaker patterns to ensure agent reliability.
- Establishing incident response and post-mortem frameworks for AI operations.
Capstone Project
- Building and deploying an agentic AI system with comprehensive monitoring and cost tracking.
- Simulating load, measuring performance, and optimising resource usage.
- Presenting the final architecture and monitoring dashboard to peers.
Summary and Next Steps
Requirements
- A robust understanding of MLOps and production machine learning systems.
- Practical experience with containerised deployments (Docker/Kubernetes).
- Familiarity with cloud cost optimisation and observability tools.
Target Audience
- MLOps engineers.
- Site Reliability Engineers (SREs).
- Engineering managers overseeing AI infrastructure.
21 Hours
Testimonials (3)
The trainer is patient and very helpful. He knows the topic well.
CLIFFORD TABARES - Universal Leaf Philippines, Inc.
Course - Agentic AI for Business Automation: Use Cases & Integration
Good mixvof knowledge and practice
Ion Mironescu - Facultatea S.A.I.A.P.M.
Course - Agentic AI for Enterprise Applications
The mix of theory and practice and of high level and low level perspectives