Get in Touch

Course Outline

Performance Concepts and Metrics

  • Latency, throughput, power consumption, and resource utilisation
  • Distinguishing between system-level and model-level bottlenecks
  • Profiling strategies for inference versus training

Profiling on Huawei Ascend

  • Utilising CANN Profiler and MindInsight
  • Kernel and operator diagnostics
  • Offload patterns and memory mapping techniques

Profiling on Biren GPU

  • Biren SDK performance monitoring capabilities
  • Kernel fusion, memory alignment, and execution queues
  • Power and temperature-aware profiling

Profiling on Cambricon MLU

  • BANGPy and Neuware performance tools
  • Kernel-level visibility and log analysis
  • Integration of the MLU profiler with deployment frameworks

Graph and Model-Level Optimisation

  • Graph pruning and quantisation strategies
  • Operator fusion and computational graph restructuring
  • Input size standardisation and batch tuning

Memory and Kernel Optimisation

  • Optimising memory layout and reuse strategies
  • Efficient buffer management across different chipsets
  • Platform-specific kernel-level tuning techniques

Cross-Platform Best Practices

  • Achieving performance portability through abstraction strategies
  • Developing shared tuning pipelines for multi-chip environments
  • Case study: tuning an object detection model across Ascend, Biren, and MLU

Summary and Next Steps

Requirements

  • Hands-on experience with AI model training or deployment pipelines
  • Working knowledge of GPU/MLU compute principles and model optimisation
  • Basic familiarity with performance profiling tools and key metrics

Target Audience

  • Performance engineers
  • Machine learning infrastructure teams
  • AI system architects
 21 Hours

Upcoming Courses

Related Categories