Whether delivered online or in person, our instructor-led live training courses on Graphics Processing Units (GPUs) combine interactive discussion with practical exercises to cover the core principles of GPU programming. These sessions are designed to help learners understand how to effectively utilize and program GPUs.
Choose between online live training (also known as remote live training) or onsite live training. The remote option is conducted via an interactive remote desktop session. Onsite training can be held directly at your premises in Durban or at NobleProg’s dedicated corporate training centres in Durban.
NobleProg -- Your Local Training Provider
Garden Court South Beach
Durban, South Africa
Way to get there
From Durban city centre: follow the beachfront/Marine Parade south toward South Beach.
By car/taxi: give the driver the address 73 O.R. Tambo Parade, South Beach, Durban 4001.
From King Shaka International Airport: the hotel is roughly 30 km ;from the airport.
This instructor-led course for Durban helps intermediate AI engineers construct and enhance neural network models via the Huawei Ascend platform and CANN toolkit. Attendees will learn to set up environments, create applications using MindSpore, and implement deployments in edge or cloud contexts.
This live, instructor-led training in Durban examines Huawei's AI ecosystem, ranging from the CANN SDK to the MindSpore framework. It is designed to assist beginner and intermediate professionals in comprehending how these components function together on Ascend hardware to streamline lifecycle management and deployment strategies.
This instructor-led, live training in Durban (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use OpenACC to program heterogeneous devices and exploit their parallelism.
By the end of this training, participants will be able to:
Set up an OpenACC development environment.
Write and run a basic OpenACC program.
Annotate code with OpenACC directives and clauses.
This instructor-led training in Durban focuses on deploying and optimising CV and NLP models using the CANN SDK for Ascend hardware. Participants will learn to convert models, integrate them into live pipelines, and enhance inference performance for real-time detection and analysis.
This instructor-led, live training in Durban (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to learn the basics of GPU programming and the main frameworks and tools for developing GPU applications.
By the end of this training, participants will be able to: Understand the difference between CPU and GPU computing and the benefits and challenges of GPU programming.
Choose the right framework and tool for their GPU application.
Create a basic GPU program that performs vector addition using one or more of the frameworks and tools.
Use the respective APIs, languages, and libraries to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the respective memory spaces, such as global, local, constant, and private, to optimize data transfers and memory accesses.
Use the respective execution models, such as work-items, work-groups, threads, blocks, and grids, to control the parallelism.
Debug and test GPU programs using tools such as CodeXL, CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize GPU programs using techniques such as coalescing, caching, prefetching, and profiling.
This instructor-led live training in Durban empowers advanced developers with the proficiency required to build, deploy, and fine-tune custom AI operators. Participants will gain mastery over CANN TIK and Apache TVM integration, enabling advanced optimisation and scheduling on Huawei Ascend hardware to achieve real-world performance gains.
This instructor-led, live training in Durban (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use different frameworks for GPU programming and compare their features, performance, and compatibility.
By the end of this training, participants will be able to:
Set up a development environment that includes OpenCL SDK, CUDA Toolkit, ROCm Platform, a device that supports OpenCL, CUDA, or ROCm, and Visual Studio Code.
Create a basic GPU program that performs vector addition using OpenCL, CUDA, and ROCm, and compare the syntax, structure, and execution of each framework.
Use the respective APIs to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the respective languages to write kernels that execute on the device and manipulate data.
Use the respective built-in functions, variables, and libraries to perform common tasks and operations.
Use the respective memory spaces, such as global, local, constant, and private, to optimize data transfers and memory accesses.
Use the respective execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test GPU programs using tools such as CodeXL, CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize GPU programs using techniques such as coalescing, caching, prefetching, and profiling.
This instructor-led training in Durban introduces CloudMatrix for scalable AI inference. It covers deploying, optimising, and monitoring models using CANN and MindSpore. Hands-on exercises address packaging, conversion, serving, and performance tuning for both real-time and batch workloads.
This live, instructor-led course in Durban delves into the foundational aspects and practical techniques for deploying AI models on Ascend edge devices via the CANN toolkit. It equips participants with the necessary skills to compile, optimise, and manage models within constrained computational environments.
This instructor-led, live training in Durban (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to install and use ROCm on Windows to program AMD GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes ROCm Platform, a AMD GPU, and Visual Studio Code on Windows.
Create a basic ROCm program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use ROCm API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use HIP language to write kernels that execute on the GPU and manipulate data.
Use HIP built-in functions, variables, and libraries to perform common tasks and operations.
Use ROCm and HIP memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use ROCm and HIP execution models to control the threads, blocks, and grids that define the parallelism.
Debug and test ROCm and HIP programs using tools such as ROCm Debugger and ROCm Profiler.
Optimize ROCm and HIP programs using techniques such as coalescing, caching, prefetching, and profiling.
This instructor-led, live training in Durban (online or onsite) targets beginner to intermediate-level developers eager to use ROCm and HIP to program AMD GPUs and harness their parallel computing power.
Upon completion of this training, participants will be able to:
Configure a development environment comprising the ROCm Platform, an AMD GPU, and Visual Studio Code.
Develop a fundamental ROCm program that executes vector addition on the GPU and retrieves results from GPU memory.
Employ the ROCm API to query device information, allocate and deallocate device memory, transfer data between host and device, launch kernels, and synchronise threads.
Utilise the HIP language to write kernels that execute on the GPU and manipulate data.
Leverage HIP built-in functions, variables, and libraries to carry out common tasks and operations.
Apply ROCm and HIP memory spaces, such as global, shared, constant, and local, to optimise data transfers and memory accesses.
Use ROCm and HIP execution models to manage the threads, blocks, and grids that define parallelism.
Debug and test ROCm and HIP programs using tools such as ROCm Debugger and ROCm Profiler.
Optimise ROCm and HIP programs through techniques including coalescing, caching, prefetching, and profiling.
This live training for Durban presents the CANN toolkit to AI framework developers. You will learn to establish environments, convert models, and deploy applications on Ascend hardware using MindSpore, TensorFlow, or PyTorch, covering the entire workflow from training to inference.
Optimise AI workloads on Ascend, Biren, and Cambricon through this hands-on training in Durban. Gain the ability to benchmark models, pinpoint bottlenecks, and apply graph, kernel, and operator-level optimisations. Fine-tune deployment pipelines to boost throughput and latency performance across these leading platforms.
Maximize neural network inference performance on Ascend AI processors with this advanced, instructor-led training in Durban. Delve into CANN's runtime architecture, utilising the Graph Engine, TIK, and TVM to master profiling, custom operator development, and resolving memory bottlenecks.
Transition CUDA applications to Chinese GPU architectures such as Huawei Ascend and Biren in Durban. This live course supports advanced developers in code translation and performance tuning, featuring practical labs for migrating CUDA codebases to new SDKs.
This instructor-led, live training in Durban (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use CUDA to program NVIDIA GPUs and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes CUDA Toolkit, an NVIDIA GPU, and Visual Studio Code.
Create a basic CUDA program that performs vector addition on the GPU and retrieves the results from the GPU memory.
Use the CUDA API to query device information, allocate and deallocate device memory, copy data between host and device, launch kernels, and synchronize threads.
Use the CUDA C/C++ language to write kernels that execute on the GPU and manipulate data.
Use CUDA built-in functions, variables, and libraries to perform common tasks and operations.
Use CUDA memory spaces, such as global, shared, constant, and local, to optimize data transfers and memory accesses.
Use the CUDA execution model to control the threads, blocks, and grids that define the parallelism.
Debug and test CUDA programs using tools such as CUDA-GDB, CUDA-MEMCHECK, and NVIDIA Nsight.
Optimize CUDA programs using techniques such as coalescing, caching, prefetching, and profiling.
This live training in Durban guides intermediate AI developers through deploying models on Ascend processors using the CANN toolkit. Learn to convert frameworks like PyTorch and TensorFlow, optimize performance, and debug issues for efficient edge and cloud inference scenarios.
This live training in Durban empowers developers with the skills to program and optimise applications on Biren AI accelerators. Participants will study the GPU architecture, set up the SDK, and translate CUDA code to Biren. The focus is on performance tuning and debugging techniques.
This instructor-led live training in Durban empowers developers with the necessary skills to build and deploy AI models using BANGPy and Neuware on Cambricon MLUs. Learners will configure environments, develop optimized models, and integrate MLU acceleration into edge and data center applications.
This instructor-led, live training in Durban (online or onsite) is aimed at beginner-level system administrators and IT professionals who wish to install, configure, manage, and troubleshoot CUDA environments.
By the end of this training, participants will be able to:
Understand the architecture, components, and capabilities of CUDA.
This instructor-led, live training in Durban (online or onsite) is aimed at beginner-level to intermediate-level developers who wish to use OpenCL to program heterogeneous devices and exploit their parallelism.
By the end of this training, participants will be able to:
Set up a development environment that includes OpenCL SDK, a device that supports OpenCL, and Visual Studio Code.
Create a basic OpenCL program that performs vector addition on the device and retrieves the results from device memory.
Use OpenCL API to query device information, create contexts, command queues, buffers, kernels, and events.
Use OpenCL C language to write kernels that execute on the device and manipulate data.
Use OpenCL built-in functions, extensions, and libraries to perform common tasks and operations.
Use OpenCL host and device memory models to optimize data transfers and memory accesses.
Use OpenCL execution model to control the work-items, work-groups, and ND-ranges.
Debug and test OpenCL programs using tools such as CodeXL, Intel VTune, and NVIDIA Nsight.
Optimize OpenCL programs using techniques such as vectorization, loop unrolling, local memory, and profiling.
This instructor-led live training in Durban (online or onsite) targets C++ developers who wish to use CUDA to accelerate applications, develop high-performance GPU kernels, and exploit parallel algorithm libraries for scientific computing, data processing, and machine learning workloads.
This instructor-led, live training in Durban (online or onsite) is aimed at C/C++ developers who wish to use CUDA to accelerate compute-intensive applications, including data processing, scientific simulations, machine learning workloads, and image processing pipelines.
This instructor-led, live training in Durban (online or onsite) is aimed at software developers, data analysts, and technical professionals who wish to use TensorFlow 2.x and Keras to build, train, and deploy deep learning models for computer vision, natural language processing, and multimodal applications.
This instructor-led, live training course in Durban covers how to program GPUs for parallel computing, how to use various platforms, how to work with the CUDA platform and its features, and how to perform various optimization techniques using CUDA. Some of the applications include deep learning, analytics, image processing and engineering applications.
Read more...
Last Updated:
Testimonials (1)
Trainers energy and humor.
Tadeusz Kaluba - Nokia Solutions and Networks Sp. z o.o.
Online GPU training in Durban, Graphics Processing Unit training courses in Durban, Weekend GPU courses in Durban, Evening Graphics Processing Unit (GPU) training in Durban, GPU instructor-led in Durban, Graphics Processing Unit instructor in Durban, GPU (Graphics Processing Unit) instructor-led in Durban, GPU private courses in Durban, Graphics Processing Unit (GPU) one on one training in Durban, Weekend Graphics Processing Unit (GPU) training in Durban, Graphics Processing Unit classes in Durban, GPU boot camp in Durban, GPU on-site in Durban, GPU trainer in Durban, Online Graphics Processing Unit training in Durban, Evening GPU (Graphics Processing Unit) courses in Durban, GPU coaching in Durban