This intensive three-day practical course is dedicated to constructing and refining high-efficiency data processing workloads utilising PySpark, Pandas, and Polars within Kubernetes-based environments.
Participants will cultivate a deep practical understanding of Spark application execution on Kubernetes, exploring how configuration decisions at the application level directly impact performance, scalability, resource usage, and operational costs. The curriculum addresses critical optimization domains such as executor sizing, memory allocation, dynamic allocation, partitioning strategies, shuffle mechanics, the small-file challenge, and the efficient handling of Parquet files.
The course also navigates common pitfalls associated with Pandas, such as memory constraints and out-of-memory errors, while introducing Polars as a high-performance alternative for specific data processing tasks. Through hands-on labs, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration strategies, and implement optimization techniques in realistic ETL and machine learning contexts.
The core focus of the course is on practical decision-making: mastering the ability to pinpoint bottlenecks, select the most appropriate tools, configure Spark effectively, and balance performance against infrastructure resource consumption and cost.
Read more...