Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Foundations (Theoretical):

  • System Architecture
  • Resilient Distributed Datasets (RDDs)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Databricks Environment (Hands-on Workshop):

  • Practical exercises utilizing the RDD API
  • Core action and transformation functions
  • PairRDD concepts
  • Join operations
  • Caching strategies
  • Practical exercises utilizing the DataFrame API
  • SparkSQL implementation
  • DataFrame operations: select, filter, group, and sort
  • UDFs (User Defined Functions)
  • Exploring the Dataset API
  • Stream processing

AWS Environment (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Key differences between AWS EMR and AWS Glue
  • Practical job implementations in both environments
  • Evaluation of advantages and limitations

Additional Topics:

  • Introduction to orchestration with Apache Airflow

Requirements

Programming experience (ideally in Python and Scala)

Fundamental knowledge of SQL

 21 Hours

Testimonials (3)

Upcoming Courses

Related Categories