Taming Big Data with Apache Spark and Python
Is this course right for you?
Frank Kane covers distributed-computing fundamentals, RDDs and the DataFrame API, Spark SQL, MLlib for distributed ML and Spark Streaming, running Spark locally and on AWS EMR, and tackles the common performance issues — partition tuning, join strategies and data skew.
What it does not do is teach Python, and it will not help if you only work with small datasets that pandas handles fine — it is Spark-specific rather than a broad data-engineering course, so start elsewhere in that case. Kane keeps it updated; on price, wait for a Udemy sale.","Tune Spark jobs to avoid data skew and slow joins"}
Compare alternatives for Taming Big Data with Apache Spark and Python
- Price
- PaidPaid, frequently discounted
- Duration
- 7 hrs
- Level
- Intermediate
- Certificate
- Course Certificate
- Price
- PaidDataCamp subscription
- Duration
- 4 hrs
- Level
- Beginner
- Certificate
- Course Certificate
- Price
- PaidPaid, frequently discounted
- Duration
- 10 hrs
- Level
- Intermediate
- Certificate
- Course Certificate
- Price
- PaidPaid, frequently discounted · lifetime access (+ small API costs)
- Duration
- 33.5 hrs
- Level
- Intermediate
- Certificate
- Course Certificate
About this course
This course covers Apache Spark with PySpark from distributed computing fundamentals through production patterns: resilient distributed datasets, the DataFrame API, Spark SQL for analytical queries, MLlib for distributed machine learning, and Spark Streaming for real-time data processing.
Instructor
Taught by Frank Kane, former Amazon engineer with extensive distributed systems experience who specializes in practical big data education.