Главная
Study mode:
on
1
Intro
2
Session Goals
3
File Formats
4
Row-wise Storage
5
Columnar (Column-wise) Storage
6
Hybrid Storage
7
Example Data
8
About: CSV
9
About: JSON
10
About: Avro
11
Inspecting: Avro
12
About: ORC
13
Structure: ORC
14
Inspecting: ORC
15
Config: ORC
16
Structure: Parquet
17
Inspecting: Parquet (1)
18
Inspecting: Parquet (2)
19
Config: Parquet
20
Case Study: Veraset
21
Looking Forward: Apache Arrow
22
Final Thoughts
Description:
Explore the Apache Spark file format ecosystem in this 29-minute video from Databricks. Dive into the critical role of storage and IO in Spark job performance and optimization. Learn about key file formats like Parquet, ORC, and Avro, their origins, core data structures, and optimization techniques. Discover how to tune SparkConf and SQLConf settings for these formats. Examine real-world industry examples showcasing file format impact on job performance and stability. Gain insights into emerging technologies like Apache Arrow. By the end, understand core concepts of prevalent file formats, relevant settings, and how to select the right format for your Spark jobs, especially for IO-intensive AI/ML workflows.

The Apache Spark File Format Ecosystem - Optimizing Storage for Performance

Databricks
Add to list
0:00 / 0:00