← Spark 모듈로

Learn · Roadmap

Spark 로드맵

원본 자료(소스)별 챕터 목차 · 총 38개 모듈

0 / 38 읽음

Spark: The Definitive Guide (Excerpts, Databricks Preview, 2017) — Chapters 2-6

Bill Chambers & Matei Zaharia
1 Spark 클러스터 아키텍처와 언어 API Chapter 2: A Gentle Introduction to Spark — Spark's Basic Architecture, Spark Applications, Spark's Language APIs (pp.3-7) 2 SparkSession, DataFrame, 파티션 Chapter 2: A Gentle Introduction to Spark — Starting Spark, SparkSession, DataFrames, Partitions (pp.7-9) 3 트랜스포메이션, 지연 평가, 액션과 Spark UI Chapter 2: A Gentle Introduction to Spark — Transformations, Lazy Evaluation, Actions, Spark UI (pp.9-11) 4 End-to-End 예제로 보는 실행 계획 (항공편 데이터) Chapter 2: A Gentle Introduction to Spark — An End to End Example, DataFrames and SQL (pp.12-21) 5 Dataset과 캐싱 — 타입 안전 API와 반복 접근 최적화 Chapter 3: A Tour of Spark's Toolset - 개요, Datasets, Caching Data for Faster Access (pp.22-26) 6 Structured Streaming 첫걸음 — 배치를 스트리밍으로 Chapter 3: A Tour of Spark's Toolset - Structured Streaming (pp.26-32) 7 MLlib로 배우는 Spark 머신러닝 파이프라인 Chapter 3: A Tour of Spark's Toolset - Machine Learning and Advanced Analytics (pp.32-38) 8 Spark 패키지 생태계와 GraphFrames Chapter 3: A Tour of Spark's Toolset - Spark's Ecosystem and Packages, GraphFrames (pp.38-43) 9 Structured API의 세 가지 얼굴: DataFrame, Dataset, SQL Chapter 4: Structured API Overview - 도입부, DataFrames and Datasets, Schemas 절 (PDF pp.44-47) 10 Catalyst 타입 시스템: untyped DataFrame vs typed Dataset Chapter 4: Overview of Structured Spark Types - Catalyst, DataFrame/Dataset 타입 차이, Columns, Rows 절 (PDF pp.45-47) 11 Structured API 실행 과정: 논리 계획에서 클러스터 실행까지 Chapter 4: Overview of Structured API Execution - Logical Planning, Physical Planning, Execution 절 (PDF pp.51-53) 12 DataFrame 스키마 정의와 컬럼·표현식 Chapter 5: Basic Structured Operations - Schemas, Columns and Expressions, Records and Rows 절 (PDF pp.54-63) 13 DataFrame 컬럼 조작: select/selectExpr, 리터럴, 추가·이름변경·제거·캐스팅 Chapter 5: DataFrame Transformations - Creating DataFrames, Select & SelectExpr, Literals, Adding/Renaming/Removing Columns, Casting 절 (PDF pp.63-75) 14 행 필터링, 고유값, 샘플링, 합치기, 정렬 Chapter 5: Filtering Rows, Getting Unique Rows, Random Samples/Splits, Concatenating and Appending Rows, Sorting Rows 절 (PDF pp.75-81) 15 결과 개수 제한, 파티셔닝(repartition/coalesce), 드라이버로 결과 수집 Chapter 5: Limit, Repartition and Coalesce, Collecting Rows to the Driver 절 (PDF pp.81-84) 16 표현식 API 지도와 불리언(Boolean) 다루기 Chapter 6: Working with Different Types of Data - Chapter Overview, Where to Look for APIs, Working with Booleans 절 (PDF pp.85-91) 17 숫자 다루기: 산술 연산과 통계 함수 Chapter 6: Working with Numbers 절 (PDF pp.91-96) 18 문자열과 정규표현식 다루기 Chapter 6: Working with Strings, Regular Expressions 절 (PDF pp.96-105) 19 날짜와 타임스탬프 다루기 Chapter 6: Working with Dates and Timestamps 절 (PDF pp.105-111) 20 null 다루기와 복합 타입(구조체·배열·맵) Chapter 6: Working with Nulls in Data, Working with Complex Types 절 (PDF pp.111-119) 21 JSON 다루기와 사용자 정의 함수(UDF) Chapter 6: Working with JSON, User-Defined Functions 절 (PDF pp.119-126)

Apache Spark Data Engineering Interview Questions (blog, 50 Q&A)

(블로그, 발행처 미상)

70 Spark Interview Questions for Data Engineers (Real Asks, 2026)

(블로그, 발행처 미상)

PySpark Tutorial: Build a Real Pipeline and Read the Spark UI

Darshil Parmar
28 로컬 Spark 세션과 파티션의 실체 — 코어 개수가 아니라 파티션 개수가 병렬성을 정한다 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — Setup, SparkSession, Partitions (문서 상단부) 29 지연 평가는 실제로 Jobs 탭에 어떻게 찍히는가 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — Spark ignored your first five lines 구간, Jobs 페이지 스크린샷 30 ANSI 모드와 캐스트 에러 — 크래시와 조용한 데이터 손실 중 무엇을 고를 것인가 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "Eight rows out of 3,384 will kill the whole job" 구간 (8~10페이지) 31 대소문자 정규화와 dropDuplicates()의 함정 — 정제는 기본값이 아니라 선택이다 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "The lower() call earns its place too" ~ "dropDuplicates() misses the duplicates that matter" 구간 (11~14페이지) 32 조인 결과는 반드시 행 수로 검증한다 — cache(), left_anti join, BroadcastHashJoin 확인 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "The join is fast because 400 rows fit in memory" 구간 (14~17페이지) 33 groupBy는 셔플이다 — spark.sql.shuffle.partitions 200과 AQE의 파티션 재조정 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "One groupBy, 200 tasks, and 151 files nobody wants" 구간 (17~22페이지) 34 AQE를 껐을 때와 켰을 때 — small files problem을 숫자로 확인하기 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "The cost of losing that safety net shows up on disk" 구간 (22~24페이지) 35 partitionBy와 파티션 프루닝 — 어떤 컬럼을 파티션 키로 고를 것인가 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "Write it out, then read it back like a stranger would" 구간 (24~26페이지) 36 FAQ 보강 — PySpark 학습 경로와 ETL 적합성 판단, 로컬 vs 클러스터 PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — Frequently asked questions 구간 (27~29페이지)

Batch vs Stream Processing: The Plain-English Guide (2026)

(블로그, 발행처 미상)