← Spark 모듈로
Learn · Roadmap
Spark 로드맵
원본 자료(소스)별 챕터 목차 · 총 38개 모듈
0 / 38 읽음
용어 사전 보기
문제 풀어보기
Spark: The Definitive Guide (Excerpts, Databricks Preview, 2017) — Chapters 2-6
Bill Chambers & Matei Zaharia
1
Spark 클러스터 아키텍처와 언어 API
Chapter 2: A Gentle Introduction to Spark — Spark's Basic Architecture, Spark Applications, Spark's Language APIs (pp.3-7)
2
SparkSession, DataFrame, 파티션
Chapter 2: A Gentle Introduction to Spark — Starting Spark, SparkSession, DataFrames, Partitions (pp.7-9)
3
트랜스포메이션, 지연 평가, 액션과 Spark UI
Chapter 2: A Gentle Introduction to Spark — Transformations, Lazy Evaluation, Actions, Spark UI (pp.9-11)
4
End-to-End 예제로 보는 실행 계획 (항공편 데이터)
Chapter 2: A Gentle Introduction to Spark — An End to End Example, DataFrames and SQL (pp.12-21)
5
Dataset과 캐싱 — 타입 안전 API와 반복 접근 최적화
Chapter 3: A Tour of Spark's Toolset - 개요, Datasets, Caching Data for Faster Access (pp.22-26)
6
Structured Streaming 첫걸음 — 배치를 스트리밍으로
Chapter 3: A Tour of Spark's Toolset - Structured Streaming (pp.26-32)
7
MLlib로 배우는 Spark 머신러닝 파이프라인
Chapter 3: A Tour of Spark's Toolset - Machine Learning and Advanced Analytics (pp.32-38)
8
Spark 패키지 생태계와 GraphFrames
Chapter 3: A Tour of Spark's Toolset - Spark's Ecosystem and Packages, GraphFrames (pp.38-43)
9
Structured API의 세 가지 얼굴: DataFrame, Dataset, SQL
Chapter 4: Structured API Overview - 도입부, DataFrames and Datasets, Schemas 절 (PDF pp.44-47)
10
Catalyst 타입 시스템: untyped DataFrame vs typed Dataset
Chapter 4: Overview of Structured Spark Types - Catalyst, DataFrame/Dataset 타입 차이, Columns, Rows 절 (PDF pp.45-47)
11
Structured API 실행 과정: 논리 계획에서 클러스터 실행까지
Chapter 4: Overview of Structured API Execution - Logical Planning, Physical Planning, Execution 절 (PDF pp.51-53)
12
DataFrame 스키마 정의와 컬럼·표현식
Chapter 5: Basic Structured Operations - Schemas, Columns and Expressions, Records and Rows 절 (PDF pp.54-63)
13
DataFrame 컬럼 조작: select/selectExpr, 리터럴, 추가·이름변경·제거·캐스팅
Chapter 5: DataFrame Transformations - Creating DataFrames, Select & SelectExpr, Literals, Adding/Renaming/Removing Columns, Casting 절 (PDF pp.63-75)
14
행 필터링, 고유값, 샘플링, 합치기, 정렬
Chapter 5: Filtering Rows, Getting Unique Rows, Random Samples/Splits, Concatenating and Appending Rows, Sorting Rows 절 (PDF pp.75-81)
15
결과 개수 제한, 파티셔닝(repartition/coalesce), 드라이버로 결과 수집
Chapter 5: Limit, Repartition and Coalesce, Collecting Rows to the Driver 절 (PDF pp.81-84)
16
표현식 API 지도와 불리언(Boolean) 다루기
Chapter 6: Working with Different Types of Data - Chapter Overview, Where to Look for APIs, Working with Booleans 절 (PDF pp.85-91)
17
숫자 다루기: 산술 연산과 통계 함수
Chapter 6: Working with Numbers 절 (PDF pp.91-96)
18
문자열과 정규표현식 다루기
Chapter 6: Working with Strings, Regular Expressions 절 (PDF pp.96-105)
19
날짜와 타임스탬프 다루기
Chapter 6: Working with Dates and Timestamps 절 (PDF pp.105-111)
20
null 다루기와 복합 타입(구조체·배열·맵)
Chapter 6: Working with Nulls in Data, Working with Complex Types 절 (PDF pp.111-119)
21
JSON 다루기와 사용자 정의 함수(UDF)
Chapter 6: Working with JSON, User-Defined Functions 절 (PDF pp.119-126)
Apache Spark Data Engineering Interview Questions (blog, 50 Q&A)
(블로그, 발행처 미상)
22
Spark 핵심 아키텍처 (Driver/Executor/DAG/RDD)
spark_ref_01.md Q1-10 (Core Concepts & Architecture)
23
RDD·DataFrame·Dataset와 데이터 처리
spark_ref_01.md Q11-20 (Data Structures & APIs)
24
성능 튜닝과 최적화 (스큐·셔플·Catalyst·Tungsten)
spark_ref_01.md Q21-30 (Performance Tuning & Optimization) + FAQ(파티셔닝/브로드캐스트 조인/성능 최적화/cache·persist/데이터 스큐) 종합, 수치 오류 정정 포함
25
운영과 배포 (클러스터 매니저·장애 복구·직렬화)
spark_ref_01.md Q31-40 (Operations & Deployment)
26
실전 코딩 시나리오 (조인·UDF·파티셔닝)
spark_ref_01.md Q41-50 (Advanced Scenarios & Coding Logic)
70 Spark Interview Questions for Data Engineers (Real Asks, 2026)
(블로그, 발행처 미상)
27
70개 질문 블로그로 보는 추가 확인 사항 — 파티셔너와 튜닝 체크리스트
70 Spark Interview Questions for Data Engineers (전체 50문항 + FAQ)
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI
Darshil Parmar
28
로컬 Spark 세션과 파티션의 실체 — 코어 개수가 아니라 파티션 개수가 병렬성을 정한다
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — Setup, SparkSession, Partitions (문서 상단부)
29
지연 평가는 실제로 Jobs 탭에 어떻게 찍히는가
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — Spark ignored your first five lines 구간, Jobs 페이지 스크린샷
30
ANSI 모드와 캐스트 에러 — 크래시와 조용한 데이터 손실 중 무엇을 고를 것인가
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "Eight rows out of 3,384 will kill the whole job" 구간 (8~10페이지)
31
대소문자 정규화와 dropDuplicates()의 함정 — 정제는 기본값이 아니라 선택이다
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "The lower() call earns its place too" ~ "dropDuplicates() misses the duplicates that matter" 구간 (11~14페이지)
32
조인 결과는 반드시 행 수로 검증한다 — cache(), left_anti join, BroadcastHashJoin 확인
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "The join is fast because 400 rows fit in memory" 구간 (14~17페이지)
33
groupBy는 셔플이다 — spark.sql.shuffle.partitions 200과 AQE의 파티션 재조정
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "One groupBy, 200 tasks, and 151 files nobody wants" 구간 (17~22페이지)
34
AQE를 껐을 때와 켰을 때 — small files problem을 숫자로 확인하기
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "The cost of losing that safety net shows up on disk" 구간 (22~24페이지)
35
partitionBy와 파티션 프루닝 — 어떤 컬럼을 파티션 키로 고를 것인가
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — "Write it out, then read it back like a stranger would" 구간 (24~26페이지)
36
FAQ 보강 — PySpark 학습 경로와 ETL 적합성 판단, 로컬 vs 클러스터
PySpark Tutorial: Build a Real Pipeline and Read the Spark UI — Frequently asked questions 구간 (27~29페이지)
Batch vs Stream Processing: The Plain-English Guide (2026)
(블로그, 발행처 미상)
37
배치와 스트림, 하나의 엔진으로 — Spark 관점에서 재구성한 처리 모델 비교
Batch vs Stream Processing: The Plain-English Guide (2026) — 전체(1~7페이지, assumption/time semantics/computation model/state 구간)를 Spark 아키텍처 관점에서 재구성
38
언제 Spark 배치, 언제 Structured Streaming인가 — 실무 판단 체크리스트
Batch vs Stream Processing: The Plain-English Guide (2026) — freshness/failure handling/cost/output nature 및 When to Use 체크리스트 구간(7~10페이지)을 Spark 파이프라인 맥락으로 재구성