Data Engineering Roadmap
Start at The shape of the problem and follow the order. Every stage names what it needs first and what you should be able to do before moving on, and the last one is where a number becomes explainable. Progress is stored locally in your browser.
Where to start
Data engineering
10 stages · 0/243 lessonsHow data gets from an operational system to an analytical consumer while staying complete, correct, fresh, explainable and affordable.
- 10/12
The shape of the problem
Start hereWhat data engineering is once the tools are removed: the journey from an application write to a number on a dashboard, and who is waiting at the end of it. OLTP and OLAP are two workloads with opposite shapes, which is why analytics moved off the production database, and ETL versus ELT is a question about where compute lives and how much raw history you keep. Everything later assumes you can name the consumer and the source of truth first.
Before moving on: Explain why analytics moved off the production database, name the consumer of a given dataset, and answer an analytical question without taking anything down.
- What Data Engineering Actually Is
- The Fundamental Data Journey
- OLTP Workloads
- OLAP Workloads
- OLTP vs OLAP
- Workload Isolation
- The OLTP to OLAP Journey
- ETL: Transform Before the Data Lands
- ELT: Load First, Transform Where the Data Lives
- ETL vs ELT: Choosing by Constraint, Not by Fashion
- Who Actually Consumes This Data
- Source of Truth
- 20/19
Where data lives
The bytes underneath analytics. Row versus column storage explains why a warehouse scans fast; Parquet, Avro and ORC show what a format stores and what that lets a reader skip; compression and encoding follow from the layout. Then the places those files live — lake, warehouse, lakehouse, open table formats, object storage — compared on what each one costs rather than on marketing. Physical layout later builds directly on this.
Before moving on: Choose between a lake, a warehouse and a lakehouse for a real workload, say what each one costs you, and explain what a columnar reader can skip that a row reader cannot.
Needs first:The shape of the problem- Row vs Column Storage
- Columnar Execution
- Why Analytical Data Compresses
- Dictionary, Run-Length, Delta and Bit Packing
- Parquet
- Parquet Internals
- The Parquet Read Path
- Avro
- ORC
- Parquet vs Avro
- CSV, JSON and Their Limits
- The Data Lake
- The Data Warehouse
- The Lakehouse
- Open Table Formats
- Lake vs Warehouse vs Lakehouse
- Object Storage as Data Infrastructure
- Separating Storage from Compute
- Data Marts
- 30/12
Getting data in
Getting data out of systems you usually do not control. Batch extracts, incremental windows on a predicate, streaming producers, and the failure recovery that decides whether a missed hour is recoverable or gone. The raw landing zone and the layered (medallion) structure come here because where data lands first decides what you can reprocess after you find a bug. This is the stage before modelling because a model is built on what actually arrived.
Before moving on: Build an ingestion path that survives a failure without losing or duplicating a day, and explain the predicate it extracts on and where the raw copy is kept.
- Data Ingestion
- Ingestion Sources
- Batch Ingestion
- Incremental Extraction
- Streaming Ingestion
- Batch vs Streaming Ingestion
- Ingestion Failure & Recovery
- The Raw Landing Zone
- Raw, Staging, Curated: Layers by Purpose
- Medallion: One Naming Convention Among Several
- Keeping Raw History: The Recovery Position and the Liability
- Where the Transformation Actually Runs
- 40/12
Modelling for questions
Facts, dimensions, grain and history. The model decides which business questions are easy, which are expensive, and which are answerable but silently wrong. Star and snowflake schemas, surrogate keys, slowly changing dimensions and snapshot tables are the vocabulary; grain is the one field that decides whether an aggregate over your table means anything. It sits after ingestion because the model is built from the raw layer, and before transformation because transformation is how the model gets built.
Before moving on: Design a model whose grain is declared, whose history is preserved where it needs to be, and whose measures do not multiply when joined.
- 50/16
Transformation and orchestration
Turning raw data into the model as a dependency graph of tested, documented SQL rather than a pile of scheduled scripts: dbt-style layering, a metrics layer, topological execution. Then the orchestrator that runs the graph — why a scheduler is not one, what a failed task in the middle of a DAG means, and why idempotency and a high-water mark are the properties that make re-running safe. Every later stage that says "re-run it" assumes this one.
Before moving on: Express a platform as a tested dependency graph, and re-run any part of it without corrupting what is already correct.
- Data Transformation
- SQL Transformations
- dbt Concepts
- The Transformation DAG
- DAGs in Data Pipelines
- Topological Execution
- Model Layering
- The Metrics Layer
- Orchestration
- Scheduler vs Orchestrator
- Airflow Concepts
- Task Dependencies
- When a Task Fails Mid-DAG
- Idempotent Data Pipelines
- Incremental Processing
- The High-Water Mark
- 60/15
Physical layout
Where bytes physically sit decides what a query must read. File size and compaction, partitioning and pruning, cardinality, clustering and bucketing are the highest-leverage and least-visible decisions in analytics. The query-engine half — coordinators and workers, predicate and projection pushdown, vectorized execution — is here because pushdown is what turns a layout into skipped bytes. It builds on the formats and object storage from Where data lives.
Before moving on: Make a query read a fraction of what it used to, and explain exactly which files, row groups and columns the engine skipped and why.
Needs first:Where data lives - 70/30
Change and streams
The continuous path. Change data capture reads a database's own log instead of polling it, and every way it silently loses or reorders changes. The durable, partitioned, replayable log — topics, keys, consumer groups, offsets, retention — is the infrastructure that carries those changes. Stream processing on top adds event time versus processing time, windows, watermarks, state and joins, and what "exactly-once" can and cannot mean for processing. It follows ingestion because streaming is the other half of batch versus streaming.
Before moving on: Capture changes from a database log, publish them to a replayable stream, and reason correctly about event time, lateness and what a window will and will not contain.
Needs first:Getting data in- Change Data Capture
- CDC vs Polling
- What a CDC Event Contains
- CDC Ordering and Transaction Boundaries
- Snapshot and Stream: the Bootstrap Problem
- CDC Failure Modes and the Retention Deadline
- CDC and Schema Drift
- The Event Log
- Message Brokers: Log-Shaped and Queue-Shaped
- Kafka as a Log, Not a Queue
- Topics and Partitions
- Event Keys and Partition Assignment
- Consumer Groups and the Parallelism Ceiling
- Retention and Replay
- Offsets and Commits
- Stream Processing
- Stateless Stream Processing
- Stateful Stream Processing
- Streaming State
- Event Time
- Processing Time
- Ingestion Time
- Late Events
- Windows
- Tumbling Windows
- Sliding Windows
- Session Windows
- Watermarks
- Stream Joins
- Exactly-Once: Input Consumption, State Update, Output Write
- 80/15
Distributed compute
Spark and its relatives from the inside: partitions, stages, tasks and the shuffle. Narrow versus wide transformations, skew, stragglers, salting and broadcast joins explain why one task in a thousand decides a job's runtime. Lazy evaluation and query optimizers connect back to the pushdown you saw in Physical layout; Flink, federated query and batch-streaming unification connect forward from Change and streams. This is where "just add more workers" stops being an answer.
Before moving on: Read an execution plan, find the shuffle, spot the skew, and say why adding workers will not help this job.
- 90/34
Trust
How anyone knows the data is correct enough to use. Quality dimensions, tests, distribution checks, freshness and reconciliation, and what every check still misses. Contracts and schema evolution say which changes are safe and which break silently; catalog and lineage make a dataset findable and its blast radius knowable; observability turns "revenue looks wrong" into an upstream walk to a cause. It comes after modelling and transformation because tests live in the DAG and a semantic change is only visible against a declared grain.
Before moving on: Say whether a number is trustworthy and prove it, name what each check on it still misses, and find out that it is wrong before a finance team does.
- Data Quality
- The Dimensions of Data Quality
- Data Tests
- Distribution Tests
- Freshness Checks
- Reconciliation
- Quality Alerting
- The Data Quality Dashboard
- Who Owns Data Quality
- Data Contracts
- Schema Evolution
- Backward Compatibility
- Forward Compatibility
- Schema Registry
- Breaking Schema Changes
- Semantic Changes
- Nullability & Defaults
- Contract Enforcement
- Metadata: Technical, Operational and Business
- The Data Catalog
- Data Lineage
- Column-Level Lineage
- Impact Analysis
- Data Ownership
- Data Discovery
- Dataset Documentation
- Data Observability
- Pipeline Observability
- Pipeline Metrics
- Freshness Monitoring
- Volume Anomalies
- Data Incidents
- Debugging a Data Incident
- Lineage Debugging
- 100/78
Production data engineering
Running the platform rather than building it. Backfills and reprocessing fix history without breaking the present; reliability adds retries, checkpoints, atomic publish and SLOs; governance covers PII, retention, access and deletion; cost engineering names the drivers — bytes scanned, shuffled, retained, work repeated — and the decisions that move each one. Architecture patterns, the cloud services, the AI-and-agents data products and the debugging lessons close the loop by walking a wrong number back to its source. Everything here re-runs, replays or reconciles something taught earlier.
Before moving on: Run a platform: backfill a range without duplicating it, govern and pay for it, and walk any number on it back to its source.
- Backfills
- What Backfills Break
- Planning a Backfill
- Validating a Backfill Before You Publish
- Reprocessing vs Retrying
- Late-Arriving Data
- Deduplication
- Upserts and Merges
- Full Refresh vs Incremental
- Replay from the Log
- Pipeline Reliability
- Atomic Publish
- Checkpointing
- Partial Failure
- Retries in Pipelines
- Pipeline SLOs
- The Freshness SLO
- Rolling Back Data
- Data Governance
- Data Classification
- PII in Pipelines
- Data Minimization
- Data Retention
- Data Access Control
- Row and Column Security
- Data Masking, Tokenisation & Encryption
- Deletion Requests
- What Actually Drives Data Platform Cost
- Scan Cost
- Compute Waste
- Storage Lifecycle
- Cost Attribution
- Cost vs Freshness
- Data Architecture Patterns
- The Central Warehouse
- The Event-Driven Data Platform
- Lambda Architecture
- Kappa Architecture
- Data Mesh
- Data Products
- Data Platform Engineering
- The Self-Service Data Platform
- Cloud Data Services
- Comparing Analytical Warehouses
- BigQuery Concepts
- Snowflake Concepts
- ClickHouse Concepts
- DuckDB Concepts
- Managed Streaming Platforms
- Choosing an Analytical Platform
- Data Engineering for Agents
- The LLM Data Pipeline
- Chunking Pipelines
- Embedding Pipelines
- Re-embedding
- Vector Data Engineering
- Evaluation Data Pipelines
- Agent Observability Data
- Feature Pipelines
- Where Did This Number Come From?
- Two Dashboards, Two Numbers
- Missing Rows
- Duplicate Rows
- Stale Dashboards
- The Pipeline Succeeded. The Data Is Wrong.
- Data Engineering Anti-Patterns
- Data Platform Anti-Patterns
- Data Engineering and Database Engineering
- Data Engineering and Distributed Systems
- Data Engineering and Backend Engineering
- Data Engineering and Cloud Infrastructure
- Data Engineering and DevOps
- Data Engineering and Observability
- Data Engineering and Security
- Data Engineering vs Its Neighbours
- The Data Loop
- Trusting Data
- What Goes Wrong Between Source and Dashboard