Data Engineering Roadmap

Start at The shape of the problem and follow the order. Every stage names what it needs first and what you should be able to do before moving on, and the last one is where a number becomes explainable. Progress is stored locally in your browser.

Where to start

Data engineering

10 stages · 0/243 lessons

How data gets from an operational system to an analytical consumer while staying complete, correct, fresh, explainable and affordable.

  1. The shape of the problem
  2. Where data lives
  3. Getting data in
  4. Modelling for questions
  5. Transformation and orchestration
  6. Physical layout
  7. Change and streams
  8. Distributed compute
  9. Trust
  10. Production data engineering
0 / 243 lessons masteredNot started 243Learning 0Practicing 0Mastered 0
  1. 1

    The shape of the problem

    Start here
    0/12

    What data engineering is once the tools are removed: the journey from an application write to a number on a dashboard, and who is waiting at the end of it. OLTP and OLAP are two workloads with opposite shapes, which is why analytics moved off the production database, and ETL versus ELT is a question about where compute lives and how much raw history you keep. Everything later assumes you can name the consumer and the source of truth first.

    Before moving on: Explain why analytics moved off the production database, name the consumer of a given dataset, and answer an analytical question without taking anything down.

  2. 2

    Where data lives

    0/19

    The bytes underneath analytics. Row versus column storage explains why a warehouse scans fast; Parquet, Avro and ORC show what a format stores and what that lets a reader skip; compression and encoding follow from the layout. Then the places those files live — lake, warehouse, lakehouse, open table formats, object storage — compared on what each one costs rather than on marketing. Physical layout later builds directly on this.

    Before moving on: Choose between a lake, a warehouse and a lakehouse for a real workload, say what each one costs you, and explain what a columnar reader can skip that a row reader cannot.

  3. 3

    Getting data in

    0/12

    Getting data out of systems you usually do not control. Batch extracts, incremental windows on a predicate, streaming producers, and the failure recovery that decides whether a missed hour is recoverable or gone. The raw landing zone and the layered (medallion) structure come here because where data lands first decides what you can reprocess after you find a bug. This is the stage before modelling because a model is built on what actually arrived.

    Before moving on: Build an ingestion path that survives a failure without losing or duplicating a day, and explain the predicate it extracts on and where the raw copy is kept.

  4. 4

    Modelling for questions

    0/12

    Facts, dimensions, grain and history. The model decides which business questions are easy, which are expensive, and which are answerable but silently wrong. Star and snowflake schemas, surrogate keys, slowly changing dimensions and snapshot tables are the vocabulary; grain is the one field that decides whether an aggregate over your table means anything. It sits after ingestion because the model is built from the raw layer, and before transformation because transformation is how the model gets built.

    Before moving on: Design a model whose grain is declared, whose history is preserved where it needs to be, and whose measures do not multiply when joined.

  5. 5

    Transformation and orchestration

    0/16

    Turning raw data into the model as a dependency graph of tested, documented SQL rather than a pile of scheduled scripts: dbt-style layering, a metrics layer, topological execution. Then the orchestrator that runs the graph — why a scheduler is not one, what a failed task in the middle of a DAG means, and why idempotency and a high-water mark are the properties that make re-running safe. Every later stage that says "re-run it" assumes this one.

    Before moving on: Express a platform as a tested dependency graph, and re-run any part of it without corrupting what is already correct.

  6. 6

    Physical layout

    0/15

    Where bytes physically sit decides what a query must read. File size and compaction, partitioning and pruning, cardinality, clustering and bucketing are the highest-leverage and least-visible decisions in analytics. The query-engine half — coordinators and workers, predicate and projection pushdown, vectorized execution — is here because pushdown is what turns a layout into skipped bytes. It builds on the formats and object storage from Where data lives.

    Before moving on: Make a query read a fraction of what it used to, and explain exactly which files, row groups and columns the engine skipped and why.

    Needs first:Where data lives
  7. 7

    Change and streams

    0/30

    The continuous path. Change data capture reads a database's own log instead of polling it, and every way it silently loses or reorders changes. The durable, partitioned, replayable log — topics, keys, consumer groups, offsets, retention — is the infrastructure that carries those changes. Stream processing on top adds event time versus processing time, windows, watermarks, state and joins, and what "exactly-once" can and cannot mean for processing. It follows ingestion because streaming is the other half of batch versus streaming.

    Before moving on: Capture changes from a database log, publish them to a replayable stream, and reason correctly about event time, lateness and what a window will and will not contain.

    Needs first:Getting data in
  8. 8

    Distributed compute

    0/15

    Spark and its relatives from the inside: partitions, stages, tasks and the shuffle. Narrow versus wide transformations, skew, stragglers, salting and broadcast joins explain why one task in a thousand decides a job's runtime. Lazy evaluation and query optimizers connect back to the pushdown you saw in Physical layout; Flink, federated query and batch-streaming unification connect forward from Change and streams. This is where "just add more workers" stops being an answer.

    Before moving on: Read an execution plan, find the shuffle, spot the skew, and say why adding workers will not help this job.

  9. 9

    Trust

    0/34

    How anyone knows the data is correct enough to use. Quality dimensions, tests, distribution checks, freshness and reconciliation, and what every check still misses. Contracts and schema evolution say which changes are safe and which break silently; catalog and lineage make a dataset findable and its blast radius knowable; observability turns "revenue looks wrong" into an upstream walk to a cause. It comes after modelling and transformation because tests live in the DAG and a semantic change is only visible against a declared grain.

    Before moving on: Say whether a number is trustworthy and prove it, name what each check on it still misses, and find out that it is wrong before a finance team does.

  10. 10

    Production data engineering

    0/78

    Running the platform rather than building it. Backfills and reprocessing fix history without breaking the present; reliability adds retries, checkpoints, atomic publish and SLOs; governance covers PII, retention, access and deletion; cost engineering names the drivers — bytes scanned, shuffled, retained, work repeated — and the decisions that move each one. Architecture patterns, the cloud services, the AI-and-agents data products and the debugging lessons close the loop by walking a wrong number back to its source. Everything here re-runs, replays or reconciles something taught earlier.

    Before moving on: Run a platform: backfill a range without duplicating it, govern and pay for it, and walk any number on it back to its source.