Data Engineering

10 notes in this chapter

Data Engineering Terms MOC

10 terms across 4 categories, the plumbing that moves data between the systems already covered in Advanced Databases (single-database internals) and ML/DL (what happens once data reaches a model).

Pipeline Fundamentals

Processing Paradigms

Storage

Orchestration & Quality


How to use this

Read this as one connected story: data changes in a source system (Change Data Capture (CDC)), gets moved by a Data Pipeline (ETL vs ELT), processed in batch or streaming (Apache Spark, Apache Kafka), lands in a lake or warehouse (Snowflake), all orchestrated and checked by Apache Airflow and Data Quality and Validation.

Suggested order if starting from zero

  1. Data Pipeline → ETL vs ELT — the shape of the whole field
  2. Data Lake vs Data Warehouse → Snowflake — where data ends up
  3. Batch vs Stream Processing → Apache Spark → Apache Kafka — how data actually moves and transforms at scale
  4. Apache Airflow → Data Quality and Validation — how it’s kept running and trustworthy
  5. Change Data Capture (CDC) — the more advanced technique, once the fundamentals click

Dig deeper