Data & Analytics
Intermediate

Data Engineering

Architect, build, and optimize scalable data platforms, distributed pipelines, and streaming engines.

Master foundational distributed computing principles, relational modeling, and real-time streaming architectures. Transition into production engineering with modern orchestrators, analytical data warehouses, and data lakehouse implementations. Build fault-tolerant, high-throughput data infrastructure capable of processing petabyte-scale workloads.

Dr. Elena Vance, Associate Professor of Distributed Systems & Former Staff Data Architect 90h + 45h 2 certificates available
Data EngineeringApache SparkDistributed SystemsApache KafkaData Modeling

Curriculum

Two complete tracks. Study either or both — each has its own exam and certificate.

A mathematically and algorithmically rigorous grounding in distributed systems theory, relational algebra, and stateful stream processing models.

Theoretical examination of low-level database storage engines, distributed consensus algorithms, and replication mechanics.

  • Log-Structured Merge-Trees and B-Tree Storage Engines Lab40 min
  • Distributed Consensus: Raft and Paxos Protocols Lab45 min
  • Replication Topologies and Conflict Resolution Lab35 min

Careers this prepares you for

  • Data Engineer
  • Distributed Systems Engineer
  • Analytics Engineer
  • Big Data Platform Architect

Your tutor

DE
Dr. Elena Vance
Associate Professor of Distributed Systems & Former Staff Data Architect