Blog

How the machinery works

Long-form notes on data systems: Spark, table formats, streaming, databases and the platforms around them.

Every post is written by me and goes as deep as I wanted it to: down to the source code where it matters. It comes from years of building and running these systems, and from the curiosity to keep asking how they work inside. I use AI to check facts against the source and to help with the prose and the diagrams.

Sketch of the topics: Kafka feeding Spark, Spark committing to Iceberg, databases, and the Kubernetes and cloud platform underneath.

the first posts are being written ↓

In progressSpark Jobs, stages and tasks How an action becomes a job, how the scheduler cuts it into stages at shuffle boundaries, and how tasks find executor cores, locality and retries included.
In progressSpark Storage, caching and memory What an executor's memory holds, how execution and cached data share one pool, and why a container gets killed while its heap looks half empty.
In progressSpark Shuffle, partitioning and joins How a shuffle moves rows between stages, how many partitions to use, and how Spark chooses and runs joins, with AQE splitting the skewed ones.