Software Engineer & Data Platform Architect

I build data platforms as software.

Batch and streaming pipelines, lakehouses and warehouses, the storage and database layers beneath them, and the frameworks, services and infrastructure around them.

every diagram here is a sketch like this ↘︎

Sketch of a data platform: sources flow through replayable ingestion into a versioned lakehouse and out through serving to analysts, apps and teams, with a control plane tracking state, lineage and retries.

§ 01 — How I work

The engineering underneath

01

Platforms other engineers build on

Frameworks, compilers and shared libraries that turn repeated pipeline work into configuration, released like a product.

02

Data that stays correct when things fail

Replayable stages, idempotent writes, automatic recovery, and releases that switch atomically and roll back in one step.

03

From engine internals to architecture

How a query engine plans a join, how a table format commits, how a log replicates: knowing both ends keeps designs as simple as the problem allows.

§ 02 — Selected systems

Core-team work

Full résumé →

§ 03 — Blog

How the machinery works

All posts →
In progress · Spark Jobs, stages and tasks How an action becomes a job, how the scheduler cuts it into stages at shuffle boundaries, and how tasks find executor cores, locality and retries included.
In progress · Spark Storage, caching and memory What an executor's memory holds, how execution and cached data share one pool, and why a container gets killed while its heap looks half empty.
In progress · Spark Shuffle, partitioning and joins How a shuffle moves rows between stages, how many partitions to use, and how Spark chooses and runs joins, with AQE splitting the skewed ones.

§ 04 — Contact

Let's talk tech.

distributed systems, architecture, engine internals