Topic

Observability & SLOs

What to measure when 'uptime' means nothing for a data platform: freshness SLIs, burn-rate alerts, lineage, and alerts that should actually page someone.

Availability is a weak first signal for a data product: every service can be green while a table serves yesterday’s answer. I prefer consumer-facing signals such as freshness, completeness, decision latency, and safe replay, each with a named owner. These notes cover what deserves to wake a person, how to connect an alert to a repair action, and why deleting a noisy alert can be a reliability improvement.

All writing

Tutorial The lakehouse spine Measured on my own hardware

Iceberg maintenance: 20 files into 1, and a rollback I actually performed

Measure what compaction buys on a small-files partition, then delete 909,968 rows and get them back — and find where the safety net stops.

  • Iceberg & the lakehouse
  • Observability & SLOs

~25 min, and the rollback itself takes under a second