Iceberg & the lakehouseTable-format engineering: schema evolution, slowly changing dimensions, snapshot expiry, compaction cadence, and repairable data history.
5 posts
PostgreSQLReplication, point-in-time recovery, promotion, connection limits, and the version-upgrade surprises that only show up in production.
6 posts
ClickHouseColumn-store analytics: MergeTree design, partition and ordering keys, materialised views, and keeping log tables from eating the disk.
1 post
Distributed databasesQuorum, sharding, failover and backup across MongoDB, ScyllaDB, Redis, MariaDB/Galera and Solr — measured on a rig rather than quoted from a docs page.
5 posts
AI agents in productionProduction agents with typed, audited outputs — and the coding-agent control plane (AGENTS.md, Cursor, Claude Code, Codex, CODEOWNERS) that keeps AI-generated PRs reviewable. LangGraph, CrewAI, Google-ADK, MCP, contracts and fallbacks.
3 posts
RAGRetrieval over corpora that keep growing: chunking, metadata filters, index choice, and the TTL and compaction policy that stops a vector store becoming a landfill.
1 post
Observability & SLOsWhat to measure when 'uptime' means nothing for a data platform: freshness SLIs, burn-rate alerts, lineage, and alerts that should actually page someone.
1 post
Data quality & contractsValidation, reconciliation, quarantine, tolerant readers, and contracts that keep imperfect producers from silently corrupting downstream decisions.
3 posts
Career & craftNotes on ownership, mentoring, review culture, and the difference between contributing to a system and being on call for it.
1 post
Self-hosted opsFree, fully OSS setups you can run yourself: a WireGuard VPN, a RustDesk remote desktop, and a Stalwart mailbox for your own domain.
5 posts
Writing
Browse by system, failure mode, or layer.
Published field notes organized by system, failure mode, and platform layer — from distributed databases and PostgreSQL to streaming, Iceberg, agents, and SLOs.