Writing

Browse by system, failure mode, or layer.

Published field notes organized by system, failure mode, and platform layer — from distributed databases and PostgreSQL to streaming, Iceberg, agents, and SLOs.

Iceberg & the lakehouseTable-format engineering: schema evolution, slowly changing dimensions, snapshot expiry, compaction cadence, and repairable data history. 5 posts PostgreSQLReplication, point-in-time recovery, promotion, connection limits, and the version-upgrade surprises that only show up in production. 6 posts ClickHouseColumn-store analytics: MergeTree design, partition and ordering keys, materialised views, and keeping log tables from eating the disk. 1 post Distributed databasesQuorum, sharding, failover and backup across MongoDB, ScyllaDB, Redis, MariaDB/Galera and Solr — measured on a rig rather than quoted from a docs page. 5 posts AI agents in productionProduction agents with typed, audited outputs — and the coding-agent control plane (AGENTS.md, Cursor, Claude Code, Codex, CODEOWNERS) that keeps AI-generated PRs reviewable. LangGraph, CrewAI, Google-ADK, MCP, contracts and fallbacks. 3 posts RAGRetrieval over corpora that keep growing: chunking, metadata filters, index choice, and the TTL and compaction policy that stops a vector store becoming a landfill. 1 post Observability & SLOsWhat to measure when 'uptime' means nothing for a data platform: freshness SLIs, burn-rate alerts, lineage, and alerts that should actually page someone. 1 post Data quality & contractsValidation, reconciliation, quarantine, tolerant readers, and contracts that keep imperfect producers from silently corrupting downstream decisions. 3 posts Career & craftNotes on ownership, mentoring, review culture, and the difference between contributing to a system and being on call for it. 1 post Self-hosted opsFree, fully OSS setups you can run yourself: a WireGuard VPN, a RustDesk remote desktop, and a Stalwart mailbox for your own domain. 5 posts

Everything is also available as RSS or JSON Feed.