A notebook on backend and systems engineering.
How every piece is built
- 01Context
- 02Constraint
- 03Approach
- 04Pitfalls
- 05Expiry
Latest40
Notes, systems write-ups and journal — one stream, newest first.
- Note4 sections
Changing Schema in Production With Zero Downtime
Running migrations without causing downtime in production: backward-compatible steps, multi-phase column changes, and avoiding table locks
- Note7 sections
Read/Write Splitting: Separating Read and Write Load
Before scaling up the database: moving read traffic to replicas, the traps replication lag creates, and when you actually need it
- SponsoreddevrazziSkip the noise, read the signalCurated tech news and synthesis for developers. The content is produced by AI agents, and every piece links back to its source.
- Note6 sections
Before You Reach for Redis: The Right Cache Strategy
The questions to ask before adding a cache, why cache invalidation is hard, and the choice between TTL and consistency
- Journal5 sections
The NoSQL Trap: Starting Because "It Has No Schema"
The consistency and data-integrity crises that projects which moved to NoSQL for the comfort of going schemaless eventually hit
- Journal4 sections
Is PostgreSQL Enough for Everything?
Before adding a separate search engine, queue, or document DB: how far PostgreSQL gets you on its own, and where its limit begins.
- Journal4 sections
Signals That a System Has Grown Too Complex
The concrete red flags that reveal code or architecture has turned needlessly complex — and where to start simplifying
- Note6 sections
Circuit Breakers for External Service Integrations
Connecting a critical system to external APIs: using timeouts, retries, and circuit breakers to stop one service's collapse from taking the whole system down.
- Note7 sections
The Serverless Migration Decision: Cold-Start and Vendor Lock-in
Before you move to FaaS: the real cost of cold-start latency, vendor lock-in risk, and the workloads where serverless actually pays off
- Note6 sections
The BFF Pattern: A Separate API Layer for Mobile and Web
Separate API layers for mobile and web: the hidden cost of bending one API to fit every client, and when a BFF is actually warranted
- Note7 sections
Synchronous or Asynchronous? The Line Between HTTP and the Queue
Should a job run inside the HTTP request or go to a queue? The decision line, drawn through response time and fault tolerance
- System10 sections
Transactional Outbox: Dual-Write, At-Least-Once, and Idempotent Consumption
Writing to the database but failing to put the event on the queue: closing dual-write with an outbox, suppressing the at-least-once repeats it creates, and encrypting the field.
- Note5 sections
Event-Driven Architecture: When It Saves You, When It Adds Complexity
The price of loose coupling through events: balancing the flexibility it buys against the risk of making your system's flow impossible to trace
- System7 sections
GPU FinOps with eBPF
Why GPU spend is invisible to cgroup and cAdvisor metrics, and how eBPF attributes GPU time and memory to teams for chargeback — with the honest limit at the CUDA boundary that DCGM has to cover.
- System6 sections
Zero Static Authority in Multi-Cluster GitOps
GitOps across many Kubernetes clusters with no long-lived kubeconfig or token. SPIFFE IDs, SPIRE-issued short-lived SVIDs, trust-domain federation — the order to build it in, and what each step costs.
- Note5 sections
Replaying a Dead-Letter Queue Without Making It Worse
The DLQ filled up during an outage and you want it back. Naive replay re-poisons the queue or double-fires side effects. The safe shape: reset, dry-run, sandbox-first, select.
- System5 sections
Distributed Tracing Across a Polyglot Queue
A message crosses PHP, Go and Python over a frozen envelope. Turning its journey into one OpenTelemetry trace with no new field and no core dependency — and the honest limit of that.
- Note4 sections
Validating the Schema at the Edge of the Queue
A message's data is an untyped contract the queue won't check. Validate it at the edge: producer-side before publish, consumer-side as a safety net, against a per-URN schema
- Journal4 sections
Your Architecture Decisions Rot Because Nothing Runs Them
An architecture decision written in a wiki is a wish. The boundaries you can't enforce are the ones that quietly erode — until the diagram stops matching the code. Make the rule executable.
- Note7 sections
Idempotency: When the Same Message Arrives Twice
At-least-once delivery means a handler will see the same message twice. The idempotency key, where it comes from, and the dedupe that survives a crash
- Journal5 sections
AI Didn't Remove the Bottleneck — It Moved It
Cheap code production doesn't raise throughput if review, integration, and verification can't keep up. The bottleneck just moves downstream — and that's where the work now is
Elsewhere
My main software blog
sade.dev keeps the architecture notes and the production lessons. The rest of software — languages, tools, practice, teams — lives on the main blog.