Architecture studio
Compare global active-active, CQRS/event-driven, read-scaling, and cell-based designs. Inject a region failure and discuss blast radius.
Move beyond boxes-and-arrows theatre. Explore how routing, state, queues, caches, replication and hardware limits behave when traffic rises or dependencies fail. Change the assumptions and watch the model respond.
Use this order in interviews and real design reviews. It prevents architecture-by-brand-name.
Users, core workflows, latency SLOs, availability, consistency and compliance boundaries.
DAU, requests/user, read/write mix, payload size, peak factor, growth and retention.
Entities, access patterns, indexes, partition keys, idempotency and ownership.
Client to edge to service to state store. Keep the first architecture boring.
CPU, connection pools, disk IOPS, lock contention, hot keys, egress or downstream quotas.
Timeouts, bounded retries, circuit breakers, backpressure, graceful degradation and recovery.
Transactions, deduplication, ordering, replay, reconciliation and consistency guarantees.
SLIs/SLOs, dashboards, tracing, alerts, load tests, runbooks and rollback strategy.
Each module gives you a mental model, a visual, and controls to test the trade-offs.
Compare global active-active, CQRS/event-driven, read-scaling, and cell-based designs. Inject a region failure and discuss blast radius.
Add servers and virtual nodes to see ownership distribution, imbalance and why moving only part of the keyspace matters.
Trip a breaker, see fast-fail behavior, advance recovery and send the single half-open probe.
Change hit ratio and workload to estimate backend QPS, approximate latency and working-set memory cost.
Model whether workers keep up with arrival rate after retry amplification. Watch backlog grow or drain over time.
Estimate peak IOPS, bandwidth, replicated storage, index memory and a first-pass node count with headroom.
Name the resource or dependency that saturates first, the signal that detects it, and the mechanism that limits the blast radius.
State the consistency contract at the API boundary. “Eventually consistent” needs a time window and a business consequence.
Explain queue drain, replica catch-up, cache warming, leader election, rollback and safe replay after an outage.
Pick a topology to inspect the request path, state placement, trade-offs, and a failure scenario. The drawing is explanatory, not a vendor-specific deployment blueprint.
Route users near the edge; be explicit about write ownership and replication lag.
Lower regional latency and fault isolation, with multiple deployment locations.
Cross-region replication, conflict policy, failover time, and cost of warm standby capacity.
Exercise regional loss, verify write ownership, observe replication lag, and measure recovery-point objective.
Every new boundary trades one kind of risk for another.
Can a request still succeed when a component or region fails?
Can reads observe stale, reordered or conflicting values?
How many network round trips and serial dependencies sit on the critical path?
Can engineers debug, deploy, roll back, and recover it at 3 a.m.?
Adjust offered traffic, choose a distribution strategy, and take nodes out of service. This simplified model exposes overload and capacity headroom, not the exact behavior of a production proxy.
Toggle node health to trigger redistribution.
Bars are actual allocated requests against the node's configured capacity.
Use realistic peak traffic, not only daily average.
Distinguish process liveness from readiness. A process can be alive but unable to serve because its DB pool is exhausted.
Use bounded queues, admission control, rate limits and priority shedding. Unbounded waiting converts overload into a latency outage.
Affinity simplifies some flows but can create skew and complicate failover. Prefer externalized session state when practical.
Map 1,000 evenly spaced sample keys to a hash ring. Increase virtual nodes to reduce ownership variance, then add or remove a physical server and inspect the rebalancing effect.
Dots on the ring represent virtual-node positions. Each test key is assigned clockwise to the next virtual node, wrapping at the end of the ring.
Change the number of servers and virtual positions.
Distributed caches, sharded key-value stores, partitioned routing, and systems that need a relatively small key remap when membership changes.
If one key accounts for a large share of traffic, all requests for that key can still hammer one owner. Consider key salting, replication or request coalescing.
Membership changes need stable hashing, ownership convergence, data migration, replica placement and a safe handoff protocol.
Generate successful and failed downstream calls. After enough consecutive failures, the breaker opens and rejects calls locally. Advance the recovery timer, then test the dependency with a single half-open probe.
Green means a request is allowed through. Red means fast-fail or downstream failure.
Latest events first.
No automatic timer: recovery is advanced manually so the transition is easy to inspect.
CLOSED: normal traffic; count failures.
OPEN: fail fast without contacting the dependency.
HALF-OPEN: allow one probe. Success closes the breaker; failure opens it again.
Estimate how hit ratio affects downstream request volume and latency. The working-set memory figure is a rough capacity estimate, not a prediction of hit rate from memory alone.
Cache hits avoid the modeled origin round trip.
Hit vs. miss split at current settings.
Change the access pattern and object footprint.
Simple operational model; stale reads are possible up to the TTL plus propagation and clock effects.
Write-through keeps cache updated on writes but adds write-path coupling. Write-around can protect cache from one-off writes.
Use request coalescing, jittered TTLs, stale-while-revalidate, and bounded refresh concurrency for hot keys.
A queue absorbs a burst; it does not make overload disappear. Model incoming jobs, worker throughput, processing time and transient failure retries to estimate whether backlog grows or drains.
The chart uses a simple deterministic arrival/service model.
Change the arrival and service rates.
Messages may be redelivered. Make side effects idempotent or deduplicate using a durable operation identifier.
Global ordering constrains parallelism. Partitioned ordering often gives a better trade-off if the business key is chosen carefully.
A DLQ is not where failures go to disappear. Define alerting, replay rules, ownership, and safe remediation.
Estimate peak database IOPS, application bandwidth, five-year replicated storage, index memory and a first-pass node count. Replace every assumption with measured workload data before treating the estimate as a deployment plan.
Start with a profile, then tune the inputs.
All rates derive from the inputs on the left.
Numbers recalculate as you change the workload.
Storage is cumulative. Index RAM follows cumulative written rows.
One query may touch many pages. Cache hit rate, query shape, indexes, page size, storage latency and write amplification change physical I/O.
Durable data, indexes, buffer pool, heap, cache and temporary spill files have different sizing drivers. Do not multiply one figure and call it done.
A cluster at 99% utilization has little room for failover, compaction, rebalancing, retries or a traffic burst. Target based on observed tail behavior.
Use the structured script to lead a system design discussion. Then test yourself with questions that reward clear contracts and realistic failure behavior.
A time-boxed sequence for a typical senior system design round.
Answer out loud before revealing the explanation.
Start from constraints and access patterns: transaction boundaries, relational integrity, query flexibility, predictable key access, scale and operational skill.
Synchronous workflows provide immediate results but tie latency and availability to dependencies. Async workflows improve isolation but introduce delivery, ordering and reconciliation complexity.
A cache reduces repeated data access with explicit freshness rules. A replica provides a database read copy but has replication lag and still carries query/connection overhead.
Vertical scaling is simpler until it is not. Sharding adds ownership, rebalancing, cross-shard query, transaction and incident-response complexity.
End-to-end “exactly once” is usually a composition of scoped guarantees, atomic writes, idempotency and deduplication. Be precise about the boundary.
Active-active is not just duplicating services. Define conflict resolution, consistency, ownership, failover fencing and split-brain behavior.
Search terms to revise before a design round.
A mechanism that makes producers slow down, reject work or limit concurrency when downstream capacity is insufficient. It protects bounded resources and prevents an overload event from becoming a system-wide collapse.
Isolation of a resource pool or concurrency budget so failure or saturation in one workload does not consume all capacity needed by other workloads.
A stable operation identifier used to recognize retries of the same logical request. Its effect depends on durable storage, atomicity, retention period and the exact scope of the operation.
A required number of votes or acknowledgements for a distributed operation. Quorum intersection can support consistency properties, but the exact guarantee depends on protocol and read/write rules.
The latency under which 95% or 99% of measured requests complete. Tail latency often reveals queueing and dependency problems hidden by averages.
A small subset of keys or partitions receives disproportionate traffic or work, causing local saturation even when cluster-wide average utilization appears healthy.
Recovery Point Objective is the acceptable data-loss window. Recovery Time Objective is the target time to restore a service. They are distinct goals and need tested procedures.
Physical writes caused per logical application write. Index maintenance, compaction, replication, journaling and storage-engine behavior can make it substantially greater than one.
A consumer that can process a duplicate delivery without duplicating the business effect, commonly using unique operation IDs, inbox/outbox patterns or atomic deduplication with the state update.
A fleet of relatively independent slices of the service and data plane, each serving a subset of customers or traffic. Cells limit blast radius but require routing, capacity management and migration between cells.
As a request depends on more components, the chance that at least one component has a slow tail increases. Parallel fan-out can make the slowest dependency dominate end-to-end latency.
Preserving a smaller, explicit set of useful behaviors when dependencies fail, such as serving stale non-critical data or disabling a secondary feature instead of failing the entire request.