SeniorDesign Lab
SYSTEMS ENGINEERING / LEARNING COCKPIT
● Offline-ready
Senior-level system design • interactive field guide

Design for failure, scale and change.

Move beyond boxes-and-arrows theatre. Explore how routing, state, queues, caches, replication and hardware limits behave when traffic rises or dependencies fail. Change the assumptions and watch the model respond.

9 learning areas6 interactive modelsNo external libraries
CLIENTS GATEWAY SERVICE A SERVICE B REPLICATED DATA REQUESTSSTATEOBSERVE • ISOLATE • RECOVER
Traffic path
Edge → Data
Every hop adds latency and failure modes
Resilience loop
Detect → Isolate
Then recover, degrade, or reject safely
Capacity basis
Peak + Headroom
Average load is not a sizing plan
Senior signal
Trade-offs
Explain why, cost, risk and failure path
Start with the problem

A practical design sequence

Use this order in interviews and real design reviews. It prevents architecture-by-brand-name.

8-step framework
01

Clarify requirements

Users, core workflows, latency SLOs, availability, consistency and compliance boundaries.

02

Quantify the load

DAU, requests/user, read/write mix, payload size, peak factor, growth and retention.

03

Define the data

Entities, access patterns, indexes, partition keys, idempotency and ownership.

04

Draw the thin path

Client to edge to service to state store. Keep the first architecture boring.

05

Find the bottleneck

CPU, connection pools, disk IOPS, lock contention, hot keys, egress or downstream quotas.

06

Design failure behavior

Timeouts, bounded retries, circuit breakers, backpressure, graceful degradation and recovery.

07

Protect correctness

Transactions, deduplication, ordering, replay, reconciliation and consistency guarantees.

08

Prove it operationally

SLIs/SLOs, dashboards, tracing, alerts, load tests, runbooks and rollback strategy.

Learning modules

Choose a system behavior to inspect

Each module gives you a mental model, a visual, and controls to test the trade-offs.

⌘
Topology

Architecture studio

Compare global active-active, CQRS/event-driven, read-scaling, and cell-based designs. Inject a region failure and discuss blast radius.

◉
Distribution

Consistent hashing

Add servers and virtual nodes to see ownership distribution, imbalance and why moving only part of the keyspace matters.

⛨
Failure

Circuit breaker

Trip a breaker, see fast-fail behavior, advance recovery and send the single half-open probe.

▦
Latency

Cache economics

Change hit ratio and workload to estimate backend QPS, approximate latency and working-set memory cost.

≋
Flow control

Queues & backpressure

Model whether workers keep up with arrival rate after retry amplification. Watch backlog grow or drain over time.

▤
Sizing

Capacity planner

Estimate peak IOPS, bandwidth, replicated storage, index memory and a first-pass node count with headroom.

The senior lens

Questions behind every architecture choice

What fails first?

Name the resource or dependency that saturates first, the signal that detects it, and the mechanism that limits the blast radius.

What is allowed to be stale?

State the consistency contract at the API boundary. “Eventually consistent” needs a time window and a business consequence.

How does it recover?

Explain queue drain, replica catch-up, cache warming, leader election, rollback and safe replay after an outage.

Architecture studio

Patterns are choices, not trophies.

Pick a topology to inspect the request path, state placement, trade-offs, and a failure scenario. The drawing is explanatory, not a vendor-specific deployment blueprint.

SVG architecture maps

Global active-active routing

Route users near the edge; be explicit about write ownership and replication lag.

Request path Compute / orchestration Data / durable state Async / replication Failed / degraded
ⓘ
Failure scenario: Route traffic away from an unhealthy region only if the surviving region has capacity and the data consistency model permits it.
Why teams choose it

Lower regional latency and fault isolation, with multiple deployment locations.

Trade-offs to defend

Cross-region replication, conflict policy, failover time, and cost of warm standby capacity.

Operational proof

Exercise regional loss, verify write ownership, observe replication lag, and measure recovery-point objective.

A useful distinction

Architecture attributes are coupled

Every new boundary trades one kind of risk for another.

Availability

Can a request still succeed when a component or region fails?

Consistency

Can reads observe stale, reordered or conflicting values?

Latency

How many network round trips and serial dependencies sit on the critical path?

Operability

Can engineers debug, deploy, roll back, and recover it at 3 a.m.?

Interactive simulation 01

Traffic distribution & load balancing

Adjust offered traffic, choose a distribution strategy, and take nodes out of service. This simplified model exposes overload and capacity headroom, not the exact behavior of a production proxy.

Request routing

Healthy capacity vs. offered load

Toggle node health to trigger redistribution.

Balanced

Per-node utilization

Bars are actual allocated requests against the node's configured capacity.

ⓘ
Every healthy node has spare capacity at the current offered load.

Workload controls

Use realistic peak traffic, not only daily average.

3,000 RPS
Offered load before rate limiting or rejection.
Real least-connections routing considers active connections, and weighted routing depends on configuration and health.
Healthy nodes
4 / 4
Available endpoints
Safe capacity
3,150
At 70% per-node target
Peak utilization
68%
Most loaded healthy node
At risk traffic
0 RPS
Above modeled capacity
Design notes

Load balancing does not create capacity

Health checks

Distinguish process liveness from readiness. A process can be alive but unable to serve because its DB pool is exhausted.

Overload strategy

Use bounded queues, admission control, rate limits and priority shedding. Unbounded waiting converts overload into a latency outage.

Sticky sessions

Affinity simplifies some flows but can create skew and complicate failover. Prefer externalized session state when practical.

Interactive simulation 02

Consistent hashing & virtual nodes

Map 1,000 evenly spaced sample keys to a hash ring. Increase virtual nodes to reduce ownership variance, then add or remove a physical server and inspect the rebalancing effect.

Partition ownership

Dots on the ring represent virtual-node positions. Each test key is assigned clockwise to the next virtual node, wrapping at the end of the ring.

Cluster topology

Change the number of servers and virtual positions.

16
More virtual nodes usually improve distribution at the cost of a larger ring and more metadata.
Physical servers
4
Ring positions
64
Ownership CV
—
Std. deviation / mean
Key remap on resize
—
Compared with previous topology
1,000 keys
ⓘ
Virtual nodes reduce variance in this model. They do not remove hot keys or skewed request cost.
Where it fits

Common uses and common mistakes

Useful for

Distributed caches, sharded key-value stores, partitioned routing, and systems that need a relatively small key remap when membership changes.

Not a magic balancing spell

If one key accounts for a large share of traffic, all requests for that key can still hammer one owner. Consider key salting, replication or request coalescing.

Production considerations

Membership changes need stable hashing, ownership convergence, data migration, replica placement and a safe handoff protocol.

Interactive simulation 03

Circuit breaker state machine

Generate successful and failed downstream calls. After enough consecutive failures, the breaker opens and rejects calls locally. Advance the recovery timer, then test the dependency with a single half-open probe.

Cascading failure control

Request path

Green means a request is allowed through. Red means fast-fail or downstream failure.

CLOSED
CLIENTAPI GATEWAYCIRCUIT BREAKERCLOSEDSERVICE BFAST FAIL • NO DOWNSTREAM CALLREQUESTDEPENDENCY
Success responses
0
Downstream failures
0
Fast-failed calls
0

Event timeline

Latest events first.

Drive the state machine

No automatic timer: recovery is advanced manually so the transition is easy to inspect.

3 consecutive failures
Success in CLOSED resets the consecutive-failure counter.
One click = one call attempt
ⓘ
Closed: calls flow to Service B. Failures accumulate until the configured threshold is reached.

State transition rules

CLOSED: normal traffic; count failures.

OPEN: fail fast without contacting the dependency.

HALF-OPEN: allow one probe. Success closes the breaker; failure opens it again.

⚠
A breaker is not a complete resilience strategy. Pair it with bounded timeouts, limited retries with jitter, bulkheads, concurrency limits, fallback behavior where correct, and telemetry. Retries without budgets can amplify the outage.
Interactive simulation 04

Cache hit ratio & origin economics

Estimate how hit ratio affects downstream request volume and latency. The working-set memory figure is a rough capacity estimate, not a prediction of hit rate from memory alone.

Latency + cost

Requests served at the edge of the data tier

Cache hits avoid the modeled origin round trip.

Origin protected
CLIENTSRPS loadCACHE80% HITTTL / evictionORIGIN DB600 RPSHITS SERVED2,400 RPSMISS PATHHIT PATH
Origin request rate
600
After cache hit reduction
Origin load reduction
80%
Compared with no cache
Estimated average latency
10 ms
Simple weighted path model
Working-set memory
19.1 GiB
Approx. cached objects only

Where the requests go

Hit vs. miss split at current settings.

Cache hitsOrigin misses

Workload assumptions

Change the access pattern and object footprint.

3,000 RPS
80%
Observed hit ratio is workload-dependent. TTL, locality, eviction policy and invalidation behavior matter.
2 ms
40 ms
4 KB
5 million
ⓘ
Model: average latency = cache lookup latency + miss ratio × origin latency. Memory is object size × working-set objects × hit ratio. Real systems also need metadata, allocator overhead, replication and fragmentation headroom.
Consistency and correctness

Cache invalidation is a product decision

TTL-based expiry

Simple operational model; stale reads are possible up to the TTL plus propagation and clock effects.

Write-through / write-around

Write-through keeps cache updated on writes but adds write-path coupling. Write-around can protect cache from one-off writes.

Stampede control

Use request coalescing, jittered TTLs, stale-while-revalidate, and bounded refresh concurrency for hot keys.

Interactive simulation 05

Queues, retries & backpressure

A queue absorbs a burst; it does not make overload disappear. Model incoming jobs, worker throughput, processing time and transient failure retries to estimate whether backlog grows or drains.

Flow control

Backlog over the next 60 seconds

The chart uses a simple deterministic arrival/service model.

Stable
Effective attempt rate
500/s
Original jobs plus retries
Worker attempt capacity
600/s
Workers × service rate
Backlog after 60s
0
Attempt-level jobs pending
Retry amplification
1.00×
Expected attempts per job
ⓘ
Worker capacity exceeds the retry-adjusted attempt rate, so the modeled backlog remains bounded.

Queue workload

Change the arrival and service rates.

500/s
12 workers
20 ms
5%
Assumes independent failures and immediate retries; production failures are often correlated, making this optimistic.
3 attempts
⚠
Production guardrails: retries need a budget, exponential backoff, jitter, idempotent handlers, dead-letter handling and a policy for poison messages. Backpressure should slow producers before the queue becomes an infinite storage bill.
Queue semantics

Separate acceptance, processing and completion

At-least-once delivery

Messages may be redelivered. Make side effects idempotent or deduplicate using a durable operation identifier.

Ordering

Global ordering constrains parallelism. Partitioned ordering often gives a better trade-off if the business key is chosen carefully.

Dead-letter queue

A DLQ is not where failures go to disappear. Define alerting, replay rules, ownership, and safe remediation.

Interactive simulation 06

Capacity planning & resource envelopes

Estimate peak database IOPS, application bandwidth, five-year replicated storage, index memory and a first-pass node count. Replace every assumption with measured workload data before treating the estimate as a deployment plan.

Sizing model

Scenario presets

Start with a profile, then tune the inputs.

Workload assumptions

All rates derive from the inputs on the left.

1,000,000
30
10%
Remaining requests are modeled as reads.
2 KB
4×
3 copies
30%
32 bytes
Illustrative bytes per stored row for the selected index/model. Actual index memory depends on keys, fill factor, pointers and engine internals.
50,000 IOPS
Node count reserves 30% headroom by targeting 70% nominal IOPS. Benchmark with your actual read/write mix.

Estimated resource envelope

Numbers recalculate as you change the workload.

First-pass estimate
Peak database IOPS
—
Read + replica write work
Peak client bandwidth
—
Payload throughput estimate
Five-year storage
—
Includes replica copies, no backup copies
Five-year index RAM
—
Primary index footprint estimate
Recommended DB nodes
—
IOPS target at 70% per node
Daily writes
—
Before internal write amplification

Five-year growth projection

Storage is cumulative. Index RAM follows cumulative written rows.

Replicated storage (TB)Index RAM (GB)
⚠
Do not size production from this alone. The model ignores compression, deletes/updates, index/table overhead beyond the selected factor, query plans, lock contention, connection-pool ceilings, replication topology, backups, failover headroom and correlated peaks. Validate with load tests and engine-specific measurements.
Capacity reasoning

Keep the units visible

IOPS is not QPS

One query may touch many pages. Cache hit rate, query shape, indexes, page size, storage latency and write amplification change physical I/O.

Storage is not memory

Durable data, indexes, buffer pool, heap, cache and temporary spill files have different sizing drivers. Do not multiply one figure and call it done.

Headroom buys recovery

A cluster at 99% utilization has little room for failover, compaction, rebalancing, retries or a traffic burst. Target based on observed tail behavior.

Interview playbook

Explain the decision, not just the diagram.

Use the structured script to lead a system design discussion. Then test yourself with questions that reward clear contracts and realistic failure behavior.

Practice mode

45-minute design rhythm

A time-boxed sequence for a typical senior system design round.

5m

Requirements + SLOs

Identify core flows, scale, p95/p99 latency targets, availability and consistency constraints.

5m

Estimation

Derive request rate, payload bandwidth, storage growth and peak multiplier. Show your units.

10m

High-level architecture

Explain every major component and why state lives where it does.

10m

Deep dive on two risks

Pick the likely bottleneck and correctness challenge; trace the data and failure paths.

10m

Reliability + evolution

Failures, backpressure, scaling triggers, migration, observability and recovery.

5m

Trade-offs + recap

Call out limitations, alternatives, costs and what you would validate first.

Knowledge check

Answer out loud before revealing the explanation.

Question 1 / 8
Consistency
What would you do when a client retries a timed-out payment request?
Decision reference

Trade-offs worth articulating

SQL vs. NoSQL

Start from constraints and access patterns: transaction boundaries, relational integrity, query flexibility, predictable key access, scale and operational skill.

Sync vs. async

Synchronous workflows provide immediate results but tie latency and availability to dependencies. Async workflows improve isolation but introduce delivery, ordering and reconciliation complexity.

Cache vs. read replica

A cache reduces repeated data access with explicit freshness rules. A replica provides a database read copy but has replication lag and still carries query/connection overhead.

Sharding vs. vertical scaling

Vertical scaling is simpler until it is not. Sharding adds ownership, rebalancing, cross-shard query, transaction and incident-response complexity.

Exactly-once claims

End-to-end “exactly once” is usually a composition of scoped guarantees, atomic writes, idempotency and deduplication. Be precise about the boundary.

Multi-region writes

Active-active is not just duplicating services. Define conflict resolution, consistency, ownership, failover fencing and split-brain behavior.

Searchable reference

System design glossary

Search terms to revise before a design round.

Backpressure

A mechanism that makes producers slow down, reject work or limit concurrency when downstream capacity is insufficient. It protects bounded resources and prevents an overload event from becoming a system-wide collapse.

Bulkhead

Isolation of a resource pool or concurrency budget so failure or saturation in one workload does not consume all capacity needed by other workloads.

Idempotency key

A stable operation identifier used to recognize retries of the same logical request. Its effect depends on durable storage, atomicity, retention period and the exact scope of the operation.

Quorum

A required number of votes or acknowledgements for a distributed operation. Quorum intersection can support consistency properties, but the exact guarantee depends on protocol and read/write rules.

p95 / p99 latency

The latency under which 95% or 99% of measured requests complete. Tail latency often reveals queueing and dependency problems hidden by averages.

Hot partition / hot key

A small subset of keys or partitions receives disproportionate traffic or work, causing local saturation even when cluster-wide average utilization appears healthy.

RPO / RTO

Recovery Point Objective is the acceptable data-loss window. Recovery Time Objective is the target time to restore a service. They are distinct goals and need tested procedures.

Write amplification

Physical writes caused per logical application write. Index maintenance, compaction, replication, journaling and storage-engine behavior can make it substantially greater than one.

Idempotent consumer

A consumer that can process a duplicate delivery without duplicating the business effect, commonly using unique operation IDs, inbox/outbox patterns or atomic deduplication with the state update.

Cell-based architecture

A fleet of relatively independent slices of the service and data plane, each serving a subset of customers or traffic. Cells limit blast radius but require routing, capacity management and migration between cells.

Tail at scale

As a request depends on more components, the chance that at least one component has a slow tail increases. Parallel fan-out can make the slowest dependency dominate end-to-end latency.

Graceful degradation

Preserving a smaller, explicit set of useful behaviors when dependencies fail, such as serving stale non-critical data or disabling a secondary feature instead of failing the entire request.

This page is a learning and estimation tool, not a substitute for architecture reviews, production measurements, formal protocol analysis or vendor-specific testing. Simplified models are labeled as such; treat outputs as hypotheses to validate.