Evolve Slack Architecture with Cells

Create a sourced Slack architecture evolution board in Excalidraw.

Introduction

30 Second Summary

A team chat can feel instant while huge amounts of work happen behind the screen. Growth exposes weak points that users never see until messages slow down or an outage spreads.

In this project, you will build an Excalidraw case-study board that shows how Slack's architecture evolves from workspace shards to failure-isolated cells. You will connect each change to the product pressure or failure that made it necessary.

What You'll Build

Your finished board will let you give a five-minute walkthrough that connects every bottleneck to a measurable architecture decision.

By the end of this project, you'll have:

  • A source-backed architecture story across four stages that lets you trace a message from client to recipient.
  • A clearly labeled simulation layer that turns workload pressure into numbers without presenting estimates as Slack production data.
  • A five-minute interview narrative that defends why every architecture change became necessary.
  • Secret Mission: Add an org-wide tenancy layer for users and channels spanning multiple workspaces.

Are there any prerequisites?

Bring a basic understanding of system design plus a browser that can access Excalidraw. No installation or paid account is required for the project path.

Before We Start

This is your moment to commit to explaining why Slack's architecture evolved under product, scale, or reliability pressure. You are building a sourced case-study board with clearly labeled simulations for a system design interview.

Set Up the Evolution Canvas

A useful architecture evolution board shows why each change became necessary. Without clear evidence boundaries, invented workload estimates can look like Slack production facts.

In this step, you'll use Excalidraw to create the four-stage canvas. You'll add reusable structures that separate documented history from hypothetical simulations.

In this step, get ready to:
  • Create a titled canvas with a local source file.
  • Lay out four empty stage panels with evidence labels.
  • Add a metric strip plus a reusable decision card.
Open and protect the canvas

The browser canvas becomes the source for your whole case study. Saving a local copy protects the editable board if browser storage is cleared.

  • Open the official Excalidraw web app in a new browser tab.
  • Select the text tool from the top toolbar.
  • Click near the upper-left area of the blank canvas.
  • Enter Evolving Slack: From Workspace Shards to Failure-Isolated Cells as the board title.

You should see the complete title near the top of the canvas. This title anchors the architecture stages that follow.

The save control is slightly tucked away in the canvas menu at the upper left. Your browser may show a file dialog during the first local save.

  • Open the canvas menu in the upper-left corner.
  • Select Save.
  • Choose a local folder in the browser file dialog.
  • Confirm the save.

Good progress. You should now have an editable .excalidraw file stored on your computer.

Why save a local copy?

Excalidraw stores drawings in browser storage. Browser storage can be cleared unexpectedly.

The local file gives you a recoverable source copy throughout the project.

Can't find the saved file?

  • Check the downloads folder selected by your browser.
  • Return to the canvas menu if no file was saved.
  • Select Save again.

Still stuck? Help me locate my saved Excalidraw file.

Build the stage and evidence layout

A left-to-right layout makes the evolution visible as a sequence of responses to pressure. Empty panels reserve space for each architecture stage without introducing future designs early.

  • Zoom out until the canvas has enough horizontal space for four large panels.
  • Select the rectangle tool from the top toolbar.
  • Draw four similarly sized rectangles in one horizontal row.

You should see four empty areas with a clear left-to-right reading order.

  • Select the text tool from the top toolbar.
  • Add Workspace Shards above the leftmost panel.
  • Add Flannel + Vitess above the second panel.

The first half of the board should now move from workspace-level sharding to a more flexible storage boundary.

  • Add Real-Time Fan-Out above the third panel.
  • Add Failure-Isolated Cells above the rightmost panel.

You should now see all four stage names in sequence. Keep the panel interiors empty for the architecture diagrams added later.

The evidence legends create a visual contract with anyone reading the board. Each future fact or calculation gets a visible source category.

  • Draw two small rectangles beneath the board title.
  • Apply a different fill color to each rectangle.

You should see two distinct color keys above the architecture panels.

  • Add Slack-documented inside the first key.
  • Add Hypothetical simulation inside the second key.

Why separate the evidence types?

The Slack-documented label identifies architecture details or historical measurements supported by Slack Engineering sources.

The Hypothetical simulation label identifies invented workloads used to test the design. The distinction keeps every claim defensible.

The metric strip gives each stage a consistent place for its measurable pressure. This makes later calculations easier to compare across the board.

  • Draw one long narrow rectangle beneath the evidence legend.
  • Divide the strip into four equal spaces with three vertical lines.

You should see four empty metric spaces in one row.

  • Add throughput to the first metric space.
  • Add connections to the second metric space.

The first half of the metric strip now covers traffic volume plus persistent client load.

  • Add fan-out to the third metric space.
  • Add recovery to the fourth metric space.

You should now see a four-part metric strip for throughput, connections, fan-out, and recovery.

Create the reusable decision card

Each stage needs a short decision record that connects pressure to an architectural response. The same field structure makes the trade-offs comparable.

  • Draw a medium rectangle beside the legend area.
  • Select the text tool from the top toolbar.
  • Add Trigger near the top of the card.
  • Add Bottleneck below the first field.

The card now starts with the pressure that causes change plus the limitation that pressure exposes.

  • Add Decision below the bottleneck field.
  • Add Trade-off below the decision field.

The middle of the card now captures the chosen response plus its cost.

  • Add Rejected alternative below the trade-off field.
  • Add Evidence at the bottom of the card.

The decision card should now contain all six required fields.

  • Drag a selection box around the card plus its field labels.
  • Use the grouping control in the actions panel.
  • Drag the grouped card slightly to the right.

All six labels should move with the card. This proves the card is reusable as one visual unit.

  • Return the card to its original position.
  • Open the canvas menu in the upper-left corner.
  • Select Save to update the local source file.

Before you inspect the full canvas, do you expect every required placeholder to be present?

  • Zoom out until the complete board is visible.
  • Inspect the four-stage row from left to right.
  • Inspect the evidence legend plus the metric strip.
  • Inspect the grouped decision card.

You should see four empty panels arranged from Workspace Shards through Failure-Isolated Cells.

You should see the Slack-documented legend plus the Hypothetical simulation legend.

The metric strip should show all four measurement categories. The reusable decision card should show all six required fields.

Does the board feel cramped?

  • Zoom out until all four panels fit across the screen.
  • Move the legends above the panel row if they overlap a stage.
  • Widen the space between panels if their headings touch.

Need another pair of eyes? Help me improve my Excalidraw board layout.

Your evolution canvas now has a protected source file plus a clear evidence system. Next, you'll use the first panel to expose the workspace-sharded bottleneck.

Expose the Workspace-Sharded Bottleneck

Your Excalidraw canvas now gives each stage of Slack's evolution its own space. This step fills the first panel with the historical baseline that everything else must respond to.

Workspace-level database sharding gives each customer a clear storage boundary. You will test how that boundary behaves when one workspace generates a large share of the workload.

In this step, get ready to:
  • Draw Slack's documented workspace-sharded baseline.
  • Trace persistence before real-time message delivery.
  • Calculate fan-out volume and expose one hot workspace shard.
Draw the documented baseline

The baseline uses MySQL shards as workspace-level storage boundaries. The Webapp monolith owns the routing decision that selects the correct shard.

  • Focus on the Workspace Shards panel in the Excalidraw board from earlier.
  • Use a rectangle to add Clients on the far left of the panel.

You should see the sending clients at the starting edge of the architecture.

  • Add a Webapp servers box to the right of the clients.
  • Add a metadata lookup box beneath the Webapp servers.

The request tier now has a separate lookup that supports shard selection.

  • Create three boxes labeled MySQL workspace shard in a row near the center-right of the panel.
  • Add a real-time message servers box to the right of the shard row.

You should now see storage separated from the real-time delivery tier.

  • Add a Connected clients box on the far right of the panel.
  • Place a Slack-documented evidence label above the diagram.

The panel now reads from request senders on the left to message recipients on the right.

  • Add the note Webapp monolith managed shard routing beneath the Webapp servers.

Good. Your first stage now shows the documented services without mixing in any simulated workload values.

How Workspace Routing Works

The Webapp monolith uses the workspace identity to select the shard holding that workspace's data. The metadata lookup supports that routing decision.

One workspace maps to one database shard in this historical model. The workspace boundary therefore determines where its storage load lands.

Trace persistence before delivery

A WebSocket keeps a persistent connection between a client and the real-time service. Slack persists each message before sending it through this real-time path.

  • Draw an arrow from Clients to Webapp servers.
  • Label that arrow with 1. client API request.

You should see the message entering the HTTP request tier before it reaches storage.

  • Place the number 2 inside the Webapp servers box.
  • Connect Webapp servers to metadata lookup with an unnumbered arrow.

The Webapp now has a visible route to the lookup that identifies the workspace shard.

  • Connect metadata lookup to the shard row with an unnumbered arrow.
  • Draw an arrow from Webapp servers to one shard labeled 3. workspace shard persistence.

The selected database shard now sits directly in the numbered message path.

  • Draw an arrow from the selected shard to real-time message servers labeled 4. real-time message server.
  • Draw an arrow from real-time message servers to Connected clients labeled 5. connected clients over WebSockets.

The complete route now ends with fan-out to connected recipients.

  • Add the note Persistence before real-time delivery beside the third arrow.

Before you follow the numbers, where do you expect durability to appear in the path? Hold that prediction while you inspect the arrows.

  • Follow the numbered path from the sending client to the connected recipients.

You should reach workspace shard persistence before real-time message server. The route now makes Slack's durability order visible.

Why Persistence Comes First

The database write creates a durable record before the real-time stack distributes the message. The delivery path can fan out only after storage succeeds.

This ordering separates message durability from the temporary state held by connected real-time services.

Calculate the hot-shard pressure

Fan-out multiplies each stored message by its average number of online recipients. The same workload can then be mapped onto the workspace storage boundary.

  • Create a rectangular simulation card in the lower part of the Workspace Shards panel.
  • Label the card Hypothetical simulation.

The simulation now has its own visual boundary beneath the documented architecture.

  • Add the warning Invented for this exercise, not Slack production data at the top of the card.
  • Add 100,000 concurrent connections as the connection assumption.

You should see the evidence warning before the first workload assumption.

  • Add 2,000 messages/s as the incoming message rate.
  • Add 40 average online recipients as the fan-out assumption.

The card now contains the three inputs needed to calculate delivery volume.

Before you multiply the values, do you expect recipient count to make delivery volume much larger than message volume? Keep your estimate in mind while you calculate.

  • Calculate the fan-out by multiplying 2,000 messages/s by 40 average online recipients.

The result is 80,000 deliveries/s. Each incoming message creates forty recipient deliveries on average.

  • Add 2,000 × 40 = 80,000 deliveries/s to the simulation card.
  • Add the assumption that one large workspace generates 25% of writes.

The delivery calculation is now paired with a concentrated workspace workload.

Before you map that share to storage, can this design split one workspace's writes across several shards? Test your prediction with the next calculation.

  • Calculate 25% of the 2,000 messages/s workload.

The result is 500 message writes/s for one workspace. That entire write rate lands on the shard assigned to that workspace.

  • Add 500 message writes/s to the simulation card.
  • Circle the selected workspace shard in the architecture diagram.

You should see one shard carrying the large workspace's complete write load.

  • Label the circled box Indivisible hot workspace shard.
  • Add the missing capability One channel cannot be moved independently of its workspace beside the circled shard.

The intended shortfall is now visible. Workspace-level routing concentrates the large tenant's writes on one storage boundary.

What the Simulation Proves

The circled shard is a hot partition in this hypothetical workload. Its pressure comes from the workspace being the indivisible sharding boundary.

The Slack-documented label identifies architecture facts from Slack Engineering. The Hypothetical simulation label keeps your invented workload separate from those historical facts.

Before the final check, which single box do you expect to receive all 500 message writes/s? Keep that prediction in mind while you inspect the panel.

  • Follow the numbered arrows from client API request to connected clients over WebSockets.
  • Recalculate 2,000 × 40 = 80,000 deliveries/s from the simulation inputs.
  • Confirm that the large workspace creates 500 message writes/s on the circled shard.
  • Confirm that the missing capability identifies one channel as inseparable from its workspace.
  • Confirm that the Flannel + Vitess panel remains empty.
  • Confirm that the Real-Time Fan-Out panel remains empty.
  • Confirm that the Failure-Isolated Cells panel remains empty.
  • Update the local copy from earlier using the same save-to-file method.

You should see the persisted message path followed by recipient fan-out. You should also see one indivisible shard carrying all 500 message writes/s from the simulated large workspace.

You found the first architectural breaking point. Next, you will separate startup pressure from storage pressure while preserving the evidence boundary you established here.

Break the Workspace Boundary

Your saved Excalidraw board now shows a hypothetical large workspace pinning 500 message writes/s to one indivisible shard in Slack's workspace model.

Growth creates another pressure. A reconnect storm forces many clients to bootstrap at once.

These pressures need separate boundaries. Flannel protects startup traffic at the edge.

Vitess gives MySQL a flexible sharding boundary beyond workspace-only application routing.

In this step, get ready to:
  • Place Flannel between clients and the core region.
  • Replace workspace-only routing with Vitess keyspaces.
  • Calculate the backend pressure from a reconnect storm.
Add Flannel at the edge

Flannel is Slack's application-level edge cache. It uses lazy loading to reduce the amount of data each client needs during startup.

Team affinity uses consistent hashing to direct related requests toward the same cache capacity.

  • Return to the empty Flannel + Vitess panel in the saved board.
  • Add the evidence label Slack-documented at the top of the panel.

The evidence label keeps this historical architecture separate from the simulation you add later.

  • Draw a client group labeled Clients on the left side of the panel.
  • Draw a large boundary labeled Core region on the right side of the panel.

You should see an open gap between the clients and the core region. That gap represents the edge.

  • Draw a box labeled Flannel inside the gap.

Flannel should now sit visibly outside the core region.

  • Connect the clients to Flannel with an arrow.
  • Connect Flannel to the core region with an arrow.

You should now see every client startup path crossing Flannel before reaching the core region.

  • Add lazy and just-in-time loading beneath Flannel.
  • Add team affinity through consistent hashing beside Flannel.

The annotations show how Flannel reduces startup work while preserving cache locality.

  • Add 4 million simultaneous connections at peak to the documented snapshot.
  • Add 600K client queries per second to the documented snapshot.

You should see both historical scale figures grouped with the Slack-documented evidence.

What does Flannel change?

Flannel reduces repeated startup traffic close to clients through lazy loading. The core region receives less bootstrap work during concentrated reconnects.

The connection count and query rate are documented historical snapshots. They do not describe Slack's current 2026 production capacity.

Replace workspace routing with Vitess

Flannel protects startup traffic. The hot workspace shard still needs a datastore boundary that can move data below the workspace level.

  • Draw a box labeled Webapp inside the core region.
  • Draw a datastore group labeled Vitess keyspaces downstream from Webapp.

You should now see Webapp leading into a separate datastore layer.

  • Connect Webapp to the Vitess keyspaces.
  • Label the datastore boundary messages table was sharded by channel ID.

The channel ID now defines the message boundary. One channel can be addressed without workspace-context routing.

  • Add application-managed workspace-only routing above Webapp.
  • Cross out the application-managed routing label.

The crossed-out label makes the architectural change visible. Flexible routing now belongs to the datastore layer.

Why are these boundaries separate?

Flannel handles startup pressure near clients. Vitess distributes database load through keyspaces and flexible sharding.

Each boundary removes a different bottleneck. Combining them into one box would hide the design decision.

  • Add 2.3 million QPS at peak to a documented database metric card.
  • Add 2M reads beneath the peak query rate.

The database card should now distinguish total query volume from read volume.

  • Add 300K writes beneath the read volume.
  • Add 2 ms median as the median query latency.

You should now see the documented workload split beside its median latency.

  • Add 11 ms p99 as the tail query latency.
  • Copy the reusable decision card into the Flannel + Vitess panel.

The panel now has a place to preserve the alternatives that Slack evaluated.

  • Record flexible sharding inside Webapp as too coupled and operationally incomplete.
  • Record duplicate shared-channel data across workspace shards as risking extra data and inconsistent histories.

The decision card should now show why the Vitess path won without making it look inevitable.

Simulate a reconnect storm

Historical scale proves that startup traffic matters. A separate hypothetical card lets you calculate the pressure without presenting invented figures as Slack production data.

  • Add a card titled Hypothetical simulation beside the Flannel diagram.
  • Add Invented for this exercise, not Slack production data beneath the title.

The card should now be visibly separated from the Slack-documented metrics.

  • Add 100,000 clients ÷ 60 seconds ≈ 1,667 reconnects/s to the card.
  • Add 2 MB as the invented bootstrap size for each client.

The first line measures reconnect frequency. The second line supplies the workload assumption needed for bandwidth.

  • Add 200 GB in one minute as the total transferred data.
  • Add about 3.33 GB/s as the average backend traffic.

You should now see the reconnect count turn into a sustained backend bandwidth requirement.

How does the reconnect calculation work?

  • Dividing 100,000 reconnecting clients across 60 seconds produces approximately 1,667 reconnects/s.
  • Multiplying 100,000 clients by the invented 2 MB bootstrap produces 200 GB in decimal units.
  • Dividing 200 GB by 60 seconds produces approximately 3.33 GB/s of average backend traffic.

Before you inspect the full panel, which pressure do you expect each new boundary to remove?

  • Trace the startup path from clients through Flannel into the core region.
  • Point to the channel-based sharding boundary inside the Vitess layer.
  • Recalculate the reconnect rate from the hypothetical card.
  • Recalculate the average backend traffic from the hypothetical card.

You should arrive at approximately 1,667 reconnects/s, 200 GB transferred in one minute, plus approximately 3.33 GB/s of average backend traffic.

  • Save the updated local copy of your Excalidraw board.

Strong work. Your second panel now separates edge startup pressure from datastore sharding pressure while keeping every invented workload visibly labeled.

Your board now explains how Slack moved beyond workspace-bound startup and storage assumptions. Next, you'll scale the real-time path across millions of persistent connections.

Scale the Real-Time Fan-Out

Your Excalidraw board now shows how Flannel protects client startup. Vitess gives Slack a flexible storage boundary.

Persistent WebSocket connections create a separate scaling problem. This stage maps the real-time services that turn each persisted message into multiplicative fan-out across connected recipients.

In this step, get ready to:
  • Map Slack's real-time services with their documented responsibilities.
  • Trace the persisted delivery path through the real-time stack.
  • Calculate message fan-out before exposing the remaining cross-AZ risk.
Map the real-time service responsibilities

Slack split real-time work across services with different ownership boundaries. Fitting every service into one panel is fiddly, so keep the delivery path in the middle of the panel.

  • Return to the empty Real-Time Fan-Out panel on your existing board.
  • Add the evidence label Slack-documented at the top of the panel.
  • Draw two client clusters on the left side of the panel with one cluster in each edge region.
  • Place Webapp to the right of the client clusters.
  • Place Vitess beside Webapp as the persisted storage layer.
  • Add separate boxes for AS, CS, GS, and PS.
  • Arrange AS near Webapp with CS at the center of the real-time stack.
  • Add CHARMs and Consul beside the CS layer.

How is real-time state divided?

  • A stateful service retains connection or channel context that later requests still need.
  • A stateless service can handle the next request without owning persistent request context.
  • An in-memory service keeps active routing or subscription data in memory for fast access.
  • Consistent hashing maps each channel ID to a Channel Server without storing one fixed routing entry for every channel.
  • Annotate AS with stateless and in-memory.
  • Annotate CS with stateful, in-memory, and hold channel history.
  • Annotate GS with stateful, in-memory, and WebSocket channel subscriptions.
  • Annotate PS with in-memory and track online users.
  • Connect CHARMs to the CS layer with a line labelled manage the Channel Server consistent hash ring.
  • Connect CS to Consul to show where Channel Servers register their current configuration.

You have cleared the hardest naming hurdle. Every real-time responsibility now has a visible home.

Trace the persisted delivery path

Message persistence protects durability before the real-time stack begins delivery. Your arrows need to preserve that ordering while showing how one channel reaches several Gateway Servers.

  • Draw the first numbered arrow from a client cluster to Webapp.
  • Draw the second numbered arrow from Webapp to Vitess with the label persisted storage.
  • Continue the third numbered arrow from the persisted-storage marker to AS.
  • Connect AS to CS with the fourth numbered arrow.
  • Label the fourth arrow channel ID through consistent hashing.
  • Draw the fifth numbered arrow from CS to every subscribed GS instance.
  • Draw the sixth numbered arrow from each GS to its recipient clients.

Why does persistence come first?

Persisting the message gives Slack a durable record before the real-time stack distributes it. The delivery system can now focus on routing the saved message to active subscribers.

AS accepts the handoff from Webapp. CS resolves the channel before each subscribed GS delivers the message to its connected clients.

  • Place about 16 million channels are served per host beside CS.
  • Place under 20 seconds beside the CS replacement note.
  • Place tens of millions of connected clients beside the client clusters.
  • Place 500ms beside the global delivery path.

The documented path now separates durable storage from connection routing. It also shows where one message starts multiplying across subscribers.

Calculate fan-out and expose the cross-AZ risk

Fan-out multiplies the incoming message rate by the average number of online recipients. A separate hypothetical card keeps that capacity exercise distinct from Slack's documented measurements.

  • Add a calculation card beneath the real-time service diagram.
  • Label the card Hypothetical simulation.
  • Enter 5 million concurrent WebSocket connections, 10,000 messages/s, and 50 average online recipients as the three assumptions.
  • Write 10,000 × 50 = 500,000 deliveries/s as the fan-out result.

Before you trace a gray failure, do you expect a degraded cross-AZ dependency to remain confined to one Availability Zone?

  • Trace the hypothetical failure by drawing an arrow across an Availability Zone boundary toward an otherwise healthy frontend.

The failure path reaches a healthy frontend through its cross-AZ dependency. This is the intended shortfall in the current design.

  • Label the arrow cross-AZ dependencies can let a gray failure affect otherwise healthy frontends.

Before you verify the stage, can the inbound message rate alone describe the work performed by the delivery stack?

  • Follow the numbered path aloud from the client through persisted storage to every recipient client.
  • Recalculate the hypothetical delivery rate from the two fan-out inputs.
  • Inspect the cross-AZ failure arrow to identify the healthy frontend exposed by the degraded dependency.

You should reach recipient clients only after the message passes through persisted storage, AS, CS, and subscribed GS instances. Your calculation should produce 500,000 deliveries/s.

Path hard to follow?

  • Check that every delivery arrow follows its number from the client to recipient clients.
  • Move the presence path below the delivery path if the lines overlap.
  • Use one colour for documented delivery arrows and a different colour for the hypothetical failure arrow.

Still stuck? Help me simplify my Slack real-time fan-out diagram without losing the documented service roles.

  • Save an updated local copy of the Excalidraw board using the same workflow from earlier.
  • Confirm the updated board file is visible in your browser's downloads list.

Strong work. Your third panel now connects service ownership, persistent delivery, multiplicative fan-out, and an unresolved failure path.

Your real-time architecture can now scale connections while making its cross-AZ weakness visible. Next, you'll contain that failure path inside isolated cells.

Contain Failure with Cells

Your Real-Time Fan-Out panel now shows how Slack persists messages before delivering them at scale. Its cross-AZ arrow exposes a remaining risk: an ambiguous network fault can affect otherwise healthy frontends.

You'll use a cellular architecture on your Excalidraw board to contain that impact inside one cell. The final panel connects failure isolation to a measurable recovery plan.

In this step, get ready to:
  • Build three Availability Zone cells with same-cell service dependencies.
  • Model weighted traffic draining around an unhealthy cell.
  • Export an interview-ready board with its failover capacity analysis.
Build three siloed cells

A cell is a complete service slice contained within one Availability Zone. Siloing keeps each slice independent during a failure.

  • Return to the empty Failure-Isolated Cells panel.
  • Draw three large cell boundaries across the panel.

You should see three distinct spaces with enough room for a replicated service stack inside each one.

  • Label the three boundaries as separate Availability Zone cells.
  • Duplicate the same user-facing service stack inside each cell.

Each cell should now look capable of serving users without borrowing a service from another cell.

  • Draw every service-to-service arrow inside its own cell boundary.
  • Add the Slack-documented evidence label beside the three cells.

Why Silo the Dependencies?

Same-cell service paths prevent a failing Availability Zone from becoming a dependency for healthy cells. The boundary turns one broad blast radius into a contained failure domain.

You should see no service-to-service arrow crossing from one cell into another. That visual boundary is the core reliability decision in this stage.

Route traffic around one failed cell

Slack uses Envoy for edge load balancing. Rotor supplies the weighted cluster configuration that controls new traffic assignments.

  • Place the Envoy edge load balancers outside the three cell boundaries.
  • Place Rotor beside Envoy outside the cell boundaries.

Envoy and Rotor should sit above the isolated stacks because they coordinate traffic across all three cells.

  • Draw control arrows from Rotor to Envoy.
  • Draw weighted cluster arrows from Envoy to each cell.

How Weighted Draining Works

Rotor updates Envoy's per-cell cluster weights. A zero weight stops new assignments to the unhealthy cell while in-flight requests continue.

  • Mark one of the three cells as unhealthy.
  • Set the unhealthy cell's cluster weight to zero.

New requests should now flow only to the two healthy cells. The unhealthy cell remains visible as the failure being contained.

Documented Drain Goals

  • Remove as much traffic as possible within 5 minutes.
  • Avoid user-visible errors during the drain.
  • Restore traffic in increments as small as 1%.
  • Add the three documented drain goals beside the weighted clusters.
  • Tag the drain-goals note with Slack-documented.

You should now see a centralized control path that directs new traffic away from one unhealthy cell. The two healthy cells keep receiving requests without gaining cross-cell service dependencies.

Finish the failure-isolation case study

Failure isolation needs spare capacity. Headroom measures how much additional traffic each healthy cell can absorb after one cell drains.

  • Add a Hypothetical simulation capacity card beneath the three cells.
  • Label the card Invented for this exercise, not Slack production data.

The evidence boundary should remain visible before anyone reads the capacity figures.

  • Add 900,000 requests/min ÷ 3 cells = 300,000 requests/min per cell as the baseline calculation.
  • Add 900,000 ÷ 2 = 450,000 requests/min as the drained-state calculation for each surviving cell.

The card should show each survivor rising from 300,000 requests/min to 450,000 requests/min.

  • Add the difference as 450,000 - 300,000 = 150,000 additional requests/min.
  • Label that increase as 50% headroom.

The completed card now turns the drain strategy into a concrete capacity requirement.

Complete the Final Decision Record

  • Trigger: A gray failure crosses Availability Zone dependencies.
  • Bottleneck: Healthy frontends still depend on degraded cross-AZ paths.
  • Decision: Replicate the user-facing stack into siloed cells.
  • Trade-off: Duplicated capacity increases operational complexity.
  • Rejected alternative: Keep cross-AZ service dependencies while recovering each service independently.
  • Evidence: Slack documented a five-minute drain target with weighted traffic control.
  • Fill the reusable decision card using the six entries above.
  • Add a text box titled Five-minute interview talk track beside the board.

Shape the Five-Minute Story

  • Minute 1 covers workspace-level routing plus the indivisible hot shard.
  • Minute 2 explains how Flannel protects startup while Vitess breaks workspace-only sharding. It rejects flexible sharding inside Webapp plus shared-channel duplication.
  • Minute 3 traces persistence through the real-time stack. It connects message rate to multiplicative fan-out.
  • Minute 4 introduces the cross-AZ gray failure. It explains how same-cell dependencies contain the impact.
  • Minute 5 closes with the documented drain target plus the hypothetical 50% survivor headroom. It names duplicated capacity plus operational complexity as the cost.
  • Write one concise prompt for each minute in the talk-track text box.

Your board should now read as one connected argument from product growth to failure containment. Each transition has a trigger plus a measurable consequence.

  • Save the updated Excalidraw board to the same local file.

Before you preview the export, which labels do you expect to be hardest to read? The export preview tests that prediction.

  • Open Excalidraw's export controls to preview the whole canvas.

You should see all four panels inside the export area. The evidence labels plus the five-minute talk track should remain readable.

  • Select PNG as the export format.
  • Export the completed architecture evolution board to your computer.

You should now have a PNG containing all four architecture stages. It should include the decision records plus every visibly labeled simulation.

Before you inspect the file, do you expect the documented evidence to remain visually distinct from the hypothetical calculations? The exported image reveals whether the evidence boundary survived.

  • Open the exported PNG from your browser's downloads list.

You should see the workspace-shard baseline, Flannel with Vitess, the real-time fan-out stack, plus three failure-isolated cells. The final panel should show one zero-weight cell beside the 50% headroom calculation.

Is the Export Cropped or Unreadable?

  • Return to the export preview if one of the four panels sits outside the captured area.
  • Increase the size of labels that become unreadable at normal image zoom.
  • Export the PNG again after the full board fits inside the preview.

Still stuck? Help me fix my Excalidraw PNG export.

Before you start the rehearsal, which architecture transition do you expect to take longest to defend? Your timed walkthrough tests that prediction.

  • Start a five-minute timer.
  • Present the four panels from left to right using the talk-track prompts.
  • Close with the documented 5 minutes cell-drain target.
  • State that each surviving cell absorbs a hypothetical 50% traffic increase.

You should finish in about five minutes. Your narrative should trace the bottleneck that forces each architectural change.

The final scenario should drain one cell within the documented 5 minutes target. Each surviving cell should absorb the hypothetical 50% increase.

That's the full case study complete. Your exported board now connects Slack's architecture history to capacity calculations plus a defensible reliability decision.

Secret mission

Design the Org-Wide Tenancy Layer

Slack's cells contain failures within infrastructure boundaries. This challenge adds an org-wide tenancy layer that boots cross-workspace data while preserving clear permission boundaries.

Clean Up Your Resources

Clean Up Your Resources

Your Excalidraw board and exported PNG live on your device. No cloud resources are running, so this project has no ongoing costs.

Choose whether to keep the case study, close the active canvas for now, or delete every copy.

Resources you used:

  • The current Excalidraw browser canvas containing the four-stage architecture board plus its org-wide tenancy extension.
  • The locally saved board titled Evolving Slack: From Workspace Shards to Failure-Isolated Cells.
  • The exported PNG of the completed architecture evolution board.

Keep everything running

No action needed. Choose this if you want to keep practicing your five-minute architecture walkthrough.

  • Keep the current Excalidraw canvas available in your browser.
  • Keep the locally saved board file as your editable copy.
  • Keep the exported PNG as your shareable copy.

Your board remains ready for more design iterations without creating ongoing costs.

Pause - I'll come back to this later

There is no running process or cloud service to pause. Closing the browser tab is enough to end the active session.

  • Close the current Excalidraw browser tab.
  • Keep the locally saved board file on your device.
  • Keep the exported PNG on your device.

The saved board preserves your editable architecture case study for your next practice session.

Delete - I don't want to use this again

Deleting these copies is permanent. The steps below remove the browser canvas plus both locally saved artefacts.

  • Switch back to the Excalidraw tab from earlier.
  • Click once on the canvas.
  • Press Cmd+A on macOS or Ctrl+A on Windows to select every object.
  • Press Backspace or Delete to remove the selected objects.
  • Confirm that the canvas is blank.
  • Close the Excalidraw browser tab.

The browser copy is now empty. Your locally saved board and exported PNG still remain on your device.

  • Press Cmd+Space on macOS or the Windows key on Windows to open system search.
  • Type Finder on macOS or File Explorer on Windows into the search bar.
  • Press Enter to open the selected file browser.

The remaining cleanup removes the two saved files from the locations you chose earlier.

  • Use the file browser search field to locate the locally saved Excalidraw board.
  • Move the board file to Trash on macOS or the Recycle Bin on Windows.
  • Use the file browser search field to locate the exported PNG.
  • Move the PNG to Trash on macOS or the Recycle Bin on Windows.
  • Empty Trash on macOS or the Recycle Bin on Windows.
  • Confirm that neither saved file appears in the file browser search results.

That removes every project artefact from your device. No further cleanup is required.

Nice Work!

Nice Work!

You did it! Your exported Excalidraw board now explains how Slack evolves from workspace shards to failure-isolated cells.

You've learned how to:

  • Build a four-stage architecture evolution board that moves from workspace shards to failure-isolated cells. Each stage connects a product pressure to a visible system change.
  • Keep source-backed evidence separate from hypothetical simulations. Use back-of-the-envelope calculations to quantify shard pressure. Apply the same method to reconnect traffic. Extend it to message fan-out. Use it again to calculate failover headroom.
  • Create decision records that capture bottlenecks. Defend trade-offs through rejected alternatives. Present the full evolution as a five-minute interview narrative.
  • Complete the Secret Mission by designing an org-wide tenancy layer. The extension models cross-workspace boot aggregation. It adds org-level permissions. It defends an incremental migration path.

Ready to quiz yourself?