MIT 6.824 Consensus & Partition Lab

Fault Tolerance & Primary-Backup Replication Workbench (Prof. Robert Morris Lecture 1)

QUORUM: 3/5 HEALTHY
Cluster Network Topology & In-Flight RPC Packets
Click a node to crash/revive. Click a link to cut/heal.
Status: Cluster Operating (Quorum 5/5 Available) Invariant: Majority = ≥ 3 Nodes
Client RPC Console (Linearizable State Machine)
State: x = 42

Submit a transactional state command. The client issues an RPC to Primary (S1). S1 must replicate to a majority quorum (≥3 nodes) before replying 'OK' and committing.

Lecture 1 Failure Scenarios

1. Minority Partition

Cut S4 & S5. Majority (3/5) remains. Cluster continues committing linearizable writes.

2. Split-Brain Prevention

Isolate S1,S2 from S3,S4,S5. Old primary has only 2 nodes < 3; writes stall safely to avoid inconsistency.

3. Leader Crash & Re-election

Kill S1. Remaining nodes elect highest term node as new leader to restore liveness.

Replicated State Machine Logs
Inspection per Node
Node Role Status Log Entries (idx:term:cmd) Commit Index Local State
Live RPC Event Stream
MIT 6.824 Core Principles

Fault Tolerance: As Prof. Morris emphasizes, the objective is to hide failures from clients so the service behaves like a single centralized computer.

Quorum Invariant: Any two majorities in a cluster of size 2F + 1 must overlap by at least one node:
Q1 ∩ Q2 ≠ ∅
For 5 nodes, quorum is 3. A partition of 2 nodes cannot commit, preventing divergent split-brain histories.

Safety vs Liveness: In a network split, safety demands rejecting or stalling requests in the minority partition rather than risking corrupted state.