Network Partition

Primary loses DCS connectivity, causing lease expiration and triggering split-brain protection and failover

RTO Timeline


Failure Model

PhaseBestWorstAverageNotes
Demoteretryloop + retryloop/2 + retryPatroni retries after detecting partition, demotes after timeout
Lease Expirationttl - loop - retryttl - loop - retryttl - loop - retryRemaining TTL time after demotion (approximately constant)
Replica Detection0looploop/2Best: Right at detection point
Worst: Just missed detection
Lock & Promote021Best: Direct lock and promote
Worst: API timeout + Promote
Health Check(rise-1) × fastinter(rise-1) × fastinter + inter(rise-1) × fastinter + inter/2Best: State changes before check
Worst: State changes right after check

Key difference between network partition and node crash:

ScenarioPatroni StatePostgreSQL StateLease HandlingSplit-brain Risk
Node Crash (Expire)Dies with nodeCompletely unavailablePassive wait for TTL expirationNone
Network Partition (This scenario)Alive but cannot access DCSMay still be running (needs active demotion)Passive wait for TTL expirationYes, needs protection

In network partition scenarios, the primary PostgreSQL may still be running and accepting writes, causing split-brain issues. Patroni solves this through active demotion: when unable to refresh Leader Key, proactively demotes PostgreSQL to read-only or shuts it down.


Timeline Analysis

Phase 1: Primary Demotion

When primary Patroni is network-partitioned from DCS, it cannot refresh Leader Key and starts retrying.

Timeline:
  Partition      Detect partition      Retry timeout      Primary demotes
     |               |                    |                    |
     |←── loop ──→|←── retry ──→|
  • Detection delay: After partition occurs, must wait for next loop_wait cycle to detect
  • Retry phase: Patroni continuously retries DCS operations during retry_timeout
  • Active demotion: After retry timeout, Patroni proactively demotes PostgreSQL (prevents split-brain)
Tdemote={retrybest (partition right before detection)loop/2+retryaverageloop+retryworst (partition right after refresh)T_{demote} = \begin{cases} retry & \text{best (partition right before detection)} \\ loop/2 + retry & \text{average} \\ loop + retry & \text{worst (partition right after refresh)} \end{cases}

Key design: Patroni requires constraint loop_wait + 2 × retry_timeout ≤ ttl to ensure primary demotes before TTL expires.

Phase 2: Lease Expiration

After primary demotion, Leader Key still exists in DCS, must wait for TTL to naturally expire.

Timeline:
  Primary demoted                   TTL expires
     |                                 |
     |←── ttl - (loop + retry) ──→|

Since the primary has demoted, waiting time during this phase is the remaining TTL time. Since partition detection and remaining TTL are negatively correlated (earlier partition means slower detection but longer remaining TTL), their sum is constant:

Texpire=ttlloopretry(approximately constant)T_{expire} = ttl - loop - retry \quad \text{(approximately constant)}

Note: Primary demotion + lease expiration total time still approximately equals ttl, same as expire failure.

Phase 3: Replica Detection

Replica wakes up in loop_wait cycle and checks Leader Key status in DCS.

Timeline:
    Lease expired      Replica wakes
       |                  |
       |←── 0~loop ─→|
  • Best case: Replica wakes right when lease expires, wait 0
  • Worst case: Replica just entered sleep when lease expires, wait loop
  • Average case: loop/2
Tdetect={0bestloop/2averageloopworstT_{detect} = \begin{cases} 0 & \text{best} \\ loop/2 & \text{average} \\ loop & \text{worst} \end{cases}

Phase 4: Lock & Promote

After replica discovers Leader Key expired, it starts the election process.

Election flow:
  ReplicaA ──→ Query replication position ──→ Compare ──→ Try lock ──→ Success
  ReplicaB ──→ Query replication position ──→ Compare ──→ Try lock ──→ Fail
  • Best case: Single replica or directly acquires lock and promotes, ≈ 0
  • Worst case: DCS API call timeout, 2s
  • Average case: 1s
Telect={0best1average2worstT_{elect} = \begin{cases} 0 & \text{best} \\ 1 & \text{average} \\ 2 & \text{worst} \end{cases}

Phase 5: Health Check

HAProxy detects new primary coming online, requires rise consecutive successful health checks.

Detection timeline:
  New primary    First check    Second check   Third check (UP)
     |              |               |               |
     |←─ 0~inter ─→|←─ fast ─→|←─ fast ─→|
  • Best case: (rise-1) × fastinter
  • Worst case: (rise-1) × fastinter + inter
  • Average case: (rise-1) × fastinter + inter/2
Thaproxy={(rise1)×fastinterbest(rise1)×fastinter+inter/2average(rise1)×fastinter+interworstT_{haproxy} = \begin{cases} (rise-1) \times fastinter & \text{best} \\ (rise-1) \times fastinter + inter/2 & \text{average} \\ (rise-1) \times fastinter + inter & \text{worst} \end{cases}

RTO Formula

Sum all phase times to get total RTO.

Since primary demotion + lease expiration ≈ ttl, network partition RTO formula is same as expire failure:

Best Case

RTOmin=ttlloop+0.1+(rise1)×fastinterRTO_{min} = ttl - loop + 0.1 + (rise-1) \times fastinter$$RTO_{min} \approx ttl - loop + (rise-1) \times fastinter$$

Average Case

RTOavg=ttl+1+inter/2+(rise1)×fastinterRTO_{avg} = ttl + 1 + inter/2 + (rise-1) \times fastinter$$RTO_{avg} = ttl + 1 + inter/2 + (rise-1) \times fastinter$$

Worst Case

RTOmax=ttl+loop+2+inter+(rise1)×fastinterRTO_{max} = ttl + loop + 2 + inter + (rise-1) \times fastinter$$RTO_{max} = ttl + loop + 2 + inter + (rise-1) \times fastinter$$

Model Calculation

Substituting the four RTO model parameters into the formulas:

pg_rto_plan:  # [ttl, loop, retry, start, margin, inter, fastinter, downinter, rise, fall]
  fast: [ 20  ,5  ,5  ,15 ,5  ,'1s' ,'0.5s' ,'1s' ,3 ,3 ]  # rto < 30s
  norm: [ 30  ,5  ,10 ,25 ,5  ,'2s' ,'1s'   ,'2s' ,3 ,3 ]  # rto < 45s
  safe: [ 60  ,10 ,20 ,45 ,10 ,'3s' ,'1.5s' ,'3s' ,3 ,3 ]  # rto < 90s
  wide: [ 120 ,20 ,30 ,95 ,15 ,'4s' ,'2s'   ,'4s' ,3 ,3 ]  # rto < 150s

Patroni constraint validation (loop + 2×retry ≤ ttl):

ModeloopretryTTLloop + 2×retryMeets constraint?
fast5520s15s✓ Safe
norm51030s25s✓ Safe
safe102060s50s✓ Safe
wide2030120s80s✓ Safe

Four mode calculation results (seconds, format: min / avg / max)

Phasefastnormsafewide
Primary Demote5 / 8 / 1010 / 13 / 1520 / 25 / 3030 / 40 / 50
Lease Expiration10153070
Replica Detection0 / 3 / 50 / 3 / 50 / 5 / 100 / 10 / 20
Lock & Promote0 / 1 / 20 / 1 / 20 / 1 / 20 / 1 / 2
Health Check1 / 2 / 22 / 3 / 43 / 5 / 64 / 6 / 8
Total16 / 23 / 2927 / 34 / 4153 / 66 / 78104 / 127 / 150

Conclusion: Network partition RTO is same as expire failure (node crash), as the bottleneck is TTL expiration time.


Split-brain Protection

The biggest risk of network partition is split-brain: old primary may still be running and accepting writes. Patroni provides multiple protection mechanisms:

1. Primary Self-Demotion

Patroni’s core protection mechanism: when unable to refresh Leader Key, proactively demotes PostgreSQL.

# Patroni pseudo-code logic
if not can_refresh_leader_key():
    retry_until(retry_timeout)
    if still_cannot_refresh():
        demote_postgresql()  # Demote to read-only or shut down

2. Linux Watchdog

If Patroni process hangs and cannot execute demotion, Linux watchdog will force system restart.

# patroni.yml configuration
watchdog:
  mode: required  # Require watchdog available
  device: /dev/watchdog
  safety_margin: 5

3. Fencing Mechanism

Can configure fencing scripts to forcibly isolate old primary (e.g., disable network interface, stop service, etc.).


Special Scenarios

Scenario A: Primary partitioned from DCS, replicas normal

This is the most common network partition scenario, the main focus of this article.

┌─────────┐         ╳         ┌─────────┐
│ Primary │ ←── Partition ──→ │   DCS   │
│ Patroni │                   │  etcd   │
└─────────┘                   └─────────┘
                                  ↑
                              Normal connection
                                  ↓
                              ┌─────────┐
                              │ Replica │
                              │ Patroni │
                              └─────────┘
  • Primary Patroni cannot refresh Leader Key → Active demotion
  • Replica normally detects TTL expiration → Elected as new primary
  • RTO ≈ Expire failure RTO

Scenario B: Primary normal, replica partitioned from DCS

┌─────────┐                   ┌─────────┐
│ Primary │ ←── Normal ──→    │   DCS   │
│ Patroni │                   │  etcd   │
└─────────┘                   └─────────┘
                                  ╳
                              Partition
                                  ╳
                              ┌─────────┐
                              │ Replica │
                              │ Patroni │
                              └─────────┘
  • Primary normally refreshes Leader Key
  • Replica cannot participate in election (but replication can continue)
  • No failover triggered, service continues normally

Scenario C: All nodes partitioned from DCS

┌─────────┐         ╳         ┌─────────┐
│ Primary │ ←── Partition ──→ │   DCS   │
│ Patroni │                   │  etcd   │
└─────────┘                   └─────────┘
                                  ╳
┌─────────┐         ╳             │
│ Replica │ ←── Partition ────────┘
│ Patroni │
└─────────┘
  • Primary demotes, replica cannot elect
  • Cluster completely unavailable
  • Requires manual intervention to restore DCS connectivity

Comparison with Other Failures

Failure TypePrimary StateLease HandlingRTOSplit-brain Risk
Expire FailureNode crashPassive wait TTL expiration16s ~ 150sNone
Crash FailurePG crash, Patroni aliveRelease after restart timeout1s ~ 111sNone
Network PartitionAlive but isolated from DCSPassive wait TTL expiration16s ~ 150sYes, needs protection
Manual SwitchoverNormal or failedDirect release/acquire1s ~ 11sNone

Key Insight: Network partition RTO is same as expire failure, but requires additional split-brain protection mechanisms. Ensuring loop_wait + 2 × retry_timeout ≤ ttl constraint is the key design to prevent split-brain.