32장. 복제·Failover 게임데이
정상 상태에서 primary/standby role, replay LSN, lag byte/time, slot retained WAL, archive success를 저장한다. 그런 다음 primary process를 중지하고 detection, election, fencing, promotion, proxy/DNS 전환, client reconnect를 측정한다.
가장 중요한 것은 old primary fencing이다. network partition이 풀렸을 때 두 primary가 write를 받으면 split brain이 된다. automation은 새 leader를 선언하기 전 old leader의 storage·network·process write 가능성을 차단해야 한다.
failover 후 read-after-write, sequence, connection pool DNS cache, prepared statement, long transaction을 확인한다. 유실 가능 transaction은 audit/outbox 기록과 대사한다. old primary를 즉시 재합류시키지 말고 rewind/rebuild 조건을 확인한다.