Files
Baradb/docs/superpowers/specs/2026-07-30-raft-network-bootstrap-design.md
T
dimgigov 16ec8b5dc4
CI / test (push) Has been cancelled
CI / verify (push) Has been cancelled
docs: raft cluster status, plans closed, CHANGELOG/README/monitoring
- Add raft-cluster-status overview (C3a/C3b/post-C3b shipped on main)
- Mark C3a/C3b design+plans done; refresh operator docs en/bg
- CHANGELOG 1.2.0 Raft section; README cluster example and status line
- monitoring.md health/metrics match real HTTP port+440 and raft series
2026-07-30 21:41:15 +03:00

35 lines
1.2 KiB
Markdown

# Networked Raft Bootstrap (C3a) — Design
Date: 2026-07-30
Status: **Done** (merged to `main`). See also `2026-07-30-raft-cluster-status.md`.
## Problem
Raft in BaraDB was half-wired for real networking. `core/raft.nim` already had
a TCP transport, but production never populated `peerAddrs`, never ran an
election timer, did not reset timers on AppendEntries, skipped state
persistence, and used unsafe partial frame reads.
## Goal
A 3-node BaraDB cluster elects a leader over TCP, maintains it with heartbeats,
and re-elects after the leader is killed — verified with real processes
(`tests/raft_e2e_test.nim`).
## Delivered
| Item | Implementation |
|------|----------------|
| Peer addresses | `BARADB_RAFT_PEERS=id@host:port``raftPeerAddrs` |
| Election timer | `timerLoop` inside `RaftNetwork.run` |
| Heartbeat reset | `processMessage` resets timer on AppendEntries with valid term |
| State persistence | `dataDir``raft_state.bin` |
| Partial reads | `recvExact` frame reassembly |
| E2E | `tests/raft_e2e_test.nim` (election + failover) |
## Non-goals (handled later)
SQL writes (C3b), DDL/forward/compact/metrics (post-C3b / C3c-lite) — all
shipped; see cluster status doc. Still open: membership, InstallSnapshot SM
payload, multi-DB raft.