04 · ENGINE

Clustering

XERJ ships an embedded Raft implementation — no etcd, no ZooKeeper, no Consul. One binary is still one binary; cluster mode is a config switch, not a separate process. Where the ring stands today, plainly: leader election, heartbeats, and the authenticated inter-node transport are real and tested; nothing in the production write path rides on the log yet. Indices, mappings, and documents are node-local — a document indexed on one node is not visible on another until the replication work lands (public tracking: xerj-org/xerj#1170). If you need a second live copy of your data today, use the WAL tap or per-node backups.

What cluster mode does not give you yet

No metadata replicationAn index created on node A does not exist on node B — no production path proposes a ClusterCommand yet.
No document replicationPUT a document on node A, GET it from node B: 404 index_not_found. Measured on a 3-node ring.
No shard routingNo forwarding to a primary, no replica reads. Each node's API serves that node's storage only.
No node-lifecycle APIsThere is no drain / activate / remove endpoint and no region split-merge. Earlier revisions of this page documented them; the binary has never shipped them.

Treat a clustered node's data durability exactly like a single node's.

What you get

Bootstrapping a 3-node cluster

Start three nodes, each with its own data dir, and point them at each other. On first startup they hold a Raft election and one becomes leader:

# /etc/xerj/node-a.toml  (run on 10.0.0.11)
[server]
data_dir     = "/var/lib/xerj"
bind_address = "10.0.0.11"

[cluster]
enabled = true
port    = 9300                            # intra-cluster Raft + search port
peers   = [
  "a=10.0.0.11:9300",                     # this node
  "b=10.0.0.12:9300",
  "c=10.0.0.13:9300",
]
tick_ms = 50                              # Raft tick interval
# /etc/xerj/node-b.toml  (run on 10.0.0.12)
[server]
data_dir     = "/var/lib/xerj"
bind_address = "10.0.0.12"

[cluster]
enabled = true
port    = 9300
peers   = [
  "a=10.0.0.11:9300",
  "b=10.0.0.12:9300",                     # this node
  "c=10.0.0.13:9300",
]
tick_ms = 50

The peers list uses "<node_id>=<host>:<port>". The current node's id is inferred from the entry whose address matches its own bind_address:port. Start all three boxes:

$ xerj --config /etc/xerj/node-a.toml     # on 10.0.0.11
$ xerj --config /etc/xerj/node-b.toml     # on 10.0.0.12
$ xerj --config /etc/xerj/node-c.toml     # on 10.0.0.13

Within a second the ring elects a leader. The honest place to watch it today is the log — the health endpoints report the node answering the request, not the ring's membership (that plumbing is part of the wiring work):

$ journalctl -u xerj -f | grep xerj_cluster
INFO xerj_cluster::transport: TCP transport listening (authenticated) node=10.0.0.11:9300
INFO xerj_cluster::raft: Starting election node=10.0.0.12:9300 term=1
INFO xerj_cluster::raft: Became leader node=10.0.0.12:9300 term=1
$ curl -sH "Authorization: ApiKey $XERJ_API_KEY" \
    http://10.0.0.11:9200/_cluster/health | jq '.status, .number_of_nodes'
"green"
1

Indexes and data are node-local today

Index creation, mappings, and documents do not cross the wire — every node's API serves that node's storage only. There is no shard assignment to other members, no replica set, and no drain/region API; the sections that used to document them were removed because the binary has never shipped them.

For a second live copy of your data, point the WAL tap at another XERJ or any ES-compatible target: it ships indexed writes near-real-time from the WAL, one-directional, without touching the write path.

Failure modes

Leader crashFollowers time out and elect a new leader — measured ~250 ms on a 3-node loopback ring, including one split-vote retry. The ring stays stable with the dead member absent; its Raft log survives on disk for the rejoin.
Hung or dead peerOne WARN on the down transition, then quiet: retries back off 100 ms → 6.4 s, and messages to live peers are never delayed behind a dead one. A peer that answers again logs "peer reachable again" at INFO.
Minority partitionThe minority side cannot reach quorum and its elections never conclude; the majority side keeps its leader. Raft safety holds — at most one leader per term.
Node rejoinRestart with the same config: the node replays its local Raft log and re-enters the election. Catch-up from the leader's log (snapshot install) is part of the wiring roadmap.

When not to cluster

Today, almost always. Until the replication wiring lands, cluster mode gives you a leader-election ring — not copies of your data. One node avoids network round-trips, needs no quorum, and is trivial to back up; add a WAL tap to a second node if you want a live spare. The default config ships with [cluster] enabled = false.

Source · engine/crates/xerj-cluster/src/raft.rs · engine/crates/xerj-cluster/src/transport.rs · engine/crates/xerj-cluster/src/runner.rs