03 · OPERATIONS

Upgrades

XERJ follows SemVer. Patch releases (0.1.x) are drop-in replacements. Minor releases (0.x.0) can introduce backward-compatible schema fields. Major releases (x.0.0) may change the on-disk format — when they do, the binary ships a one-shot migration tool.

Single-node upgrade

# 1. Back up (see Backup & restore)
$ curl -sXPOST -H "Authorization: ApiKey $XERJ_API_KEY" \
    http://127.0.0.1:8080/v1/admin/backup \
    -d '{"destination":"s3://my-backups/xerj/pre-upgrade"}'

# 2. Fetch and verify the new release — extraction is chained to the checksum,
#    so an archive that fails to verify is never unpacked over the running node
$ (
    ver=1.0.0-rc.18                             # the release you are upgrading to
    target=x86_64-unknown-linux-musl            # your platform's target triple
    stage="xerj-${ver}-${target}"
    asset="${stage}.tar.gz"
    base="https://github.com/xerj-org/xerj/releases/download/v${ver}"

    curl -fsSLO "$base/$asset"
    curl -fsSLO "$base/$asset.sha256"

    want=$( { sha256sum "$asset" 2>/dev/null || shasum -a 256 "$asset"; } | cut -d ' ' -f 1 )
    printf %s "$want" | LC_ALL=C grep -qE '^[0-9a-f]{64}$' \
      || { echo "no working SHA-256 tool — refusing to install an unverified $asset" >&2; exit 1; }

    tr -d '\r' < "$asset.sha256" \
      | LC_ALL=C grep -qxF -e "$want  $asset" -e "$want *$asset" \
      || { echo "CHECKSUM MISMATCH for $asset — do not install it" >&2; exit 1; }

    tar xzf "$asset"
  )

# 3. Replace the binary in place (reached only when step 2 verified and unpacked)
$ sudo systemctl stop xerj
$ sudo install -m 0755 "xerj-1.0.0-rc.18-x86_64-unknown-linux-musl/xerj" /usr/local/bin/xerj
$ sudo systemctl start xerj

# 4. Confirm the node reports the version you just installed
$ curl -s http://127.0.0.1:8080/v1/health | jq .version
"1.0.0-rc.18"

Downtime is the time it takes the WAL to replay — usually a few seconds for workloads in the low millions of live docs.

Cluster upgrade

The cluster wire protocol has no version negotiation: a node on a new wire version cannot talk to a node on the old one, in either direction. Upgrading a cluster is a full stop/start of all nodes in one maintenance window, not a rolling restart — there is no drain/activate API and no shard migration to sequence around.

# On every node, within the same window:
$ ssh node-<id> sudo systemctl stop xerj
$ ssh node-<id> sudo install -m 0755 xerj-new /usr/local/bin/xerj
$ ssh node-<id> sudo systemctl start xerj
# The ring re-elects a leader on its own once a quorum is back up

Because index data is node-local today, each node's data directory is also each node's upgrade responsibility — snapshot per the backup guide before touching the binary.

On-disk format migrations

A major release that changes the segment format ships a migrate subcommand. Run it after stopping the old binary and before starting the new one:

$ sudo systemctl stop xerj
$ xerj migrate --from 1.x --to 2.x --data-dir /var/lib/xerj
Scanning segments... 842 found
Rewriting segment 1/842: logs/seg-000123 ... ok
...
Migration complete. Verify with: xerj verify --data-dir /var/lib/xerj
$ sudo systemctl start xerj

The migrator is idempotent — if it crashes halfway, re-run it. It never touches source segments until the replacement is fully written and fsynced.

Rolling back

Patch releases roll back by reinstalling the previous binary. Across a format migration, rolling back requires the pre-migration backup — there is no inverse migrator.

Source · engine/crates/xerj-server/src/main.rs