<!-- generated by build_articles.py — edit content/answers/check-codebase-index-is-complete.md instead -->
---
title: "How do I check my code index actually finished?"
canonical: "https://xerj.org/answers/check-codebase-index-is-complete"
updated: "2026-08-23"
source: "content/answers/check-codebase-index-is-complete.md"
---

# My codebase indexer says indexed but I don't see my code. How do I check it actually finished?

**TL;DR** — Reconcile 4 numbers before you trust an "indexed" status: files on disk, the `--dry-run` plan, `xerj autoindex status`, and the document count in `_cat/indices`. A captured XERJ run turned 64 non-hidden files into a 3-file plan, and every missing file had a named ignore rule.

## Agent prompt

```text
Act as a coding agent. Read https://xerj.org/llms.txt, count the files under the target folder, run xerj autoindex with --dry-run, run it for real, then reconcile the dry-run file count with xerj autoindex status and with a _cat/indices count before you report the codebase as indexed.
```

## Commands

### Command 1

Note: Print the plan and the ignore accounting without writing anything.

```sh
xerj autoindex ./my-repo --url http://127.0.0.1:9200 --prefix comp --state-dir ./state-comp --dry-run
```

### Command 2

Note: Index the folder and print the final document counts.

```sh
xerj autoindex ./my-repo --url http://127.0.0.1:9200 --prefix comp --state-dir ./state-comp
```

### Command 3

Note: Read the journal state and the live indices for this run.

```sh
xerj autoindex status --url http://127.0.0.1:9200 --state-dir ./state-comp
```

## The 4 numbers that must agree

A completeness check is arithmetic, not a status word. Count the files on disk, read the planned file count, read the journal state, then read the index count, and account for every difference.

| Number | Where it comes from | Captured value |
| --- | --- | --- |
| Files on disk | `find . -type f` | 65 total, 64 non-hidden |
| Files in the plan | `autoindex --dry-run` | 3 |
| Files done | `xerj autoindex status` | 3 files done, `FINISHED` |
| Documents live | `_cat/indices` | set by the emission model below, not by the file count |

The gap between 64 and 3 is the interesting one, and `autoindex` explains it rather than hiding it.

## Read the plan before the run

A dry run prints the plan and the ignore accounting, and writes nothing to the node. Run it first on any folder that later looks under-indexed.

```sh
xerj autoindex ./my-repo --url http://127.0.0.1:9200 --prefix comp --state-dir ./state-comp --dry-run
```

The captured dry run named every exclusion with an exact file count.

```text
autoindex: 3 files (0 MB) under .../fixtures/code
autoindex: ignore rules: skipped 2 files and pruned 3 directories (60 non-hidden files inside them); 1 ignore file read
autoindex: ignore rules:   <built-in>:target/ — 1 directory pruned (30 non-hidden files inside)
autoindex: ignore rules:   <built-in>:node_modules/ — 1 directory pruned (25 non-hidden files inside)
autoindex: ignore rules:   .gitignore:coverage/ — 1 directory pruned (5 non-hidden files inside)
```

3 planned files plus 2 skipped files plus 60 files inside pruned directories accounts for 65 paths. Every exclusion carries a rule name and a count.

## Documents outnumber files, and that is normal

XERJ writes one document per extracted row, and for most families that includes a whole-file document alongside the rows. Source code goes one step further. For each source file XERJ writes the file document, carrying the full text and a `defs` list of everything the file declares. It then writes one more document for every declaration it found, keyed by a `code:<line>:<name>` locator. That is what makes a constant or a one-line signature retrievable on its own, instead of only inside the class or method enclosing it. A declaration captured twice at the same line and name collapses to a single document.

The document count for a folder of source files therefore follows the number of declarations in it. No fixed ratio to the file count exists to carry over from another run. Read the count from `_cat/indices` for your own run, and reconcile it against the count `xerj autoindex status` prints.

State the convention whenever you quote a count, because a reader comparing a file count with a much larger document count will otherwise assume an error.

## Ask the journal, not your memory

`xerj autoindex status` reads the state directory the run wrote and prints both the journal state and the live indices.

```sh
xerj autoindex status --url http://127.0.0.1:9200 --state-dir ./state-comp
```

The journal prints as one line: the journal path, the root it indexed, how many files are done, how many documents were written, and either `FINISHED` or `in progress`. A run that wrote a graph adds an indented `graph:` line naming the edge count, the edges index and the brain. Below that, `status` lists every live index carrying your prefix with its document count, read from the node rather than from the journal.

Compare those two sides rather than trusting `FINISHED` on its own. The journal says what the run believed it wrote; the live index list says what the node actually holds.

## Finish with a query you can predict

A count proves quantity; a query proves retrievability. Pick a symbol you know exists in the tree and search the `defs` field for it.

```sh
curl -s -XPOST http://127.0.0.1:9200/comp-*/_search -H 'content-type: application/json' -d '{"query":{"match":{"defs":"merge_segments"}},"size":10,"_source":["ax_path","language","defs"],"track_total_hits":true}'
```

The captured response returned 2 hits, `src/lib.rs` in Rust and `src/ingest.py` in Python, each with its full `defs` list.

## Two size checks that fail on XERJ

An Elasticsearch user checking index size reaches for 2 forms that fail here. `GET /_cat/indices?format=json&bytes=b` still returned `58.5kb`, not a raw byte count. `GET /{index}/_stats/store` returned 404.

Use the plain statistics form instead, then read `_all.total.store.size_in_bytes`.

```sh
curl -s http://127.0.0.1:9200/comp-docs/_stats
```

## One more surprise in `_cat/indices`

`autoindex` creates one `.xerj-memory-<brain>-edges` index per indexed folder, named after the folder. A first look at `_cat/indices` therefore shows more indices than folders you indexed. XERJ is single-node, so all of them live in one process.

## FAQ

### My codebase indexer says indexed but I don't see my code. How do I check it actually finished?

Reconcile 4 numbers: files on disk, the dry-run plan, the autoindex status journal, and the index document count. A gap between the first two is always an ignore rule.

### How can I tell if folder indexing is still running or if it died?

Run xerj autoindex status with the same --state-dir. The journal reports files done, records and its state; the captured output read 3 files done and FINISHED.

### How do I verify a local code index actually contains my files?

Query for a symbol you know exists. The captured query on the defs field for merge_segments returned 2 hits, src/lib.rs in Rust and src/ingest.py in Python, each with its path.

### The indexer exited. Did it finish or just stop?

Read the terminal line and then the journal. Exit 0 with reason=completed means nothing was refused, and exit 3 with reason=completed-with-junk means the run finished and refused at least one file.

### Why does the index hold more documents than files?

XERJ writes one document per extracted row, and for source files it adds one document per declaration on top of the whole-file document. The document count therefore follows how many declarations your code holds, not a fixed multiple of the file count, so read it from your own run.

### Why does the plan hold fewer files than the folder?

Ignore rules pruned them. The captured dry run named target/, node_modules/ and a .gitignore entry, with an exact count of files inside each.

### Does _cat/indices report the index size in bytes?

No. Even with bytes=b it returned human-formatted values such as 58.5kb, so read GET /{index}/_stats and use _all.total.store.size_in_bytes.

## Related

- [How do I stop my agent from reading the whole repo into context?](/answers/index-monorepo-for-agent)
- [How should an agent figure out what's in a messy data folder before searching?](/answers/catalog-files-with-autoindex-map)
- [How do I read autoindex progress?](/answers/read-autoindex-progress)
