<!-- generated by build_articles.py — edit content/answers/search-html-export.md instead -->
---
title: "Search a help-center you saved as HTML"
canonical: "https://xerj.org/answers/search-html-export"
updated: "2026-08-21"
source: "content/answers/search-html-export.md"
---

# I saved a help-center as HTML. How do I search it like the real help-center search?

**TL;DR** — XERJ `autoindex` indexes a static HTML export straight from local disk. In our 3-page capture XERJ extracted a `title` and a `headings` array from each page, and one `match` query on `body` returned 3 hits. A network watch observed 0 non-loopback peers during the run.

## Agent prompt

```text
Act as a coding agent. Read https://xerj.org/llms.txt, start a XERJ node with --insecure, run xerj autoindex ./html-export --prefix web on a folder of saved HTML pages, then POST a match query on body to /web-*/_search and report ax_path, title and the headings array for every hit without fetching any URL.
```

## Commands

### Command 1

Note: Index a folder of saved HTML pages from local disk.

```sh
xerj autoindex ./html-export --url http://127.0.0.1:9410 --prefix web --state-dir ./state-web --progress plain --disable-feedback
```

### Command 2

Note: Search the page text and return the title and headings of each hit.

```sh
curl -s -XPOST 'http://127.0.0.1:9410/web-*/_search' -H 'content-type: application/json' -d '{"query":{"match":{"body":"checkpoint journal"}},"size":10,"_source":["ax_path","title","headings","section"],"track_total_hits":true}'
```

## Index the export from disk

`xerj autoindex` indexes the HTML files that are already on disk and fetches no URL. XERJ has no fetcher and follows no `href`, so a saved site export is the input, not a live site.

The command below indexes a folder of 3 static pages.

```sh
xerj autoindex ./html-export --url http://127.0.0.1:9410 --prefix web --state-dir ./state-web --progress plain --disable-feedback
```

## Fields the HTML family produced

The HTML family produced `title`, `headings` and `body`, on top of the 7 `ax_*` provenance fields. The `headings` field is an array of the heading text in document order, so `recovery.html` carried `["Recovery procedure", "Overview"]`.

That field set distinguishes HTML from plain text in XERJ. A Markdown file lands in the `txt-prose` family and gets no `headings` array, as [the Markdown answer](/answers/index-markdown-into-elasticsearch-api) shows.

## The query and its 3 hits

One `match` query on `body` for `checkpoint journal` returned 3 hits, one per page, ranked by BM25. The captured `_source` carried `ax_path`, `title` and `headings` for every hit.

```sh
curl -s -XPOST 'http://127.0.0.1:9410/web-*/_search' -H 'content-type: application/json' -d '{"query":{"match":{"body":"checkpoint journal"}},"size":10,"_source":["ax_path","title","headings","section"],"track_total_hits":true}'
```

| `ax_path` | `title` | `headings` | `_score` |
| --- | --- | --- | --- |
| `recovery.html` | Recovery procedure | `["Recovery procedure", "Overview"]` | `0.61400104` |
| `index.html` | Export index | `["Runbook export", "Overview"]` | `0.57417387` |
| `glossary.html` | Glossary | `["Glossary", "Overview"]` | `0.14597225` |

## Proof that no URL was fetched

A network watch over the XERJ node and the whole harness process tree observed 0 non-loopback peers for the whole run. The watch polled `/proc/net/tcp` and `/proc/net/tcp6` and cross-referenced the socket inodes of up to 4 process ids in the watched tree.

The method has one honest limit, and the capture states it. The watch is a sampler, not a packet capture, and it took 4 samples at a 0.05 second interval across 0.32 seconds. A connection that opens and closes inside one gap can escape the sampler.

## Limits worth knowing before you start

XERJ indexes files from a filesystem walk on 1 single-node process, so the export must exist on disk first. A hidden dotfile and a dotted directory are always skipped, and `--no-ignore` does not turn that off.

XERJ elects `body` as a `semantic_text` field on document datasets. The default embedder is lexical feature hashing, and the neural embedder is opt-in through `--embed-mode neural`. A `match` query on a `semantic_text` field runs BM25, not kNN.

## FAQ

### I saved a bunch of docs pages as HTML. How do I search that?

Index the folder from disk and send a `match` query on `body`. The HTML family produces `title`, `headings` and `body`, so a hit names the page and its section headings.

### How do I search my Confluence HTML export?

The same way. XERJ reads the saved files from disk and fetches no URL, and the Confluence export article covers that space layout in detail.

### How do I search a wget of vendor docs?

Point `autoindex` at the download directory. Every page becomes a document with `title`, `headings` and `body`, and one `match` query on `body` covers the whole mirror.

### Does site navigation drown the real hits?

It can. Every saved page repeats its nav and footer text in `body`, so a common word matches many pages. Query a distinctive phrase, or read `title` and `headings` to keep the answer inside the article text.

### Does XERJ download the pages itself?

No. XERJ indexes HTML files that are already on disk and fetches no URL. Our network watch observed 0 non-loopback peers during the run.

### Can I see which headings a page has?

Yes. Ask for `headings` in `_source`. In our capture `recovery.html` returned `["Recovery procedure", "Overview"]` in document order.

### Is the zero-network result a packet capture?

No. The capture is a sampler at a 0.05 second interval, so a connection opening and closing inside one gap would be missed. The capture states that limit.

## Related

- [I have a folder of PDFs, Word docs, and markdown. How do I search all of them at once?](/answers/search-file-contents-in-a-folder)
- [How do I search a Confluence export?](/answers/search-confluence-html-export)
- [Can ChatGPT search a folder on my laptop, or do I need something else?](/answers/give-chatgpt-claude-local-file-access)
- [Pagefind vs indexing a local HTML dump for an agent?](/compare/xerj-vs-pagefind)
