Skip to content

Snapshot tests

We use snapshot tests to ensure that our HTML-to-text conversion is working as expected.

There is one split per source format: web, wiki, ar5iv, stackexchange, and dclm_hq. Each split has an inputs/ directory of source documents and an expected/ directory of markdown. The web split nests its markdown one level deeper, under expected/resiliparse/, because the extractor is part of the snapshot's identity. tests/test_snapshot.py asserts on the first four; dclm_hq has stored snapshots that generate_expected.py regenerates but no test compares.

Running the tests

To run the snapshot tests, run uv run pytest tests/test_snapshot.py.

Adding a test case

To add a test case, do the following:

Pro-tip: rather than writing the expected file by hand, run uv run python tests/snapshots/generate_expected.py to regenerate the expected outputs for every split, then review and edit the result.

If it's reasonable, try to add a unit test as well. This will help ensure that the conversion is correct. If you've made a change that you think is correct, you can update the snapshots by copying tests/snapshots/web/outputs/resiliparse/ over tests/snapshots/web/expected/resiliparse/. This will overwrite the expected output with the new output. You should review these changes before committing them.