argand-site-registry/docs/BENCHMARKING.md
nicweyand 557ba7cd69
All checks were successful
Standalone registry checks / check (push) Successful in 5m58s
release: implement site registry v0.5
2026-09-13 14:42:39 -04:00

4.3 KiB

Benchmarking and recovery

scripts/benchmark.py runs one import, build, audit verification, or evaluation command without a shell and creates a new evidence directory. REPORT.json records exact source-manifest and source-object hashes, hardware, wall/CPU time, peak RSS, database/record/fact growth, hashes of stdout/stderr, optional measured artifacts, and a hash of the argument vector. Source bytes are never copied into the report. The command refuses to overwrite an existing directory.

Use three profiles consistently:

  • small: synthetic or source-shaped fixtures used for correctness and rapid regression checks;
  • medium: a documented provider subset large enough to exercise checkpoints, bounds and projection behavior; and
  • provider: one complete current official provider object in its native format.

Example import measurement:

python3 scripts/benchmark.py --profile provider \
  --output /data/benchmarks/ror-v2.12-import \
  --source-manifest /data/cache/ror/source.json \
  --database /data/ror-writer.sqlite --expanded-bytes 362619018 -- \
  argand-site-registry import --database /data/ror-writer.sqlite \
    --input /data/cache/ror/SHA256 --manifest /data/cache/ror/source.json

For build measurements, label exact regular files with repeated --measured-path NAME=PATH. Run evaluation with at least 1,000 representative cases in one process so authenticated startup/copy time remains separate from the report's native per-query p50/p95. Store reports outside Git because commands and snapshots can reveal private operational structure. A report establishes one machine and object only; synthetic numbers are never provider-capacity evidence.

Version 0.5 provider canary

The 2026-09-13 canary used official ROR release v2.12-2026-08-25, Zenodo record 22099990. The compressed object was 36,246,232 bytes, SHA-256 5779c7baf71771fd8ea829201e7bd4343a3c68ff36c595f480b3a00292f78931, and matched provider MD5 ce8807691455d4ada3216c31408e9e1a. Its JSON and CSV members expanded to 362,619,018 bytes.

The final streaming import produced 137,398 records and 914,439 facts in 62.30 seconds (9.47 user, 13.24 system), with 25,032 KiB peak RSS and about 14,678 facts/second. It grew the PSL-initialized writer by 828,682,240 bytes to 829,505,536 bytes. Reimport replayed the complete pinned source in 0.025 seconds without changing any record, fact, or database byte count.

A compact build combining that ROR source with the current PSL projected 137,398 entities, 131,591 properties and 133,398 edges. The pre-compaction database was 1,345,097,728 bytes; the runtime database was 824,705,024 bytes, a 38.688% reduction. The initial compact build took 58.36 seconds with 45,140 KiB peak RSS; a second build reused the same bundles and produced byte-identical database and receipt bytes. The two cold bundles totaled 686,468,048 bytes and held 137,399 records plus 914,440 facts. Warm-cache full streaming audit verification took 3.03 seconds with 24,076 KiB peak RSS.

A 1,000-case safe-abstention evaluation passed 1,000/1,000 with native reusable reader latency of 72 microseconds p50 and 96 microseconds p95. The warm-cache full process took 0.83 seconds, including authentication and the private copy of the 786.5 MiB SQLite file. Serving deployments must reuse the verified Registry and must measure cold startup on their own storage.

Hardware: Linux 6.18.51-1-lts x86_64, AMD Ryzen 9 7900X (12 cores, 24 logical CPUs), 61 GiB RAM, ext4. These numbers are a single local canary and not a public service capacity claim.

Recovery matrix

The acceptance suite sends real process termination after two distinct committed checkpoints and proves the same import completes exactly once on restart. The provider canary also killed the final binary at a 2,304-record checkpoint via a bounded-runner fault, recovered its hot journal, and reached the exact clean-import record/fact totals. Other tests cover handled parser failure rollback, SQLite growth-cap (SQLITE_FULL) behavior, truncated/compressed inputs, mutated cache objects, duplicate records and JSON keys, archive-member drift, missing/truncated/substituted cold bundles, existing publication targets, failed atomic writes, and unchanged previous generations after update failure. Provider objects and benchmark output remain outside the repository.