Evaluation
Performance on a held-out test set of 139 images that never appear in the reference index. Scoring parameters were fit by leave-one-out on the reference split only, so nothing here was tuned on the test set. This page is the working draft of the Task 5 Internal Testing and Validation Report.
Read these numbers as a baseline, not a field claim
The test images come from the same museum and photographic sources as the reference set: clean studio photographs of catalogued objects. Field photographs of seized or in-transit objects will score lower. The Task 6 simulated field test with CPEOC experts is where concordance on realistic inputs gets measured. With only 12 control images, the false-alarm rate has a wide interval.
Benchmark generated 2026-09-25 · corpus f2de7bac7cf69146 · model sha256 583fd1110a514667… · build-time latency p50 168 ms / p95 273 ms (Node v24.11.0, win32, CPU (onnxruntime)). Reproduce with npm run data:index.