Prototype for simulated field testing. Results are investigative leads for a person to review. They are never legal determinations.

CPIT

Evaluation

Performance on a held-out test set of 139 images that never appear in the reference index. Scoring parameters were fit by leave-one-out on the reference split only, so nothing here was tuned on the test set. This page is the working draft of the Task 5 Internal Testing and Validation Report.

Top-1 accuracy

85.6%

95% CI 79%–90%, n=139

Top-3 accuracy

95.0%

True category among the first three

Source country correct

90.6%

127 designated-list images

Missed detections

0.0%

Designated objects ranked as control

False alarms

16.7%

Control objects flagged, n=12

Calibration error

0.035

Expected calibration error (lower is better)

Per-category results
Precision: when the tool says this category, how often it is right. Recall: how often it finds this category.
CategorynPrecisionRecall
Scythian gold and bronze ornamentUkraine8
89%
100%
Cucuteni–Trypillia painted ceramicsUkraine13
100%
92%
Eastern-rite icons and ecclesiastical objectsUkraine6
100%
100%
Faience shabtis and amuletsEgypt11
73%
100%
Coffins, cartonnage and funerary masksEgypt12
80%
67%
Stone stelae and relief fragmentsEgypt11
100%
82%
Moche modelled and stirrup-spout ceramicsPeru9
58%
78%
Nasca polychrome ceramicsPeru10
88%
70%
Andean textilesPeru10
100%
90%
Black- and red-figure painted potteryGreece9
80%
89%
Shang and Zhou bronzes (vessels and weapons)China13
93%
100%
Han–Tang tomb ceramics and figuresChina5
57%
80%
Khmer stone and bronze sculptureCambodia10
88%
70%
Out-of-scope controlControl12
100%
83%
Confusion matrix
Rows: true category. Columns: top prediction. Hover a cell for its count.
UA-SCYUA-TRYUA-ICOEG-SHAEG-FUNEG-RELPE-MOCPE-NASPE-TEXGR-PAICN-RITCN-TANKH-KHMCTL
UA-SCY
8
UA-TRY
12
1
UA-ICO
6
EG-SHA
11
EG-FUN
3
8
1
EG-REL
1
1
9
PE-MOC
1
7
1
PE-NAS
1
7
1
1
PE-TEX
9
1
GR-PAI
1
8
CN-RIT
13
CN-TAN
1
4
KH-KHM
1
1
1
7
CTL
1
1
10
UA-SCY
Scythian gold and bronze ornament
UA-TRY
Cucuteni–Trypillia painted ceramics
UA-ICO
Eastern-rite icons and ecclesiastical objects
EG-SHA
Faience shabtis and amulets
EG-FUN
Coffins, cartonnage and funerary masks
EG-REL
Stone stelae and relief fragments
PE-MOC
Moche modelled and stirrup-spout ceramics
PE-NAS
Nasca polychrome ceramics
PE-TEX
Andean textiles
GR-PAI
Black- and red-figure painted pottery
CN-RIT
Shang and Zhou bronzes (vessels and weapons)
CN-TAN
Han–Tang tomb ceramics and figures
KH-KHM
Khmer stone and bronze sculpture
CTL
Out-of-scope control
Calibration
Does a stated confidence of 70% mean the tool is right about 70% of the time? These are grouped held-out predictions.
Stated confidencePredictionsMean confidenceObserved accuracy
0%–20%0——
20%–40%836%25%
40%–60%1451%57%
60%–80%2170%71%
80%–100%9695%98%
Ablation: what each signal contributes
Fused score = 0.6 × exemplar similarity + 0.4 × description similarity, softmax temperature 0.01, k = 1.
ConfigurationTop-1Top-3False alarmsECE
Fused (shipped)85.6%95.0%16.7%0.035
Exemplars only86.3%94.2%16.7%0.043
Descriptions only (zero-shot)53.2%80.6%8.3%0.083

Descriptions alone are weak. Their value is in explaining a result in words and in lowering calibration error, not in raw accuracy.

Model selection
Candidate on-device encoders, evaluated on the same corpus, split and procedure. Only permissively licensed weights were eligible.
ModelLicenceSizeTop-1Top-3False alarmsECE
DINOv2 ViT-S/14 (Meta), 8-bit quantised Apache-2.024.5 MB89.3%95.7%25.0%0.114
CLIP ViT-B/32 (OpenAI), 8-bit quantised ShippedMIT89.1 MB85.6%95.0%16.7%0.035

DINOv2 is smaller and slightly more accurate on top-1. CLIP was chosen for three reasons: its calibration error is about three times lower, it raised fewer false alarms, and it scores images against text descriptions, which supplies the plain-language rationale for each result. The top-1 gap (about five test images) is within the confidence interval. Apple MobileCLIP was excluded before testing because of its research-only licence.

Benchmark generated 2026-09-25 · corpus f2de7bac7cf69146 · model sha256 583fd1110a514667… · build-time latency p50 168 ms / p95 273 ms (Node v24.11.0, win32, CPU (onnxruntime)). Reproduce with npm run data:index.