Every accuracy figure VerifAI publishes is measured on a blind, held-out corpus of 4,820 clips and reported with a two-sided Wilson 95% confidence interval. If the lower bound ever falls under 98%, this page - and the site's headline number - auto-downgrades. No exceptions.
The Wilson 95% CI lower bound (98.0%) sits above the 98% floor. That is the statistical guarantee behind the "98% accuracy" line in the introduction video: if we ran the same corpus 20 more times, at least 19 of them would land above 98.0%.
A single blended number can hide a weak family. Here is the same corpus split by generator class - the failure surface any bank actually cares about.
| Family | Samples | Recall | Precision | Specificity |
|---|---|---|---|---|
| lip sync | 820 | 97.80% | 98.20% | - |
| face swap | 1,400 | 99.10% | 98.90% | - |
| audio clone | 500 | 98.20% | 98.60% | - |
| full synthesis | 1,100 | 98.60% | 98.40% | - |
| authentic control | 1,000 | - | - | 98.59% |
Every confirmed analyst decision (review queue, verdict override, benchmark-lab clip) updates per-signal true/false positive-rate EMAs and re-derives a bounded weight multiplier (clamped to 0.5x-1.5x). Every change is written here as append-only evidence.
Anyone claiming 100% detection accuracy on synthetic video is either lying or has not been evaluated adversarially. Generator classes evolve monthly. We publish the headline at 98% to leave calibrated headroom for a novel family appearing between corpus refreshes, and we hard-cap per-scan calibration_confidence at 0.99. The 26-layer SHIELD watches for accuracy regression: if a nightly re-run falls below the CI lower bound by more than the CI width, a security_events row is written and this page renders the new, lower number on next request.