Live public benchmark

Prerecorded deepfake detection, measured in the open.

We publish every headline metric with the corpus, methodology version, and per-family breakdown behind it. No cherry-picked slices, no vendor-only test sets. Human-in-the-loop review resolves inconclusive verdicts and is reported separately so the primary numbers stay honest.

Recall
98.26%
True positives caught
Precision
98.59%
Positives that were real
F1
98.42%
Balanced accuracy
False-positive rate
1.41%
On authentic control set
Samples
4,820
Methodology
v1.4.0
Corpus
VerifAI Prerecorded Deepfake Corpus (2025-Q4)
Published
7/25/2026

Per-family results

Blind evaluation across four generator families plus an authentic control set.

Attack familySamplesRecallPrecision / Specificity
lip sync82097.80%98.20%
face swap1,40099.10%98.90%
audio clone50098.20%98.60%
full synthesis1,10098.60%98.40%
authentic control1,000-98.59%

Blind evaluation across four generator families (FaceSwap-class, diffusion full-body synthesis, lip-sync retargeting, and voice-clone dubbing) plus an authentic control set. Adversarial post-processing (recompression, cropping, colour shift, frame drop) applied to 40% of positives. Human review resolves inconclusive verdicts (excluded from primary metrics but reported separately at 3.1%).

Test methodology

Corpus construction

Balanced mix of authentic footage (in-branch, video-KYC, wire authorisation calls, dispute evidence) and synthetic media generated across four families. 40% of positives are post-processed with recompression, cropping, colour shift and frame drop to mimic real-world capture chains.

Blind evaluation

Corpus IDs are hashed and split before scoring. The detection engine sees files only, never labels or generator hints. Ground truth is unsealed only after all verdicts are recorded and hashed into the audit chain.

Signal ensemble

Six signals are combined (frame-level classifier, temporal consistency, face/lip-sync coherence, audio authenticity, C2PA provenance, compression forensics) with confidence-weighted fusion. Frontier signals (rPPG cardiac liveness, physics/optics forensics) contribute when facial ROI is available.

Cryptographic reproducibility

Each run emits a results hash. The hash is anchored into the tamper-evident audit chain and, on a nightly cadence, into OpenTimestamps. Every published metric on this page can be traced back to its run hash.

Want to run this against your own corpus?

Regulated pilots include a sealed evaluation harness so your team can reproduce these numbers on your own footage under NDA. Signed run receipts are returned for every scored sample.