Audio deepfake benchmark: 95.4% accuracy.
Evaluation across 2,000 samples and 12 commercial TTS providers.
- Dataset
- MLADD-v3
- Language
- English Holdout
- Samples
- 2,000 total
- Distribution
- 1,000 real / 1,000 fake
- Generators
- 12 commercial TTS
Results at a glance.
Area under curve (AUC)
Higher is betterTested across twelve commercial voice generators.
95.4% overall accuracy across a balanced 2,000-sample holdout.
Bars show the synthetic sample coverage contributed by each generator—not invented per-provider scores.
Complete result table.
| Model / version | Dataset / split | AUC Higher is better | Accuracy Higher is better | TN | FP | FN | TP | Provenance |
|---|---|---|---|---|---|---|---|---|
| DetectifAIv3 | MLADD-v3English / Holdout | 0.99 | 95.4% | 954 | 46 | 46 | 954 | MEASUREDArchived |
| RawNet2archived reproduction | MLADD-v3English / Holdout | 0.87 | 82.0% | 820 | 180 | 180 | 820 | REPRODUCEDArchived |
| Siamese-Netarchived reproduction | MLADD-v3English / Holdout | 0.55 | 57.0% | 570 | 430 | 430 | 570 | REPRODUCEDArchived |
| SSL-W2V2archived reproduction | MLADD-v3English / Holdout | 0.53 | 54.5% | 545 | 455 | 455 | 545 | REPRODUCEDArchived |
| SSL-WavLMarchived reproduction | MLADD-v3English / Holdout | 0.51 | 51.5% | 515 | 485 | 485 | 515 | REPRODUCEDArchived |
Confusion matrices.
Primary result
DetectifAI v3
Baseline comparison
RawNet2 archived reproduction
Provenance key
Every value carries its source.
Reported by the original author or official repository, with a direct source.
Executed by DetectifAI in a documented environment with a pinned implementation.
A DetectifAI model or internal evaluation, clearly labelled and scoped.
A benchmark is more than one score.
Limitations
- The recovered archive did not include a pinned code version or executable evaluation package.
- Hardware, runtime, numerical precision, and timing conditions were not supplied.
- A named evidence owner, reviewer, and current approval record were not supplied.
- Per-provider estimates and carried-forward language rows are excluded.
The method is part of the result.
EER
Lower is betterThe operating point where false acceptance and false rejection rates are equal.
AUC
Higher is betterArea under the receiver operating characteristic curve across thresholds.
Accuracy
Higher is betterThe share of evaluated samples classified correctly at a stated threshold.
FPR / FNR
Lower is betterFalse-positive and false-negative rates at the documented operating point.
Dataset
Identify the corpus, version, access status, sources, languages, generators, and duration distribution.
+
Preprocessing
Document sample rate, channel handling, segmentation, normalisation, and every transformation before inference.
+
Splits
Separate training, validation, and held-out evaluation while exposing speaker or generator overlap.
+
Metrics
Define direction, threshold selection, aggregation, and operational interpretation for every reported metric.
+
Inference
Pin the model, checkpoint, runtime, clip aggregation, and warm or cold execution path.
+
Hardware
Disclose hardware, numerical precision, batch size, audio duration, and measurement protocol.
+
Reproducibility
Connect results to code, configuration, source records, evaluation date, owner, and reviewer.
+
Evidence before assertion.
Publication requires a complete, reviewed evidence package.
- Evaluation configurationRequired
- Pinned code versionRequired
- Model and checkpoint referencesRequired
- Dataset manifestRequired
- Raw prediction recordRequired
- Metric calculation outputRequired
- Hardware and runtime recordRequired
- Approval and review recordRequired