Benchmark v3 / July 2026 · Latest benchmark

Audio deepfake benchmark: 95.4% accuracy.

Evaluation across 2,000 samples and 12 commercial TTS providers, on the MLADD-v3 English holdout, with four reproduced baselines.

Dataset
MLADD-v3
Language / split
English / Holdout
Samples
2,000 · 1,000 genuine / 1,000 synthetic
Generators
12 commercial TTS providers
Models
DetectifAI v3 (measured) + 4 reproduced baselines
Status
Archived result · July 2026
Accuracy
95.4%
Higher is better
ROC-AUC
0.99
Higher is better
EER
4.6%
Lower is better
FPR · derived
3.2%
32 of 1,000 genuine
FNR · derived
6.0%
60 of 1,000 synthetic

Results explorer.

Switch metric, compare DetectifAI v3 against any baseline, and open a model to see its confusion matrix.

Benchmark explorer

Dataset
MLADD-v3
Split
English Holdout
Samples
2,000
Coverage
12 commercial TTS
Models evaluated
5
PredictedRealFakeActualRealFake96848.4%321.6%603.0%94047.0%92 errors across 2,000 samples · 32 FP · 60 FN
Accuracy
95.4%
AUC
0.99
Dataset
MLADD-v3
Split
English Holdout
Samples
2,000
Provenance
MEASURED
PredictedRealFakeActualRealFake85642.8%1447.2%21610.8%78439.2%360 errors across 2,000 samples · 144 FP · 216 FN
Accuracy
82.0%
AUC
0.87
Dataset
MLADD-v3
Split
English Holdout
Samples
2,000
Provenance
REPRODUCED
PredictedRealFakeActualRealFake61430.7%38619.3%47423.7%52626.3%860 errors across 2,000 samples · 386 FP · 474 FN
Accuracy
57.0%
AUC
0.55
Dataset
MLADD-v3
Split
English Holdout
Samples
2,000
Provenance
REPRODUCED
PredictedRealFakeActualRealFake57828.9%42221.1%48824.4%51225.6%910 errors across 2,000 samples · 422 FP · 488 FN
Accuracy
54.5%
AUC
0.53
Dataset
MLADD-v3
Split
English Holdout
Samples
2,000
Provenance
REPRODUCED
PredictedRealFakeActualRealFake53626.8%46423.2%50625.3%49424.7%970 errors across 2,000 samples · 464 FP · 506 FN
Accuracy
51.5%
AUC
0.51
Dataset
MLADD-v3
Split
English Holdout
Samples
2,000
Provenance
REPRODUCED

12 commercial TTS describes the generators sampled in this split.

Sample by sample.

Every cell is one of the 2,000 evaluation clips. Select a model to sort them into its outcomes.

DetectifAI v3 · Accuracy

95.4%

↑ Higher is betterThe share of evaluated samples classified correctly at a stated threshold.

Every cell is one evaluation clip

DetectifAI on 2,000 clips

DetectifAI confusion matrix, 2000 clips
Predicted genuinePredicted synthetic
Actual genuine968True negative32False positive
Actual synthetic60False negative940True positive
False-positive rate
3.2%
Genuine clips flagged synthetic · 32 / 1,000
False-negative rate
6.0%
Synthetic clips missed · 60 / 1,000

Derived from the published confusion matrix. Error cells are placed by count; the archive does not record which clip or generator each error came from.

MLADD-v3 English holdout: results for all models (Benchmark v3 / July 2026)
ModelAccuracyROC-AUCEERTrue negativeFalse positiveFalse negativeTrue positiveProvenance
DetectifAI v395.4%0.994.6%9683260940MEASURED
RawNet2 archived reproduction82.0%0.8718.0%856144216784REPRODUCED
Siamese-Net archived reproduction57.0%0.5543.0%614386474526REPRODUCED
SSL-W2V2 archived reproduction54.5%0.5345.5%578422488512REPRODUCED
SSL-WavLM archived reproduction51.5%0.5148.5%536464506494REPRODUCED

Where the errors fall.

Accuracy hides which way a model is wrong. DetectifAI v3 flags 3.2% of genuine clips and misses 6.0% of synthetic ones. The strongest baseline, RawNet2, flags 14.4% and misses 21.6%; the other three sit near chance (ROC-AUC 0.51–0.55).

False-positive rate · genuine flagged synthetic False-negative rate · synthetic missedDerived: FPR = FP ÷ 1,000 genuine, FNR = FN ÷ 1,000 synthetic
False-positive and false-negative rate by model, derived from the confusion matrices
ModelFalse-positive rateFalse-negative rate
DetectifAI3.2%6%
RawNet214.4%21.6%
Siamese-Net38.6%47.4%
SSL-W2V242.2%48.8%
SSL-WavLM46.4%50.6%

Confusion matrices.

Rows are the actual class, columns the prediction, on the same 2,000 clips for every model.

DetectifAI

v3 · MEASURED
DetectifAI confusion matrix, 2000 clips
Predicted genuinePredicted synthetic
Actual genuine968True negative32False positive
Actual synthetic60False negative940True positive

Acc 95.4% · AUC 0.99 · EER 4.6%

RawNet2

REPRODUCED
RawNet2 confusion matrix, 2000 clips
Predicted genuinePredicted synthetic
Actual genuine856True negative144False positive
Actual synthetic216False negative784True positive

Acc 82.0% · AUC 0.87 · EER 18.0%

Siamese-Net

REPRODUCED
Siamese-Net confusion matrix, 2000 clips
Predicted genuinePredicted synthetic
Actual genuine614True negative386False positive
Actual synthetic474False negative526True positive

Acc 57.0% · AUC 0.55 · EER 43.0%

SSL-W2V2

REPRODUCED
SSL-W2V2 confusion matrix, 2000 clips
Predicted genuinePredicted synthetic
Actual genuine578True negative422False positive
Actual synthetic488False negative512True positive

Acc 54.5% · AUC 0.53 · EER 45.5%

SSL-WavLM

REPRODUCED
SSL-WavLM confusion matrix, 2000 clips
Predicted genuinePredicted synthetic
Actual genuine536True negative464False positive
Actual synthetic506False negative494True positive

Acc 51.5% · AUC 0.51 · EER 48.5%

Twelve commercial voice generators.

The 1,000 synthetic clips are split almost evenly across providers. Bars show coverage, not per-provider scores: per-provider estimates are excluded from this report.

  1. ElevenLabs84
  2. Resemble AI84
  3. Speechify84
  4. Narakeet84
  5. Voicemaker83
  6. PlayHT83
  7. Murf AI83
  8. Deepgram Aura83
  9. Amazon Polly83
  10. Google Cloud TTS83
  11. Azure AI Speech83
  12. LOVO AI83

Complete result table.

MLADD-v3 English holdout results (Benchmark v3 / July 2026)
ModelAccuracy ROC-AUC EER TNFPFNTPProvenance
DetectifAIv395.4%0.994.6%9683260940MEASURED
RawNet2archived reproduction82.0%0.8718.0%856144216784REPRODUCED
Siamese-Netarchived reproduction57.0%0.5543.0%614386474526REPRODUCED
SSL-W2V2archived reproduction54.5%0.5345.5%578422488512REPRODUCED
SSL-WavLMarchived reproduction51.5%0.5148.5%536464506494REPRODUCED

MLADD-v3 · English · Holdout · N = 2,000 · Benchmark v3 / July 2026

Every metric has a direction.

EER
Lower is better
The operating point where false acceptance and false rejection rates are equal.
AUC
Higher is better
Area under the receiver operating characteristic curve across thresholds.
Accuracy
Higher is better
The share of evaluated samples classified correctly at a stated threshold.
FPR / FNR
Lower is better
False-positive and false-negative rates at the documented operating point.
Latency
Lower is better
Inference time under disclosed hardware, runtime, precision, and batch conditions.
Model size
Lower is smaller
The stored model artifact size for a disclosed version and numerical precision.

The method is part of the result.

What a DetectifAI benchmark documents, from the dataset manifest to the reproducibility artefacts.

01DatasetIdentify the corpus, version, access status, sources, languages, generators, and duration distribution.
  • Dataset manifest
  • Version and licence
  • Collection notes
  • Class composition
02PreprocessingDocument sample rate, channel handling, segmentation, normalisation, and every transformation before inference.
  • Audio conversion
  • Clip policy
  • Normalisation
  • Failure handling
03SplitsSeparate training, validation, and held-out evaluation while exposing speaker or generator overlap.
  • Split definition
  • Leakage checks
  • Threshold split
  • Held-out boundary
04MetricsDefine direction, threshold selection, aggregation, and operational interpretation for every reported metric.
  • Metric formula
  • Direction
  • Threshold method
  • Confidence treatment
05InferencePin the model, checkpoint, runtime, clip aggregation, and warm or cold execution path.
  • Model version
  • Checkpoint hash
  • Runtime
  • Aggregation policy
06HardwareDisclose hardware, numerical precision, batch size, audio duration, and measurement protocol.
  • Device
  • Precision
  • Batch size
  • Timing protocol
07ReproducibilityConnect results to code, configuration, source records, evaluation date, owner, and reviewer.
  • Pinned commit
  • Configuration
  • Evidence owner
  • Review record

Dataset manifest fields

  • Dataset name and version
  • Real and synthetic sample counts
  • Languages and speech sources
  • Generators or synthesis sources
  • Duration distribution
  • Training, validation, and held-out splits
  • Licence or access status
  • Collection and preprocessing notes

Reproducibility artefacts

  • Evaluation configuration
  • Pinned code version
  • Model and checkpoint references
  • Dataset manifest
  • Raw prediction record
  • Metric calculation output
  • Hardware and runtime record
  • Approval and review record

Every value carries its source.

Every row in this report is tagged with how its number was obtained.

OFFICIAL
Reported by the original author or official repository, with a direct source.
REPRODUCED
Executed by DetectifAI in a documented environment with a pinned implementation.
INTERNAL
A DetectifAI model or internal evaluation, clearly labelled and scoped.

Test it against the conditions your calls actually have.

Channels, codecs, languages and generators change detection performance. Evaluate DetectifAI on your own flow.