Audio deepfake benchmark: 95.4% accuracy.

Evaluation across 2,000 samples and 12 commercial TTS providers.

Dataset
MLADD-v3
Language
English Holdout
Samples
2,000 total
Distribution
1,000 real / 1,000 fake
Generators
12 commercial TTS

Results at a glance.

Detection accuracy

Higher is better
0%50%100%

Area under curve (AUC)

Higher is better
0.000.501.00

Tested across twelve commercial voice generators.

95.4% overall accuracy across a balanced 2,000-sample holdout.

95.4%Overall accuracy

Bars show the synthetic sample coverage contributed by each generator—not invented per-provider scores.

Complete result table.

MLADD-v3 audio deepfake detection benchmark results
Model / versionDataset / splitAUC Higher is betterAccuracy Higher is betterTNFPFNTPProvenance
DetectifAIv3MLADD-v3English / Holdout0.9995.4%9544646954MEASUREDArchived
RawNet2archived reproductionMLADD-v3English / Holdout0.8782.0%820180180820REPRODUCEDArchived
Siamese-Netarchived reproductionMLADD-v3English / Holdout0.5557.0%570430430570REPRODUCEDArchived
SSL-W2V2archived reproductionMLADD-v3English / Holdout0.5354.5%545455455545REPRODUCEDArchived
SSL-WavLMarchived reproductionMLADD-v3English / Holdout0.5151.5%515485485515REPRODUCEDArchived

Confusion matrices.

Primary result

DetectifAI v3

N = 2,000
Predicted realPredicted fake
Actual realActual synthetic

Baseline comparison

RawNet2 archived reproduction

N = 2,000
Predicted realPredicted fake
Actual realActual synthetic

Provenance key

Every value carries its source.

OFFICIAL

Reported by the original author or official repository, with a direct source.

REPRODUCED

Executed by DetectifAI in a documented environment with a pinned implementation.

INTERNAL

A DetectifAI model or internal evaluation, clearly labelled and scoped.

A benchmark is more than one score.

Limitations

  • The recovered archive did not include a pinned code version or executable evaluation package.
  • Hardware, runtime, numerical precision, and timing conditions were not supplied.
  • A named evidence owner, reviewer, and current approval record were not supplied.
  • Per-provider estimates and carried-forward language rows are excluded.

The method is part of the result.

MetricDirectionDefinition

EER

Lower is better

The operating point where false acceptance and false rejection rates are equal.

AUC

Higher is better

Area under the receiver operating characteristic curve across thresholds.

Accuracy

Higher is better

The share of evaluated samples classified correctly at a stated threshold.

FPR / FNR

Lower is better

False-positive and false-negative rates at the documented operating point.

Dataset

Identify the corpus, version, access status, sources, languages, generators, and duration distribution.

Dataset manifestVersion and licenceCollection notesClass composition

Preprocessing

Document sample rate, channel handling, segmentation, normalisation, and every transformation before inference.

Audio conversionClip policyNormalisationFailure handling

Splits

Separate training, validation, and held-out evaluation while exposing speaker or generator overlap.

Split definitionLeakage checksThreshold splitHeld-out boundary

Metrics

Define direction, threshold selection, aggregation, and operational interpretation for every reported metric.

Metric formulaDirectionThreshold methodConfidence treatment

Inference

Pin the model, checkpoint, runtime, clip aggregation, and warm or cold execution path.

Model versionCheckpoint hashRuntimeAggregation policy

Hardware

Disclose hardware, numerical precision, batch size, audio duration, and measurement protocol.

DevicePrecisionBatch sizeTiming protocol

Reproducibility

Connect results to code, configuration, source records, evaluation date, owner, and reviewer.

Pinned commitConfigurationEvidence ownerReview record

Evidence before assertion.

Publication requires a complete, reviewed evidence package.

  • Evaluation configurationRequired
  • Pinned code versionRequired
  • Model and checkpoint referencesRequired
  • Dataset manifestRequired
  • Raw prediction recordRequired
  • Metric calculation outputRequired
  • Hardware and runtime recordRequired
  • Approval and review recordRequired