DetectifAI
v3 · MEASURED| Predicted genuine | Predicted synthetic | |
|---|---|---|
| Actual genuine | 968True negative | 32False positive |
| Actual synthetic | 60False negative | 940True positive |
Acc 95.4% · AUC 0.99 · EER 4.6%
Benchmark v3 / July 2026 · Latest benchmark
Evaluation across 2,000 samples and 12 commercial TTS providers, on the MLADD-v3 English holdout, with four reproduced baselines.
Switch metric, compare DetectifAI v3 against any baseline, and open a model to see its confusion matrix.
12 commercial TTS describes the generators sampled in this split.
Every cell is one of the 2,000 evaluation clips. Select a model to sort them into its outcomes.
MLADD-v3 english holdout set: 1,000 genuine and 1,000 synthetic clips from 12 commercial TTS providers. DetectifAI v3 is measured; the four baselines are reproduced by DetectifAI from archived implementations.
Values come from the recovered July 2026 archive, which did not include the code version, hardware or approval record the current publication gate requires. Read the limitations.
DetectifAI v3 · Accuracy
95.4%
↑ Higher is betterThe share of evaluated samples classified correctly at a stated threshold.
Every cell is one evaluation clip
| Predicted genuine | Predicted synthetic | |
|---|---|---|
| Actual genuine | 968True negative | 32False positive |
| Actual synthetic | 60False negative | 940True positive |
Derived from the published confusion matrix. Error cells are placed by count; the archive does not record which clip or generator each error came from.
| Model | Accuracy | ROC-AUC | EER | True negative | False positive | False negative | True positive | Provenance |
|---|---|---|---|---|---|---|---|---|
| DetectifAI v3 | 95.4% | 0.99 | 4.6% | 968 | 32 | 60 | 940 | MEASURED |
| RawNet2 archived reproduction | 82.0% | 0.87 | 18.0% | 856 | 144 | 216 | 784 | REPRODUCED |
| Siamese-Net archived reproduction | 57.0% | 0.55 | 43.0% | 614 | 386 | 474 | 526 | REPRODUCED |
| SSL-W2V2 archived reproduction | 54.5% | 0.53 | 45.5% | 578 | 422 | 488 | 512 | REPRODUCED |
| SSL-WavLM archived reproduction | 51.5% | 0.51 | 48.5% | 536 | 464 | 506 | 494 | REPRODUCED |
Accuracy hides which way a model is wrong. DetectifAI v3 flags 3.2% of genuine clips and misses 6.0% of synthetic ones. The strongest baseline, RawNet2, flags 14.4% and misses 21.6%; the other three sit near chance (ROC-AUC 0.51–0.55).
| Model | False-positive rate | False-negative rate |
|---|---|---|
| DetectifAI | 3.2% | 6% |
| RawNet2 | 14.4% | 21.6% |
| Siamese-Net | 38.6% | 47.4% |
| SSL-W2V2 | 42.2% | 48.8% |
| SSL-WavLM | 46.4% | 50.6% |
Rows are the actual class, columns the prediction, on the same 2,000 clips for every model.
| Predicted genuine | Predicted synthetic | |
|---|---|---|
| Actual genuine | 968True negative | 32False positive |
| Actual synthetic | 60False negative | 940True positive |
Acc 95.4% · AUC 0.99 · EER 4.6%
| Predicted genuine | Predicted synthetic | |
|---|---|---|
| Actual genuine | 856True negative | 144False positive |
| Actual synthetic | 216False negative | 784True positive |
Acc 82.0% · AUC 0.87 · EER 18.0%
| Predicted genuine | Predicted synthetic | |
|---|---|---|
| Actual genuine | 614True negative | 386False positive |
| Actual synthetic | 474False negative | 526True positive |
Acc 57.0% · AUC 0.55 · EER 43.0%
| Predicted genuine | Predicted synthetic | |
|---|---|---|
| Actual genuine | 578True negative | 422False positive |
| Actual synthetic | 488False negative | 512True positive |
Acc 54.5% · AUC 0.53 · EER 45.5%
| Predicted genuine | Predicted synthetic | |
|---|---|---|
| Actual genuine | 536True negative | 464False positive |
| Actual synthetic | 506False negative | 494True positive |
Acc 51.5% · AUC 0.51 · EER 48.5%
The 1,000 synthetic clips are split almost evenly across providers. Bars show coverage, not per-provider scores: per-provider estimates are excluded from this report.
| Model | Accuracy ↑ | ROC-AUC ↑ | EER ↓ | TN | FP | FN | TP | Provenance |
|---|---|---|---|---|---|---|---|---|
| DetectifAIv3 | 95.4% | 0.99 | 4.6% | 968 | 32 | 60 | 940 | MEASURED |
| RawNet2archived reproduction | 82.0% | 0.87 | 18.0% | 856 | 144 | 216 | 784 | REPRODUCED |
| Siamese-Netarchived reproduction | 57.0% | 0.55 | 43.0% | 614 | 386 | 474 | 526 | REPRODUCED |
| SSL-W2V2archived reproduction | 54.5% | 0.53 | 45.5% | 578 | 422 | 488 | 512 | REPRODUCED |
| SSL-WavLMarchived reproduction | 51.5% | 0.51 | 48.5% | 536 | 464 | 506 | 494 | REPRODUCED |
MLADD-v3 · English · Holdout · N = 2,000 · Benchmark v3 / July 2026
What a DetectifAI benchmark documents, from the dataset manifest to the reproducibility artefacts.
Every row in this report is tagged with how its number was obtained.
Channels, codecs, languages and generators change detection performance. Evaluate DetectifAI on your own flow.