The first multilingual commercial-TTS benchmark.
The original archived comparison of DetectifAI models and four open-source baselines across English and Hindi audio.
- Dataset
- MLADD-v2 early evaluation
- Reported scale
- 400 audio files
- Languages
- English and Hindi
- Synthesis sources
- 4 commercial TTS providers
- Status
- Archived · superseded by v3
Reported performance across both languages.
These figures are restored from the original preorder-site report. DetectifAI rows are highlighted; baseline rows remain neutral.
Hindi
EER · lower is betterEnglish holdout
200 samples · 100 real / 100 synthetic
| Model | Type | AUC Higher is better | EER Lower is better | Accuracy Higher is better |
|---|---|---|---|---|
| DetectifAI v2 | DetectifAI | 0.8664 | 22.00% | 78.00% |
| RawNet2 | Baseline | 0.8715 | 18.00% | 82.00% |
| DetectifAI v1 | DetectifAI | 0.8687 | 26.00% | 74.00% |
| Siamese-Net | Baseline | 0.5477 | 43.00% | 57.00% |
| SSL-W2V2 | Baseline | 0.5305 | 45.50% | 54.50% |
| SSL-WavLM | Baseline | 0.5084 | 48.50% | 51.50% |
Hindi holdout
202 samples · 100 real / 102 synthetic
| Model | Type | AUC Higher is better | EER Lower is better | Accuracy Higher is better |
|---|---|---|---|---|
| DetectifAI v2 | DetectifAI | 0.9943 | 2.65% | 97.35% |
| SSL-WavLM | Baseline | 0.9998 | 0.49% | 99.50% |
| SSL-W2V2 | Baseline | 0.9975 | 1.98% | 98.02% |
| DetectifAI v1 | DetectifAI | 0.9838 | 6.93% | 93.07% |
| RawNet2 | Baseline | 0.9465 | 12.87% | 87.13% |
| Siamese-Net | Baseline | 0.9285 | 15.35% | 84.65% |
Performance by commercial synthesis source.
Equal error rate is shown for the four providers included in the original reports.
English per-provider EER
| Model | Narakeet | Resemble AI | Speechify | Voicemaker | Overall |
|---|---|---|---|---|---|
| DetectifAI v2DetectifAI | 12.00% | 32.00% | 22.00% | 12.00% | 22.00% |
| RawNet2Baseline | 10.50% | 21.50% | 28.00% | 7.50% | 18.00% |
| DetectifAI v1DetectifAI | 23.50% | 30.00% | 32.00% | 21.00% | 26.00% |
| Siamese-NetBaseline | 43.00% | 40.50% | 40.50% | 43.50% | 43.00% |
| SSL-W2V2Baseline | 39.50% | 40.50% | 39.00% | 47.50% | 45.50% |
| SSL-WavLMBaseline | 46.50% | 45.00% | 48.00% | 47.00% | 48.50% |
Hindi per-provider EER
| Model | Narakeet | Resemble AI | Speechify | Voicemaker | Overall |
|---|---|---|---|---|---|
| DetectifAI v2DetectifAI | 0.10% | 0.20% | 0.00% | 0.00% | 2.65% |
| SSL-WavLMBaseline | 0.00% | 0.00% | 1.00% | 0.00% | 0.49% |
| SSL-W2V2Baseline | 0.50% | 4.00% | 0.00% | 0.00% | 1.98% |
| DetectifAI v1DetectifAI | 7.50% | 7.50% | 10.00% | 1.00% | 6.93% |
| RawNet2Baseline | 7.00% | 15.00% | 19.50% | 5.70% | 12.87% |
| Siamese-NetBaseline | 7.50% | 7.00% | 23.50% | 9.56% | 15.35% |
Cross-language comparison.
The language gap is the absolute difference between English and Hindi EER.
| Model | Type | English EER | Hindi EER | Average EER | Language gap |
|---|---|---|---|---|---|
| DetectifAI v2 | DetectifAI | 22.00% | 2.65% | 12.33% | 19.35 pp |
| RawNet2 | Baseline | 18.00% | 12.87% | 15.44% | 5.13 pp |
| DetectifAI v1 | DetectifAI | 26.00% | 6.93% | 16.47% | 19.07 pp |
| SSL-W2V2 | Baseline | 45.50% | 1.98% | 23.74% | 43.52 pp |
| SSL-WavLM | Baseline | 48.50% | 0.49% | 24.50% | 48.01 pp |
| Siamese-Net | Baseline | 43.00% | 15.35% | 29.18% | 27.65 pp |
Read the archive with its limits.
The recovered report describes 400 files, while its language tables total 402 samples.
A pinned code version, hardware record, and executable evaluation package were not included.
Threshold-selection details and confidence intervals were not supplied.
The results predate the current DetectifAI model and should not be read as current production performance.
Historical continuity only. This report is not a current deployment, certification, or regulatory claim.