A larger multilingual robustness evaluation.
The second archived report expanded the DetectifAI model family while retaining the English and Hindi commercial-TTS stress tests.
- Dataset
- MLADD-v2 expanded evaluation
- Reported scale
- 1,600 audio files
- Languages
- English and Hindi
- Synthesis sources
- 4 commercial TTS providers
- Status
- Archived · superseded by v3
Reported performance across both languages.
These figures are restored from the original preorder-site report. DetectifAI rows are highlighted; baseline rows remain neutral.
Hindi
EER · lower is betterEnglish holdout
800 samples · 400 real / 400 synthetic
| Model | Type | AUC Higher is better | EER Lower is better | Accuracy Higher is better |
|---|---|---|---|---|
| DetectifAI v2 | DetectifAI | 0.9630 | 11.00% | 89.00% |
| DetectifAI v2-MoE | DetectifAI | 0.9420 | 13.00% | 87.00% |
| RawNet2 | Baseline | 0.8715 | 18.00% | 82.00% |
| DetectifAI v1 | DetectifAI | 0.8664 | 22.00% | 78.00% |
| DetectifAI v1-AX | DetectifAI | 0.8687 | 26.00% | 74.00% |
| Siamese-Net | Baseline | 0.5477 | 43.00% | 57.00% |
| SSL-W2V2 | Baseline | 0.5305 | 45.50% | 54.50% |
| SSL-WavLM | Baseline | 0.5084 | 48.50% | 51.50% |
Hindi holdout
808 samples · 400 real / 408 synthetic
| Model | Type | AUC Higher is better | EER Lower is better | Accuracy Higher is better |
|---|---|---|---|---|
| SSL-WavLM | Baseline | 0.9998 | 0.49% | 99.50% |
| SSL-W2V2 | Baseline | 0.9975 | 1.98% | 98.02% |
| DetectifAI v2 | DetectifAI | 0.9952 | 1.97% | 98.02% |
| DetectifAI v2-MoE | DetectifAI | 0.9839 | 5.94% | 94.06% |
| DetectifAI v1-AX | DetectifAI | 0.9838 | 6.93% | 93.07% |
| DetectifAI v1 | DetectifAI | 0.9632 | 8.42% | 91.58% |
| RawNet2 | Baseline | 0.9465 | 12.87% | 87.13% |
| Siamese-Net | Baseline | 0.9285 | 15.35% | 84.65% |
Performance by commercial synthesis source.
Equal error rate is shown for the four providers included in the original reports.
English per-provider EER
| Model | Narakeet | Resemble AI | Speechify | Voicemaker | Overall |
|---|---|---|---|---|---|
| DetectifAI v2DetectifAI | 8.00% | 16.50% | 12.50% | 7.00% | 11.00% |
| DetectifAI v2-MoEDetectifAI | 10.50% | 19.00% | 14.50% | 8.00% | 13.00% |
| RawNet2Baseline | 10.50% | 21.50% | 28.00% | 7.50% | 18.00% |
| DetectifAI v1DetectifAI | 18.00% | 27.00% | 22.00% | 15.00% | 22.00% |
| DetectifAI v1-AXDetectifAI | 20.00% | 30.00% | 25.00% | 19.00% | 26.00% |
| Siamese-NetBaseline | 43.00% | 40.50% | 40.50% | 43.50% | 43.00% |
| SSL-W2V2Baseline | 39.50% | 40.50% | 39.00% | 47.50% | 45.50% |
| SSL-WavLMBaseline | 46.50% | 45.00% | 48.00% | 47.00% | 48.50% |
Hindi per-provider EER
| Model | Narakeet | Resemble AI | Speechify | Voicemaker | Overall |
|---|---|---|---|---|---|
| SSL-WavLMBaseline | 0.00% | 0.00% | 1.00% | 0.00% | 0.49% |
| SSL-W2V2Baseline | 0.50% | 4.00% | 0.00% | 0.00% | 1.98% |
| DetectifAI v2DetectifAI | 0.50% | 2.00% | 1.50% | 0.50% | 1.97% |
| DetectifAI v2-MoEDetectifAI | 2.50% | 7.00% | 4.00% | 3.00% | 5.94% |
| DetectifAI v1-AXDetectifAI | 3.00% | 8.50% | 6.00% | 4.50% | 6.93% |
| DetectifAI v1DetectifAI | 4.00% | 10.00% | 7.00% | 5.00% | 8.42% |
| RawNet2Baseline | 7.00% | 15.00% | 19.50% | 5.70% | 12.87% |
| Siamese-NetBaseline | 7.50% | 7.00% | 23.50% | 9.56% | 15.35% |
Cross-language comparison.
The language gap is the absolute difference between English and Hindi EER.
| Model | Type | English EER | Hindi EER | Average EER | Language gap |
|---|---|---|---|---|---|
| DetectifAI v2 | DetectifAI | 11.00% | 1.97% | 6.49% | 9.03 pp |
| DetectifAI v2-MoE | DetectifAI | 13.00% | 5.94% | 8.97% | 7.06 pp |
| RawNet2 | Baseline | 18.00% | 12.87% | 15.44% | 5.13 pp |
Read the archive with its limits.
The recovered report describes 1,600 files, while its language tables total 1,608 samples.
Confusion counts in the source were normalised differently from the stated holdout sizes and are therefore omitted here.
A pinned code version, hardware record, and current review approval were not included.
The results predate DetectifAI v3 and should not be read as current production performance.
Historical continuity only. This report is not a current deployment, certification, or regulatory claim.