A larger multilingual robustness evaluation.

The second archived report expanded the DetectifAI model family while retaining the English and Hindi commercial-TTS stress tests.

Historical report · v2 · March 2026
Dataset
MLADD-v2 expanded evaluation
Reported scale
1,600 audio files
Languages
English and Hindi
Synthesis sources
4 commercial TTS providers
Status
Archived · superseded by v3

Reported performance across both languages.

These figures are restored from the original preorder-site report. DetectifAI rows are highlighted; baseline rows remain neutral.

English

EER · lower is better
0%25%50%

Hindi

EER · lower is better
0%25%50%

English holdout

800 samples · 400 real / 400 synthetic

Archived as reported
English historical benchmark results
ModelTypeAUC Higher is betterEER Lower is betterAccuracy Higher is better
DetectifAI v2DetectifAI0.963011.00%89.00%
DetectifAI v2-MoEDetectifAI0.942013.00%87.00%
RawNet2Baseline0.871518.00%82.00%
DetectifAI v1DetectifAI0.866422.00%78.00%
DetectifAI v1-AXDetectifAI0.868726.00%74.00%
Siamese-NetBaseline0.547743.00%57.00%
SSL-W2V2Baseline0.530545.50%54.50%
SSL-WavLMBaseline0.508448.50%51.50%

Hindi holdout

808 samples · 400 real / 408 synthetic

Archived as reported
Hindi historical benchmark results
ModelTypeAUC Higher is betterEER Lower is betterAccuracy Higher is better
SSL-WavLMBaseline0.99980.49%99.50%
SSL-W2V2Baseline0.99751.98%98.02%
DetectifAI v2DetectifAI0.99521.97%98.02%
DetectifAI v2-MoEDetectifAI0.98395.94%94.06%
DetectifAI v1-AXDetectifAI0.98386.93%93.07%
DetectifAI v1DetectifAI0.96328.42%91.58%
RawNet2Baseline0.946512.87%87.13%
Siamese-NetBaseline0.928515.35%84.65%

Performance by commercial synthesis source.

Equal error rate is shown for the four providers included in the original reports.

English per-provider EER

English equal error rate by synthesis provider
ModelNarakeetResemble AISpeechifyVoicemakerOverall
DetectifAI v2DetectifAI8.00%16.50%12.50%7.00%11.00%
DetectifAI v2-MoEDetectifAI10.50%19.00%14.50%8.00%13.00%
RawNet2Baseline10.50%21.50%28.00%7.50%18.00%
DetectifAI v1DetectifAI18.00%27.00%22.00%15.00%22.00%
DetectifAI v1-AXDetectifAI20.00%30.00%25.00%19.00%26.00%
Siamese-NetBaseline43.00%40.50%40.50%43.50%43.00%
SSL-W2V2Baseline39.50%40.50%39.00%47.50%45.50%
SSL-WavLMBaseline46.50%45.00%48.00%47.00%48.50%

Hindi per-provider EER

Hindi equal error rate by synthesis provider
ModelNarakeetResemble AISpeechifyVoicemakerOverall
SSL-WavLMBaseline0.00%0.00%1.00%0.00%0.49%
SSL-W2V2Baseline0.50%4.00%0.00%0.00%1.98%
DetectifAI v2DetectifAI0.50%2.00%1.50%0.50%1.97%
DetectifAI v2-MoEDetectifAI2.50%7.00%4.00%3.00%5.94%
DetectifAI v1-AXDetectifAI3.00%8.50%6.00%4.50%6.93%
DetectifAI v1DetectifAI4.00%10.00%7.00%5.00%8.42%
RawNet2Baseline7.00%15.00%19.50%5.70%12.87%
Siamese-NetBaseline7.50%7.00%23.50%9.56%15.35%

Cross-language comparison.

The language gap is the absolute difference between English and Hindi EER.

Cross-language equal error rate comparison
ModelTypeEnglish EERHindi EERAverage EERLanguage gap
DetectifAI v2DetectifAI11.00%1.97%6.49%9.03 pp
DetectifAI v2-MoEDetectifAI13.00%5.94%8.97%7.06 pp
RawNet2Baseline18.00%12.87%15.44%5.13 pp

Read the archive with its limits.

  • The recovered report describes 1,600 files, while its language tables total 1,608 samples.

  • Confusion counts in the source were normalised differently from the stated holdout sizes and are therefore omitted here.

  • A pinned code version, hardware record, and current review approval were not included.

  • The results predate DetectifAI v3 and should not be read as current production performance.

Historical continuity only. This report is not a current deployment, certification, or regulatory claim.