All papers

Research paper / Working preprint / 2026

Do generator artifacts survive a language change?

Do Generator Artifacts Survive a Language Change? A Language-Disjoint Study of Forensic Representations for Speech Deepfakes

  • Source tracing
  • Forensic embeddings
  • Cross-lingual control
Training
12 languages, 62 generators
Holdout
6 unseen languages
Clips
5.43M train / 1.36M holdout
Audio
~17,400 hours
Protocol
Language-disjoint retrieval

What the paper asks.

Speech deepfake forensics rests on the premise that a synthesis system leaves a persistent, generator-specific trace in its output. That premise is rarely tested against its most obvious confound: the language being spoken.

A representation that appears to encode which generator produced a clip may in practice be encoding which language, corpus and phone inventory it contains, because generators are almost always evaluated on the languages they were trained on.

This study separates the two with a language-disjoint protocol. A forensic embedding is trained with a supervised contrastive objective whose label is generator identity and whose positives are drawn across languages, on synthetic speech only, from 12 languages. It is then evaluated on 6 languages removed from the corpus before training, with generator identity held fixed, so the evaluation changes exactly one variable.

Generator retrieval holds at 91.0% top-1 and 98.8% top-5 on languages never observed, against 94.2% in-language and a 3.7% floor for an embedding carrying no generator information: a 3.2-point cost for a complete linguistic substitution.

Read more detail

Training uses 5,431,240 synthetic clips from 62 generators across 12 languages (about 13,950 hours). No real speech enters the pipeline at any stage, which removes real/fake separability as a shortcut and forces the representation's capacity onto inter-generator distinctions.

The holdout is not a convenience sample: Russian, Japanese, Korean, Swedish, Thai and Vietnamese span four writing systems and three language families, and include two tonal languages where only one training language is tonal. No holdout audio enters the pipeline at any point, not for training, not for validation, not for front-end statistics.

The paper introduces the language entanglement ratio, the across-language centroid spread of a generator normalized by the inter-generator distance scale. On unseen languages it halves over training (0.127 to 0.064) while holdout retrieval simultaneously rises, so language is being factored out rather than traded against discriminability.

Residual confusions on the holdout are organized by vendor and model family, not by language: ChatGPT-TTS with OpenAI TTS-1 HD, MOSS-1.7B with MOSS-8B, Qwen3-0.6B with Qwen3-1.7B. That is the error structure a genuine forensic trace predicts and a linguistic shortcut does not.

The paper is equally explicit about where the effect is weaker. Intra-cluster compactness transfers poorly, per-language accuracy varies by 13.7 points across the holdout, the separation ratio collapses from 173.5 to 16.7 for reasons that are an artifact of the metric rather than of generalization, and the study evaluates seen generators in unseen languages, which is a different and strictly easier question than open-set attribution.

Headline measurements.

91.0%Top-1 generator retrieval on unseen languages
98.8%Top-5 on unseen languages
3.2 ptsCost of a complete language substitution
24.6xAbove the 3.7% no-information floor
0.064Language entanglement ratio on the holdout
62Synthesis systems as class labels
6Languages excluded before training
0%Real speech used at any stage

Fix the generator. Change the language.

The representation

130M-parameter spectrogram transformer, 256-D L2-normalized forensic embedding

  • Supervised contrastive loss, temperature 0.07
  • Class label is the synthesis system, not real vs fake
  • Positives are same-generator clips drawn across languages
  • Negatives include 15 hard generators in every step
  • Fake-only training: no real/fake axis to fall back on

The control

  • 18 languages split 12 / 6 before the corpus is indexed
  • Holdout spans 4 scripts, 3 families, 2 tonal languages
  • Generators held fixed: every holdout system was seen in training
  • Evaluated purely by leave-one-out cosine retrieval, no test-time classifier
  • 81,920 embeddings per split per epoch, no parameter fitting

Whatever generator structure survives a complete linguistic substitution cannot be explained by the model having memorized the acoustics of the test language.

Three evaluation tiers.

Top-1 generator retrieval at the terminal checkpoint. Train to in-language costs 0.27 points; in-language to language-disjoint costs 3.17. Bars start at 80%.

TR / train94.46%top-1 generator retrieval

12 seen languages, seen clips

rho-lang 0.0164

The ceiling of what the objective can reach on data it optimized against.

VA / in-language94.19%top-1 generator retrieval

12 seen languages, unseen clips

rho-lang 0.0204

Clip-level novelty costs 0.27 points. Almost the entire gap is language, not unseen data.

HO / language-disjoint91.02%top-1 generator retrieval

6 unseen languages, same generators

rho-lang 0.0636

One variable changes. Top-5 stays at 98.80% and MRR at 0.948.

Findings.

The trace survives the substitution

Under a complete replacement of the linguistic content, top-1 generator retrieval falls from 94.19% in-language to 91.02%, a gap of 3.17 points, while top-5 holds at 98.80% and MRR at 0.948. Against the analytical floor of 3.70% for an embedding carrying no generator information, the holdout result is 24.6x chance and retains 96.5% of the recoverable in-language performance. The near-saturated top-5 matters operationally: when the nearest neighbour is wrong, the correct generator is almost always still in the shortlist.

91.02%
Holdout top-1
98.80%
Holdout top-5
0.965
Transfer index vs in-language

Geometry: position transfers, tightness does not

The rows of the results table do not degrade uniformly, and the pattern is the informative part. Inter-generator scale is essentially unaffected: d-inter falls 1.8% from validation to holdout, so generators do not collapse toward each other when the language changes. Intra-generator compactness degrades 10x. Clusters stay in the right places but become much looser, and that single term drives the separation ratio from 173.5 to 16.7 without any loss of inter-generator structure.

1.8%
Drop in inter-generator distance
10x
Inflation in intra-cluster spread
173.5 to 16.7
Separation ratio, a misleading metric

Language is factored out, not merely tolerated

The language entanglement ratio is the across-language centroid spread of a generator divided by the inter-generator distance scale, so it says how far moving a generator to another language displaces it relative to the distance to its neighbours. On the holdout it halves over the analysed window, from 0.127 at epoch 6 to 0.064 at epoch 19, while holdout retrieval rises 3.1 points. Nothing in the objective ever saw those languages, so the reduction cannot come from fitting them.

0.127 to 0.064
rho-lang on unseen languages
39%
Reduction in absolute language spread
+3.1 pts
Holdout top-1 over the same window

Errors follow lineage, not language

Of the 351 generator pairs on the holdout, the great majority are well separated. The residual dark structure is organized by provenance: ChatGPT-TTS with OpenAI TTS-1 HD as the closest pair, MOSS-1.7B with MOSS-8B and Qwen3-0.6B with Qwen3-1.7B as tight same-family blocks, and a cluster of kokoro, RVC, griffin_lim and minimax_speech-02-turbo that plausibly share a waveform-generation stage. If the embedding were encoding language, confusions would cluster by which holdout language each system emits. They do not.

351
Generator pairs on the holdout
Vendor / family
Axis the residual errors follow
xtts v1.1 / v2
Same lineage, still separated

Per-language difficulty is a gallery effect

Aggregate holdout accuracy conceals a 13.7-point spread across the six unseen languages. Thai reaches 97.2% and Vietnamese 96.0%, despite both being tonal and typologically remote from an Indo-European-heavy training set, while Russian sits at 85.5% and Korean at 83.5%. Typological distance does not predict forensic difficulty; the number of generators competing within each language does. The same effect is starker in-language, where Portuguese scores lowest at 84.7% despite sitting in the same data-volume band as languages that saturate near 99.8%.

13.7 pts
Spread across holdout languages
97.2% / 83.5%
Thai best, Korean worst
Gallery density
The confound, not data volume

What the study does not establish

The paper states its limits as plainly as its results. No trained ablations or independent baselines were run, so cross-language positives coincide with language robustness rather than being shown to cause it. The holdout shares its label space with training, so this says nothing about open-set attribution. All audio comes from one collection pipeline, so corpus-level factors are untested, and language remains confounded with speaker and content. One seed, one configuration, and the final epoch is not the best one: holdout top-1 peaks at 91.46% at epoch 17.

5
Ablations named as required before submission
Seen generators
Label space of the holdout
1 seed
No confidence intervals reported

Partitioning by language costs almost nothing and tells you what accuracy cannot.

3.2 ptsCost of replacing the language entirely
1 passCost of running the control

Any multilingual forensic corpus can be split by language, and the resulting in-language to language-disjoint contrast costs one extra evaluation pass while answering a question aggregate open-set numbers cannot: how much of the reported forensic performance is linguistic bookkeeping. Source tracing systems are increasingly proposed for investigative use, where a questioned recording may well be in a language absent from any available training corpus. These results suggest at least some current representations would survive that control, but that is a hypothesis to test per system, not an assumption.

Read the paper.

Full text, 14 pages, rendered here on the page. No download.

1/ 14
100%
Page 1 of 14
Page 2 of 14
Page 3 of 14
Page 4 of 14
Page 5 of 14
Page 6 of 14
Page 7 of 14
Page 8 of 14
Page 9 of 14
Page 10 of 14
Page 11 of 14
Page 12 of 14
Page 13 of 14
Page 14 of 14