Research paper / Working preprint / 2026
Do generator artifacts survive a language change?
Do Generator Artifacts Survive a Language Change? A Language-Disjoint Study of Forensic Representations for Speech Deepfakes
- Source tracing
- Forensic embeddings
- Cross-lingual control
- Training
- 12 languages, 62 generators
- Holdout
- 6 unseen languages
- Clips
- 5.43M train / 1.36M holdout
- Audio
- ~17,400 hours
- Protocol
- Language-disjoint retrieval
What the paper asks.
Speech deepfake forensics rests on the premise that a synthesis system leaves a persistent, generator-specific trace in its output. That premise is rarely tested against its most obvious confound: the language being spoken.
A representation that appears to encode which generator produced a clip may in practice be encoding which language, corpus and phone inventory it contains, because generators are almost always evaluated on the languages they were trained on.
This study separates the two with a language-disjoint protocol. A forensic embedding is trained with a supervised contrastive objective whose label is generator identity and whose positives are drawn across languages, on synthetic speech only, from 12 languages. It is then evaluated on 6 languages removed from the corpus before training, with generator identity held fixed, so the evaluation changes exactly one variable.
Generator retrieval holds at 91.0% top-1 and 98.8% top-5 on languages never observed, against 94.2% in-language and a 3.7% floor for an embedding carrying no generator information: a 3.2-point cost for a complete linguistic substitution.
Read more detail+
Training uses 5,431,240 synthetic clips from 62 generators across 12 languages (about 13,950 hours). No real speech enters the pipeline at any stage, which removes real/fake separability as a shortcut and forces the representation's capacity onto inter-generator distinctions.
The holdout is not a convenience sample: Russian, Japanese, Korean, Swedish, Thai and Vietnamese span four writing systems and three language families, and include two tonal languages where only one training language is tonal. No holdout audio enters the pipeline at any point, not for training, not for validation, not for front-end statistics.
The paper introduces the language entanglement ratio, the across-language centroid spread of a generator normalized by the inter-generator distance scale. On unseen languages it halves over training (0.127 to 0.064) while holdout retrieval simultaneously rises, so language is being factored out rather than traded against discriminability.
Residual confusions on the holdout are organized by vendor and model family, not by language: ChatGPT-TTS with OpenAI TTS-1 HD, MOSS-1.7B with MOSS-8B, Qwen3-0.6B with Qwen3-1.7B. That is the error structure a genuine forensic trace predicts and a linguistic shortcut does not.
The paper is equally explicit about where the effect is weaker. Intra-cluster compactness transfers poorly, per-language accuracy varies by 13.7 points across the holdout, the separation ratio collapses from 173.5 to 16.7 for reasons that are an artifact of the metric rather than of generalization, and the study evaluates seen generators in unseen languages, which is a different and strictly easier question than open-set attribution.
Headline measurements.
Fix the generator. Change the language.
The representation
130M-parameter spectrogram transformer, 256-D L2-normalized forensic embedding
- Supervised contrastive loss, temperature 0.07
- Class label is the synthesis system, not real vs fake
- Positives are same-generator clips drawn across languages
- Negatives include 15 hard generators in every step
- Fake-only training: no real/fake axis to fall back on
The control
- 18 languages split 12 / 6 before the corpus is indexed
- Holdout spans 4 scripts, 3 families, 2 tonal languages
- Generators held fixed: every holdout system was seen in training
- Evaluated purely by leave-one-out cosine retrieval, no test-time classifier
- 81,920 embeddings per split per epoch, no parameter fitting
Whatever generator structure survives a complete linguistic substitution cannot be explained by the model having memorized the acoustics of the test language.
Three evaluation tiers.
Top-1 generator retrieval at the terminal checkpoint. Train to in-language costs 0.27 points; in-language to language-disjoint costs 3.17. Bars start at 80%.
12 seen languages, seen clips
rho-lang 0.0164
The ceiling of what the objective can reach on data it optimized against.
12 seen languages, unseen clips
rho-lang 0.0204
Clip-level novelty costs 0.27 points. Almost the entire gap is language, not unseen data.
6 unseen languages, same generators
rho-lang 0.0636
One variable changes. Top-5 stays at 98.80% and MRR at 0.948.
Findings.
The trace survives the substitution
Under a complete replacement of the linguistic content, top-1 generator retrieval falls from 94.19% in-language to 91.02%, a gap of 3.17 points, while top-5 holds at 98.80% and MRR at 0.948. Against the analytical floor of 3.70% for an embedding carrying no generator information, the holdout result is 24.6x chance and retains 96.5% of the recoverable in-language performance. The near-saturated top-5 matters operationally: when the nearest neighbour is wrong, the correct generator is almost always still in the shortlist.
- 91.02%
- Holdout top-1
- 98.80%
- Holdout top-5
- 0.965
- Transfer index vs in-language
Geometry: position transfers, tightness does not
The rows of the results table do not degrade uniformly, and the pattern is the informative part. Inter-generator scale is essentially unaffected: d-inter falls 1.8% from validation to holdout, so generators do not collapse toward each other when the language changes. Intra-generator compactness degrades 10x. Clusters stay in the right places but become much looser, and that single term drives the separation ratio from 173.5 to 16.7 without any loss of inter-generator structure.
- 1.8%
- Drop in inter-generator distance
- 10x
- Inflation in intra-cluster spread
- 173.5 to 16.7
- Separation ratio, a misleading metric
Language is factored out, not merely tolerated
The language entanglement ratio is the across-language centroid spread of a generator divided by the inter-generator distance scale, so it says how far moving a generator to another language displaces it relative to the distance to its neighbours. On the holdout it halves over the analysed window, from 0.127 at epoch 6 to 0.064 at epoch 19, while holdout retrieval rises 3.1 points. Nothing in the objective ever saw those languages, so the reduction cannot come from fitting them.
- 0.127 to 0.064
- rho-lang on unseen languages
- 39%
- Reduction in absolute language spread
- +3.1 pts
- Holdout top-1 over the same window
Errors follow lineage, not language
Of the 351 generator pairs on the holdout, the great majority are well separated. The residual dark structure is organized by provenance: ChatGPT-TTS with OpenAI TTS-1 HD as the closest pair, MOSS-1.7B with MOSS-8B and Qwen3-0.6B with Qwen3-1.7B as tight same-family blocks, and a cluster of kokoro, RVC, griffin_lim and minimax_speech-02-turbo that plausibly share a waveform-generation stage. If the embedding were encoding language, confusions would cluster by which holdout language each system emits. They do not.
- 351
- Generator pairs on the holdout
- Vendor / family
- Axis the residual errors follow
- xtts v1.1 / v2
- Same lineage, still separated
Per-language difficulty is a gallery effect
Aggregate holdout accuracy conceals a 13.7-point spread across the six unseen languages. Thai reaches 97.2% and Vietnamese 96.0%, despite both being tonal and typologically remote from an Indo-European-heavy training set, while Russian sits at 85.5% and Korean at 83.5%. Typological distance does not predict forensic difficulty; the number of generators competing within each language does. The same effect is starker in-language, where Portuguese scores lowest at 84.7% despite sitting in the same data-volume band as languages that saturate near 99.8%.
- 13.7 pts
- Spread across holdout languages
- 97.2% / 83.5%
- Thai best, Korean worst
- Gallery density
- The confound, not data volume
What the study does not establish
The paper states its limits as plainly as its results. No trained ablations or independent baselines were run, so cross-language positives coincide with language robustness rather than being shown to cause it. The holdout shares its label space with training, so this says nothing about open-set attribution. All audio comes from one collection pipeline, so corpus-level factors are untested, and language remains confounded with speaker and content. One seed, one configuration, and the final epoch is not the best one: holdout top-1 peaks at 91.46% at epoch 17.
- 5
- Ablations named as required before submission
- Seen generators
- Label space of the holdout
- 1 seed
- No confidence intervals reported
Partitioning by language costs almost nothing and tells you what accuracy cannot.
Any multilingual forensic corpus can be split by language, and the resulting in-language to language-disjoint contrast costs one extra evaluation pass while answering a question aggregate open-set numbers cannot: how much of the reported forensic performance is linguistic bookkeeping. Source tracing systems are increasingly proposed for investigative use, where a questioned recording may well be in a language absent from any available training corpus. These results suggest at least some current representations would survive that control, but that is a hypothesis to test per system, not an assumption.
Read the paper.
Full text, 14 pages, rendered here on the page. No download.













