Generator traces across languages
A language-disjoint evaluation of forensic representations. Train on 12 languages, then test the same generators on six unseen languages.
We build models that detect synthetic speech and verify speaker identity. Use them across calls, recordings and voice agents, in the infrastructure you control.
A familiar voice can now be generated from a short recording. Banks and contact centers need to verify the speech behind an account change, payment or identity check.
24 ms per clip, with 97.3% of deepfakes detected. Results from the Podonos benchmark across 23 systems.
Higher is more accurate. Further right is faster.
Processing speed · 1× = real time
Average time per audio clip. Lower is faster.
Reported timings · bars use a log scale
Accuracy scored by Podonos. Detectif.ai timings are vendor-reported; hardware and service conditions vary. The latency ranking includes the 11 commercial systems with published timing.
Podonos source · September 2026Follow one cloned-voice call from the raw line signal to a decision. Every stage below is drawn from the same few seconds of audio.
The customer leg arrives as 8,000 samples a second, compressed with µ-law. Detection starts from this signal, exactly as the network delivers it.
Four specialized models for speech authenticity, speaker identity, call classification and conversation.
Analyze live calls and recordings for synthetic, cloned or replayed speech. Return a synthetic score and identify the segments that require review.
Synthetic speech detected
Stream live calls from your telephony platform, add signals to a voice agent, score recordings, or run the detector on the handset itself.
Send the caller’s audio as it arrives and act on the signal in your own routing, without changing what the agent hears.
// Stream the customer leg of a live call
const session = await detectif.calls.stream({
callId: "7F3A-21C9",
audio: { encoding: "mulaw", sampleRate: 8000 },
checks: ["authenticity", "speaker"],
speakerReference: customer.voiceprintId,
});
customerLeg.on("audio", (frame) => session.send(frame));
session.on("signal", (signal) => {
if (signal.authenticity.synthetic >= 0.5) {
fraudDesk.route(call, signal); // your policy
}
});
{
"callId": "7F3A-21C9",
"at": 6.85,
"authenticity": { "synthetic": 0.92, "segments": [[5.50, 10.08]] },
"speaker": { "match": 0.86 }
}Example syntax. Production interfaces are provided during evaluation.
Classify the recipient, verify the caller and run the approved workflow. Transfer the call and its context when a specialist is needed.
Identify a person, IVR, voicemail or another agent and select the appropriate route.
Inspect authenticity and compare the enrolled voice before a sensitive request.
Check permitted actions, tool access and additional verification requirements.
Continue in scope, or pass the conversation and the evidence to your team.
Use the hosted API, deploy in a private VPC, run on your own servers or embed detection on a device. Compare the data path and operating requirements.
Stream the caller’s audio to Detectif.ai over an encrypted connection. Receive a signal in your application and apply your own decision policy.
A straightforward starting point for evaluation and cloud-based call flows.
Add voice checks to onboarding, payments, account recovery and customer service. Explore the call flow for each application.
Add voice authenticity to remote onboarding and identity verification. Inspect the applicant’s speech alongside your document and liveness checks, then enroll a voice reference when the identity process is complete.
“I’d like to open an account.”
Studies in forensic representations, language transfer and deployment. Explore the methods, results and limitations.
All researchA language-disjoint evaluation of forensic representations. Train on 12 languages, then test the same generators on six unseen languages.
One detector, seven deployment configurations. Measure processing capacity and detection quality on GPU, server CPU and Android.
Straight answers, with the evidence behind each one.
Voice deepfake detection identifies speech that was generated or cloned by AI, or replayed from a recording, instead of spoken live by a person. Detectif.ai’s models score audio for these signs and return an authenticity signal that an application or agent can act on.
By scoring the customer’s side of the call for synthetic speech while it is happening, alongside speaker verification. A good clone can match a voiceprint, so authenticity has to be checked separately before sensitive actions such as payments, credential resets or changes to contact details.
Telephone audio is the design target: 8 kHz speech after codecs, packet loss, resampling and background noise, which is how fraud actually arrives. Performance in any deployment depends on channel, language and generator, and is measured on your own traffic during evaluation.
On the independent Podonos audio deepfake detection benchmark, Detectif.ai reported a mean latency of 24 ms per clip, the lowest of any commercial system that submitted timing. Latency is vendor-reported and depends on hardware and deployment.
Podonos scored Detectif.ai at 94.47% accuracy on 4,524 files against labels it holds privately, 9th of 23 systems. 97.3% of synthetic files were caught and 8.4% of genuine files were flagged. Results on your traffic are measured in an evaluation before deployment.
The models are built to be language-agnostic. In Detectif.ai’s cross-lingual study, forensic representations trained on 12 languages reached 91.0% top-1 generator retrieval on 6 languages held out of training. That study measures attribution among known generators; detection performance per language is confirmed during evaluation.
Yes. The same models can be hosted, deployed inside your own infrastructure, or run on device. The deployment study measured one detector on an NVIDIA L4 GPU, a server CPU and an Android handset, where 99.65% of clips were processed faster than real time.
Retention is set per deployment. Audio can be analysed without being stored, or the flagged segment can be kept with the model record for review when your policy and regulations allow it.
A 30-minute walkthrough of your highest-risk call flow, then an evaluation on audio from your own lines.