TechCrunchFeatured in TechCrunch · Read the story

Audio authentication models.
Built for real conversations.

We build models that detect synthetic speech and verify speaker identity. Use them across calls, recordings and voice agents, in the infrastructure you control.

Latency per clip
24msPer clip · vendor-reported
Detection accuracy
94.47%Scored by Podonos
Telephone channel
8kHzModel design target
Call screeningCALL 7F3A-21C9Inbound PSTN · G.711 µ-law · 8 kHz

As featured in

Press and recognition

Ecosystem support

Programmes and ecosystem partners

A few seconds of audio can clone a customer’s voice.

A familiar voice can now be generated from a short recording. Banks and contact centers need to verify the speech behind an account change, payment or identity check.

  1. Account takeover through the contact center

    Attack
    A caller uses a cloned voice to pass voice authentication or convince an agent, then resets credentials or changes the registered number.
    Control
    Aletheia scores the customer leg before any account change.
  2. Payment authorisation and impersonation

    Attack
    A cloned executive, relative or customer asks for an urgent transfer, or confirms a high-value payment by phone.
    Control
    High-risk instructions from a flagged voice are stepped up or blocked.
  3. Remote onboarding and identity verification

    Attack
    Synthetic or replayed speech is used to open accounts or pass liveness prompts during remote identity verification.
    Control
    Deepfake detection runs alongside Eidos and liveness checks.

#1 in commercial latency.

24 ms per clip, with 97.3% of deepfakes detected. Results from the Podonos benchmark across 23 systems.

Latency per clip
24ms#1 among commercial systems
Processing speed
2.3×the next-fastest commercial system
Deepfakes detected
97.3%of synthetic audio in the test
Files scored
4,524Every file. None declined.
Detection accuracy
94.47%#9 of 23 systems

Speed and detection accuracy

Higher is more accurate. Further right is faster.

Detectif.aiOther commercial systemsAccurate and fast: 90%+ accuracy, 10×+ real time
70%80%90%100%0×25×50×75×100×125×Accurate and fastDetectif.ai: 94.47% accuracy, 119.0× real-time speed, 24 ms per clip.Detectif.aiNII Synthetiq Audio v0.8-Beta: 89.57% accuracy, 41.7× real-time speed, 91 ms per clip.Corsound AI: 87.79% accuracy, 28.6× real-time speed, 180 ms per clip.Corsound AIPindrop: 95.05% accuracy, 13.2× real-time speed, 282 ms per clip.Hive: 83.53% accuracy, 2.9× real-time speed, 881 ms per clip.HiveWhispeak: 97.7% accuracy, 2.6× real-time speed, 1.1 s per clip.Reality Defender: 71.27% accuracy, 0.7× real-time speed, 5.7 s per clip.Reality Defender

Processing speed · 1× = real time

94.47% accuracyat 119× real-time speed

Commercial latency ranking

Average time per audio clip. Lower is faster.

  1. 1Detectif.ai24 ms
  2. 2Pella Research57 ms
  3. 3NII Synthetiq Audio91 ms
  4. 4Corsound AI180 ms
  5. 5Pindrop282 ms
  6. 6Resemble DETECT-World399 ms
  7. 7Fennura553 ms
  8. 8Hive881 ms
  9. 9Whispeak1.1 s
  10. 10Resemble AI1.2 s
  11. 11Reality Defender5.7 s

Reported timings · bars use a log scale

Accuracy scored by Podonos. Detectif.ai timings are vendor-reported; hardware and service conditions vary. The latency ranking includes the 11 commercial systems with published timing.

Podonos source · September 2026

How the models read a phone call.

Follow one cloned-voice call from the raw line signal to a decision. Every stage below is drawn from the same few seconds of audio.

Capture the call audio

The customer leg arrives as 8,000 samples a second, compressed with µ-law. Detection starts from this signal, exactly as the network delivers it.

Input
8 kHz mono · µ-law

Our models.

Four specialized models for speech authenticity, speaker identity, call classification and conversation.

Aletheia · Truth

Deepfake detection

Analyze live calls and recordings for synthetic, cloned or replayed speech. Return a synthetic score and identify the segments that require review.

  • Synthetic speech
  • Voice cloning
  • Replay detection
Input
Live or recorded audio
Output
Synthetic score + flagged segments
Explore deepfake detection
AuthenticityExample
Flagged speech
Synthetic score0.93

Synthetic speech detected

Applications
  • Live call screening
  • Recording analysis
  • Remote onboarding

Integrations.

Stream live calls from your telephony platform, add signals to a voice agent, score recordings, or run the detector on the handset itself.

Fork the customer leg from your telephony stack.

Send the caller’s audio as it arrives and act on the signal in your own routing, without changing what the agent hears.

  • SIP and RTP media forks
  • Contact-center audio streams
  • Agent desktop alerts
stream-call.js
// Stream the customer leg of a live call
const session = await detectif.calls.stream({
  callId: "7F3A-21C9",
  audio: { encoding: "mulaw", sampleRate: 8000 },
  checks: ["authenticity", "speaker"],
  speakerReference: customer.voiceprintId,
});

customerLeg.on("audio", (frame) => session.send(frame));

session.on("signal", (signal) => {
  if (signal.authenticity.synthetic >= 0.5) {
    fraudDesk.route(call, signal); // your policy
  }
});
Signal returnedapplication/json
{
  "callId": "7F3A-21C9",
  "at": 6.85,
  "authenticity": { "synthetic": 0.92, "segments": [[5.50, 10.08]] },
  "speaker": { "match": 0.86 }
}

Example syntax. Production interfaces are provided during evaluation.

Voice agents.

Classify the recipient, verify the caller and run the approved workflow. Transfer the call and its context when a specialist is needed.

  1. Noesis

    Classify the recipient

    Identify a person, IVR, voicemail or another agent and select the appropriate route.

  2. Aletheia + Eidos

    Check the caller

    Inspect authenticity and compare the enrolled voice before a sensitive request.

  3. Workflow rules

    Apply access controls

    Check permitted actions, tool access and additional verification requirements.

  4. Lexis

    Act or hand off

    Continue in scope, or pass the conversation and the evidence to your team.

Lexis · payment confirmation
ConversationExample
  1. NoesisHuman recipient · conversation route selected
  2. Lexis

    Thanks for calling. How can I help today?

  3. Caller

    I need to move $5,000 to a new payee.

  4. Access rulesNew payee · additional verification required
  5. AletheiaSynthetic score 0.93 · voice flagged
  6. EidosSpeaker match 0.86 · authenticity check failed
  7. Lexis

    I’ll connect you with a specialist to review this request.

  8. OutcomePayment held · call and evidence sent to a specialist

Deployment options.

Use the hosted API, deploy in a private VPC, run on your own servers or embed detection on a device. Compare the data path and operating requirements.

Where the audio goesAudio crosses into our cloud
Your systems
Your telephonySIP · contact center · application
Encrypted audio
Detectif.ai cloud
Detectif.ai runtimeManaged and scaled by us
Signal response
Your systems
Your policy engineContinue · step up · review
Your infrastructureDetectif.ai model runtime
Hosted API

Start with a managed runtime.

Stream the caller’s audio to Detectif.ai over an encrypted connection. Receive a signal in your application and apply your own decision policy.

Audio processed in
Detectif.ai-managed cloud
Operated by
Detectif.ai
Network path
Encrypted connection to Detectif.ai
Release control
Runtime and model updates managed by us

A straightforward starting point for evaluation and cloud-based call flows.

Study reference: 14.4 ms median per clip on an NVIDIA L4 GPU, FP16.Read the deployment study

Use cases.

Add voice checks to onboarding, payments, account recovery and customer service. Explore the call flow for each application.

Choose a call flow
Banks · credit unions · fintech

Voice checks for KYC onboarding.

Add voice authenticity to remote onboarding and identity verification. Inspect the applicant’s speech alongside your document and liveness checks, then enroll a voice reference when the identity process is complete.

  1. 01Receive the applicant’s speech
  2. 02Check for synthetic or replayed audio
  3. 03Complete KYC, then enroll a reference
Example
Identity document
Government-issued IDPassport or driver’s license
Details maskedIllustrative example
Caller request
“I’d like to open an account.”
Voice authenticitySynthetic score · 0.04
Speaker enrollmentEnroll after your KYC checks
Next stepContinue your onboarding checks
Built for your team
  • Banks
  • Lenders and credit unions
  • Insurers
  • Payments and fintech
  • Contact centers and BPOs
  • Voice AI platforms

Voice model research.

Studies in forensic representations, language transfer and deployment. Explore the methods, results and limitations.

All research
Editorial visualization of speech traces passing through glass planes in an acoustic research environment.
Speech forensics2026 · Working preprint

Generator traces across languages

A language-disjoint evaluation of forensic representations. Train on 12 languages, then test the same generators on six unseen languages.

91.0%Top-1 generator retrieval on languages excluded from training
62 synthesis systems · 18 languages14 pagesRead study
Editorial visualization of an audio waveform connecting server hardware and a mobile device in a quiet architectural environment.
Model deployment2026 · Working preprint

Deepfake detection across devices

One detector, seven deployment configurations. Measure processing capacity and detection quality on GPU, server CPU and Android.

99.65%Of handset test clips processed faster than real time
7 configurations · 4,000 clips each25 pagesRead study

Questions from risk, fraud and engineering teams.

Straight answers, with the evidence behind each one.

What is voice deepfake detection?

Voice deepfake detection identifies speech that was generated or cloned by AI, or replayed from a recording, instead of spoken live by a person. Detectif.ai’s models score audio for these signs and return an authenticity signal that an application or agent can act on.

How do banks detect cloned voices on phone calls?

By scoring the customer’s side of the call for synthetic speech while it is happening, alongside speaker verification. A good clone can match a voiceprint, so authenticity has to be checked separately before sensitive actions such as payments, credential resets or changes to contact details.

Does it work on real phone-line audio?

Telephone audio is the design target: 8 kHz speech after codecs, packet loss, resampling and background noise, which is how fraud actually arrives. Performance in any deployment depends on channel, language and generator, and is measured on your own traffic during evaluation.

How fast is Detectif.ai?

On the independent Podonos audio deepfake detection benchmark, Detectif.ai reported a mean latency of 24 ms per clip, the lowest of any commercial system that submitted timing. Latency is vendor-reported and depends on hardware and deployment.

How accurate is it?

Podonos scored Detectif.ai at 94.47% accuracy on 4,524 files against labels it holds privately, 9th of 23 systems. 97.3% of synthetic files were caught and 8.4% of genuine files were flagged. Results on your traffic are measured in an evaluation before deployment.

Which languages does it support?

The models are built to be language-agnostic. In Detectif.ai’s cross-lingual study, forensic representations trained on 12 languages reached 91.0% top-1 generator retrieval on 6 languages held out of training. That study measures attribution among known generators; detection performance per language is confirmed during evaluation.

Can it run on-premises or on a device?

Yes. The same models can be hosted, deployed inside your own infrastructure, or run on device. The deployment study measured one detector on an NVIDIA L4 GPU, a server CPU and an Android handset, where 99.65% of clips were processed faster than real time.

Is call audio stored?

Retention is set per deployment. Audio can be analysed without being stored, or the flagged segment can be kept with the model record for review when your policy and regulations allow it.

Test it on your own call traffic.

A 30-minute walkthrough of your highest-risk call flow, then an evaluation on audio from your own lines.

Product filmA live call: agent, speaker verification, decision.Watch on X