Back to Index
PyTorchRedis PubSubWebRTCFastAPI

VoiceTrace

The Problem

Financial fraud relies on cheap, real-time AI voice clones. Current detection models were either too slow for live phone calls or too massive to run without expensive GPU clusters. We needed a system capable of analyzing live WebRTC and telephony streams continuously without introducing perceptible lag.

The Hard Constraint

We had a strict 350ms latency budget to intercept, process, and classify live audio streams on a pure CPU architecture, while keeping all data in ephemeral memory to comply with strict privacy laws.

Architecture

[WebRTC / Twilio Stream] │ ▼ (1-sec chunks) [In-Memory Redis PubSub] │ ▼ [Dynamic Batch Worker] ──► [AASIST-L (Deepfake)] ──► [ECAPA-TDNN (Speaker)] ──► [SpeechBrain (Liveness)] │ ▼ [Risk Score Aggregation] │ ▼ (Returned < 350ms) [Client Dashboard]

Results & Caveats

The Outcome

Achieved strict sub-second latency running entirely on CPU. Delivered a fully functional, zero-cost, real-time deepfake detection system that successfully intercepts live streams, proving the architecture for the Smart India Hackathon.

Honest Caveats

While the architecture proved the concept is viable, it is not ready for deployment at a massive enterprise. To handle 10,000+ simultaneous calls, the Python asyncio backend would bottleneck and require migration to a multithreaded language like Go or Rust.