The Problem
Financial fraud relies on cheap, real-time AI voice clones. Current detection models were either too slow for live phone calls or too massive to run without expensive GPU clusters. We needed a system capable of analyzing live WebRTC and telephony streams continuously without introducing perceptible lag.
The Hard Constraint
We had a strict 350ms latency budget to intercept, process, and classify live audio streams on a pure CPU architecture, while keeping all data in ephemeral memory to comply with strict privacy laws.
Architecture
Results & Caveats
The Outcome
Achieved strict sub-second latency running entirely on CPU. Delivered a fully functional, zero-cost, real-time deepfake detection system that successfully intercepts live streams, proving the architecture for the Smart India Hackathon.
Honest Caveats
While the architecture proved the concept is viable, it is not ready for deployment at a massive enterprise. To handle 10,000+ simultaneous calls, the Python asyncio backend would bottleneck and require migration to a multithreaded language like Go or Rust.