Modulate Raises $25M to Fight Deepfake Voices With Next-Generation Audio AI

Blog WriterCybersecurity News - Original News Source is cybersecuritynews.com


Modulate has secured $25 million in funding to expand its audio-native AI platform as enterprises confront deepfake voice fraud, impersonation, and unsafe automated conversations.

Future Ventures led the round, joined by Hyperplane and Lakestar, bringing the company’s funding total to $60 million.

The investment will accelerate AI research, engineering, developer tools, partnerships, and deployment options for organizations integrating voice intelligence into security and communications workflows.

The financing arrives as convincing synthetic speech makes telephone cues unreliable. Instead of examining only a transcript, Modulate’s technology analyzes how something is spoken, including emotion, tone, emphasis, intent, conversational behavior, and evidence of synthetic generation.

That distinction matters in vishing and executive-impersonation attacks, where the words may appear harmless while urgency, manipulation, or an artificial voice exposes the threat.

Modulate Raises $25M Fund

According to the announcement published by Modulate, the company’s offering centers on Velma, a real-time conversation-understanding platform designed to detect events such as fraud attempts, harassment, customer dissatisfaction, policy violations, and failures by voice AI agents.

Its underlying Ensemble Listening Model, or ELM, coordinates more than 100 specialized audio models rather than routing every task through one foundation model. Modulate says this design can deliver up to 1,000 times greater inference efficiency while reducing the computing, memory, and energy needed for audio analysis.

Deepfake detection is one of the platform’s most security-relevant capabilities. Modulate reported an average equal error rate of 1.104% across 14 Speech DF Arena datasets as of August 19, 2026, corresponding to 98.9% detection accuracy and first place on that benchmark snapshot. The 316-million-parameter model can begin returning streaming verdicts after 2.5 seconds of speech.

However, Modulate cautions that production systems rarely operate at the equal-error threshold; customers should tune thresholds using their own labeled recordings because stronger detection can produce more false alarms.

Modulate says its models now inspect more than 10 million hours of audio monthly and have processed over 600 million hours overall. The company also reached first place on Hugging Face’s Open ASR Leaderboard for transcription, ranking first among 88 evaluated models in July. Its batch transcription API starts at $0.03 per hour, while deepfake detection is listed at $0.25 per hour.

The capital will help Modulate extend APIs and software development kits, build industry-specific models, add partner integrations, and support deployment environments.

The company is targeting fraud prevention, healthcare security, contact-center oversight, gaming and social-platform moderation, child safety, and AI voice-agent supervision.

Real-time analysis could allow defenders to challenge suspicious callers, escalate risky sessions, or stop harmful conversations before damage occurs rather than reviewing recordings afterward.

For security teams, treat the technology as a detection layer rather than proof of identity. Benchmark leadership does not guarantee equivalent performance on compressed telephone audio, unfamiliar languages, background noise, replay attacks, or new voice generators.

Modulate notes that the arena does not measure streaming behavior or end-to-end narrowband telephony and primarily covers English and Mandarin.

Chief executive Carter Huffman said voice is becoming a primary AI interface, creating problems that transcripts cannot solve. With this round, Modulate is betting that audio understanding will become foundational infrastructure for authenticating callers, monitoring automated agents, and detecting manipulation.

As synthetic voices improve, combining audio AI with multifactor verification, transaction controls, and human review will remain essential to prevent a detection score from becoming a single point of failure.

Cut every SOC alert investigation by 21 min. Power your SOC with instant IOC context for immediate response: Integrate TI Lookup into your SOC

The post Modulate Raises $25M to Fight Deepfake Voices With Next-Generation Audio AI appeared first on Cyber Security News.