Modulate raises $25 million for audio-native AI expansion
Modulate says it will use a $25 million financing led by Future Ventures to support audio-native model research, developer tools and partnerships.
Modulate said it raised $25 million in new funding led by Future Ventures, with Hyperplane and Lakestar participating. The company plans to use the financing to expand its audio-native models, developer tools and partnerships. Modulate said the investment brings its total funding to $60 million, though it did not disclose the round’s stage, valuation or ownership terms.
Modulate says the capital is earmarked for AI and machine-learning research, product and engineering work, developer relations and partnerships. It also plans to expand the APIs, software development kits, integrations, industry-specific models and deployment options available to developers.
Modulate’s pitch centers on models that inspect the original audio instead of relying only on a transcript. Its Velma platform analyzes signals such as tone, emotion, intent, emphasis and whether speech appears synthetic. It then combines those signals to identify events including possible fraud, harassment or a voice agent failing during a conversation. Modulate says the system can operate in real time.
The company says Velma uses its Ensemble Listening Model architecture, dynamically selecting and combining more than 100 specialized audio models. Modulate claims the design has demonstrated up to 1,000 times greater efficiency than using one large model for the same work. The evidence reviewed for this article did not include an independent reproduction of that efficiency claim.
Modulate also says its models analyze more than 10 million hours of audio each month and have processed more than 600 million hours in total. It describes adoption across fraud prevention, voice-agent supervision, customer experience, and trust and safety. The sources reviewed, however, did not provide customer-level measurements supporting those usage totals.
The company has also claimed first-place positions on Hugging Face benchmarks for transcription and deepfake speech detection, along with 98.9% accuracy for its deepfake detector. The public Open ASR Leaderboard reviewed for this article did not expose the ranking table. The research also did not independently verify the deepfake ranking or accuracy figure. Modulate lists batch transcription at three cents per audio hour.
According to Modulate, the financing will also support hiring in research, engineering and developer relations. The company did not disclose a closing date for the investment.
More news

IBM and Marist launch AI incubator with campus IBM z17

New Jersey fines DataOne $1.07 million over 62 gas generators

Samsung affiliates commit $1 billion to KKR-backed Helix
