Modulate Secures $25M to Scale AI Music Detection
Modulate raised $25 million in new funding led by Future Ventures to expand its AI audio detection tools, as streaming services face a surge in

Modulate raised $25 million in new funding led by Future Ventures, bringing its total funding to $60 million. The audio-native AI company, based in Boston, announced the round on Monday, September 28, with participation from existing investors Hyperplane and Lakestar. The capital will support AI research, product development, engineering, and partnerships.
Leadership and company background
Modulate was founded in 2017 by Carter Huffman and Mike Pappas, who met as physics undergraduates at MIT. Pappas, who was CEO when the company's AI music detection API launched in June, has transitioned to Chairman. Huffman, previously the Chief Technology Officer, has now taken over as CEO.
The company says its models analyze more than 10 million hours of audio each month and have processed over 600 million hours in total. It currently employs between 40 and 45 people.
AI music detection API capabilities and market positioning
Modulate markets a self-serve AI music detection API to streaming services and distributors, priced at $0.07 per hour of audio. The tool is designed to flag AI-generated tracks at the point of ingestion before they can dilute royalty payouts for human artists.
The API runs three models across a recording. One model identifies which parts of a clip contain music, speech, or both. Two separate AI detectors then score the vocal and instrumental components. It returns a result for every four-second window and a verdict for the entire clip.
Modulate claims its synthetic voice detection is 98.9% accurate on public benchmark data. It pitches the technology as superior to rival tools, which it says score a track only once and can misfire on human recordings that have been autotuned or compressed. While Modulate has not named a music industry customer, it said at the June launch that it was already seeing interest from major labels and distribution platforms.
Broader AI audio technology and applications
Beyond music, Modulate's core technology aims to help machines understand emotion, tone, intent, and synthetic speech in conversations. Its flagship platform, Velma, is designed to recognize these signals and identify higher-level events like suspected fraud, harassment, or customer dissatisfaction.
Velma is powered by the Ensemble Listening Model (ELM) architecture, which coordinates over 100 specialized audio models. Modulate says ELM is up to 1,000 times more efficient than using a single large model, reducing computing power and energy use. The company claims Velma delivers twice the accuracy of traditional large language models in detecting true positives and generates seven times fewer false positives. It can operate in real time.
This technology has wider applications. It is used in fraud and deepfake detection, AI agent supervision, customer experience, and trust and safety. Modulate's transcription and deepfake detection technologies are ranked number one on public Hugging Face benchmarks. A separate transcription API is priced at $0.03 per hour for batch processing.
Carter Huffman highlighted the unique challenges of audio. "Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript," he said. Steve Jurvetson of Future Ventures added, "Modulate has gained a significant technical lead in audio-native AI, and the market opportunity is expanding quickly."
Modulate plans to add 10 more employees in the coming months.





