Owen Song releases Inflect-Micro-v2, a 9.36M-parameter open text-to-speech model on Hugging Face
Independent developer Owen Song released Inflect-Micro-v2, a 9.36-million-parameter open text-to-speech model on Hugging Face that he claims beats larger compact rivals.
Independent developer Owen Song has released Inflect-Micro-v2, an open text-to-speech model with about 9.36 million parameters, on Hugging Face under an Apache-2.0 license. It is small enough to run on a CPU yet, Song claims, competitive with larger compact rivals.
The model generates 24kHz mono audio from one fixed synthetic male voice, handles punctuation-aware long text, and runs on CPU or CUDA. Its deployable weights total 9,356,513 parameters, about 37.5MB in FP32, tiny next to typical speech models.
Song’s model card reports a 66.2% preference rate (21 wins, 10 losses and 3 ties) in an anonymous community listening study against compact baselines including KittenTTS Nano, Piper Low and Supertonic 3, each of which carries more parameters. A companion model, Inflect-Nano-v2, uses about 3.97 million parameters (3,966,721).
Those results are the developer’s own, drawn from a small, self-run listening test of 34 total judgments with no independent replication, so the claim that it beats bigger models should be read as a single-source benchmark rather than an established one.
The appeal is efficiency. A sub-10-million-parameter model that runs without a GPU could suit on-device or low-cost voice applications, if the quality holds up beyond a small preference study.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
