AWS adds Qwen3-TTS voice cloning to SageMaker JumpStart
AWS has added Qwen3-TTS voice cloning to SageMaker JumpStart and documented how to deploy the models on managed real-time inference endpoints.
AWS has added Qwen3-TTS voice-cloning models to Amazon SageMaker JumpStart, giving customers a managed route to deploy the publicly available models behind real-time inference endpoints.
The September 25 release includes Qwen3-TTS-12Hz-1.7B-Base, Qwen3-TTS-12Hz-1.7B-CustomVoice and Qwen3-ASR-1.7B. AWS’s walkthrough uses the Base variant to clone a voice from a user-supplied audio sample and transcript. CustomVoice instead relies on predefined speakers.
This differs from the open, modular speech-to-speech pipeline DataPhoenix covered in July. There, Qwen3-TTS was only the spoken-output component alongside separate speech-recognition, language-model and inference systems. The AWS release provides a managed deployment path specifically for Qwen3-TTS voice cloning.
AWS’s workflow starts with a JumpStartModel configured with the model ID huggingface-ttsvoiceclone-qwen3-tts-12hz-1-7b-base. The developer pins version 1.0.1, accepts the model’s end-user license agreement and calls deploy. The example selects an ml.g6.4xlarge instance and sets SM_VLLM_GPU_MEMORY_UTILIZATION to 0.45. JumpStart supplies the model artifacts and a pre-built serving container, so no custom inference handler is required in the example.
The client invokes the endpoint through the SageMaker runtime. Each example request includes the text to synthesize, a base64-encoded 24 kHz mono WAV reference clip, a transcript of that clip, task_type set to Base, and the custom attribute route=/v1/audio/speech. The endpoint returns audio, shown in the walkthrough as a 24 kHz mono WAV file. In this setup, voice cloning conditions generation on the sample and transcript rather than retraining the model.
The Qwen model card says the Base model can clone a voice from three seconds of user audio and supports streaming generation. It lists Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. AWS also demonstrates a cross-lingual request using an English reference clip and transcript to generate Chinese speech. AWS’s claim that the output preserves the reference speaker’s identity was not independently evaluated.
AWS describes two serving stages on the same GPU: a “talker” generates speech tokens, then code2wav converts them into a waveform. The company says one ml.g6.4xlarge instance with a 24 GB Nvidia L4 GPU is sufficient for the 1.7-billion-parameter model when each stage is capped at 45% GPU memory utilization, leaving roughly 10% as a buffer. That is AWS’s sizing guidance, not an independently reproduced benchmark.
The post’s example startup logs show 3.66 GiB for the talker weights, 0.45 GiB for the code2wav weights and 6.08 GiB for the talker’s key-value cache, with a 56,928-token cache budget. AWS does not publish measured end-to-end latency, throughput or concurrency-test results for the endpoint.
AWS says this hosting model ties cost to compute usage instead of per-character API pricing. It also says reference audio and generated speech remain in the customer’s AWS account and customer-managed SageMaker endpoint. The post provides neither a worked cost comparison nor an independent audit of the data-control claim, and it does not list regional availability for the JumpStart entries.
Customers can scale by adding endpoint instances, moving to a larger GPU or configuring SageMaker automatic scaling, according to AWS. The release adds another workload to the same endpoint layer as AWS’s previously documented routing for SageMaker real-time endpoints. SageMaker publishes hardware and invocation metrics to CloudWatch, including GPU and GPU-memory utilization, concurrent requests, model and overhead latency, and error counts.
AWS instructs users to delete the endpoint, endpoint configuration and model after testing to avoid ongoing charges.
More news

xAI launches Team Bots for shared workflows

AWS adds xAI’s Grok 4.7 to Amazon Bedrock

OpenAI adds $5 million and up to $5 million in credits to Lenfest AI program
