Local voice conditioning
Record or import a short, clean reference clip and keep the voice profile inside the local workspace.
Offline Windows studio · v1.0.0
XTTS Portable turns text and a consenting reference voice into WAV audio with a complete local runtime—no cloud synthesis, no account and no voice upload.
5.21 GiB Windows x64 ZIP64 · Non-commercial use only · Persian speech is experimental.
A complete local studio
The portable package brings the interface, Python runtime, audio stack and XTTS-v2 model together in one Windows folder.
Record or import a short, clean reference clip and keep the voice profile inside the local workspace.
Type text or import common text, subtitle, markup and data formats, then create a standard WAV output.
Use NVIDIA CUDA when available and fall back to CPU when the selected GPU path cannot run.
Start in German and switch to the Persian UI without rearranging the core workspace.
Sensitive by nature
A voice sample is biometric and personal data. XTTS Portable is built for on-device processing and does not require uploading your recording, text or generated audio to a synthesis service.
System requirements
XTTS Portable trades a large download and heavier local processing for privacy and independence from an online service.
The package contains its own application runtime. Extract it to a normal writable folder.
Keep at least 12 GB free for the extracted package, temporary work and generated audio.
A CUDA-capable GPU improves speed. CPU mode works but model loading and synthesis can be much slower.
A microphone or a 6–20 second clean voice file gives the model a better reference.
The local model is large. A cold start can take well over a minute on modest hardware.
Once the complete package is extracted, everyday generation can remain offline.
Use responsibly
The technical ability to imitate a voice does not grant permission to use somebody's identity or to mislead an audience.
XTTS Portable · v1.0.0
Download the complete Windows package for private, non-commercial voice work.