This format audio_name.wav#audio_text#lang_code is not a valid reference-audio entry format, and combined with your clone requirement it will fail.
The correct reference-audio format requires only 2 parts separated by #: filename.wav#text spoken in the audio. There is no third lang_code field. Adding a third field makes the software treat the whole string as the spoken text, so the reference audio cannot be located correctly and dubbing fails.
Correct steps:
- Prepare a clean, single-speaker WAV audio, 3-10 seconds, mono, no background noise, no leading/trailing silence.
- Copy it into the software directory folder
f5-tts. - Open
Menu -> TTS Settings -> Set reference audio and fill in exactly:
myaudio1.wav#the exact text spoken in this audio - Save, then select
myaudio1.wav in the Voice role dropdown on the main interface.
Additional requirements for reliable cloning (already implied by your clone tag):
- In
Menu -> Tools/Options -> Advanced Options -> ASR Settings, set VAD:max speech duration(s) to 6-10 and VAD:min speech duration(ms) to 3000-4000. Reference clips shorter than 3 seconds are very likely to fail cloning. - Enable
Separate Vocal/BGM in More... to improve reference audio quality. - Use an AI translation channel (DeepSeek/OpenAI) with
Send full SRT format checked.
Note: clone only works inside the Video Translation function, not in the standalone Text-to-Speech panel.
请查阅相关文档: