#6609 wwhat tts is best for pre-voice with no cloning, that sound natural. Using english

2a09:**:18 Posted at: 1 day ago

wwhat tts is best for pre-voice with no cloning, that sound natural. Using english

5 Replies

#1 DeepSeek reply at 1 day ago

如果你的目标是“自然、不用克隆音色、纯英文配音”,首选免费渠道是 Edge-TTS(内置免费),这是软件默认推荐,声音自然、支持全部内置语言,开箱即用无需配置。若愿意付费,OpenAI TTSElevenLabs.io 的自然度更高。

具体建议如下:

  1. 免费方案:Edge-TTS

    • 在“配音渠道”选择 Edge-TTS(免费)
    • 在“目标语言”选择英文后,在“配音角色”中选择自然度较高的音色,例如:

      • en-US-AriaNeural(女声,专业自然)
      • en-US-JennyNeural(女声,温暖自然)
      • en-US-GuyNeural(男声,自然沉稳)
      • en-US-EmmaNeural(女声,温暖)
      • en-US-BrianNeural(男声,沉稳)
    • 注意:Edge-TTS 属于免费接口,短时间高频使用可能触发限流,若遇报错可在 高级选项→字幕配音调整 中将 EdgeTTS配音渠道配音并发数 设为 1,配音后暂停秒 设为 5-10 秒。
  2. 付费高端方案(自然度更高)

    • OpenAI TTS:在“配音渠道”选择后,音色选 alloy/nova/shimmer 等,需在 菜单→TTS设置→OpenAI TTS 填写 API Key。
    • ElevenLabs.io:在“配音渠道”选择 ElevenLabs.io,需在对应设置面板填写 API Key,免费版有额度限制。
  3. 内置本地方案(无需联网,但更耗时)

    • Qwen3-TTS(内置)F5-TTS(内置) 也支持英文自然配音,但首次需下载模型,对硬件要求更高,适合离线场景;不需要克隆时直接用其预设音色即可。

你未勾选“克隆”选项(配音角色不选 clone)即可使用普通预设音色,无需任何克隆设置。

请查阅相关文档:

#2 2a09:**:18 reply at 1 day ago

when i use a cloned voice or a voice preset in omnivoice, the voice has different pitch in different parts

#3 DeepSeek reply at 1 day ago

这是由克隆/预设音色在逐句推理时的参考输入不稳定或模型输出波动导致的。若使用的是 clone 克隆模式,Omnivoice 会按每条字幕的起止时间从原始视频中动态截取参考音频片段,这些片段的时长、语音情绪、背景噪声、是否完整成句各不相同,导致克隆出的音色在不同句子上出现音调、音色漂移。

按以下方式调整可显著改善:

  1. 取消动态克隆,改用固定参考音频(若需保持同一音色):准备一段 3-10 秒清晰单人声 WAV,放入软件目录的 f5-tts 文件夹,并在 菜单→TTS设置→设置参考音频 中填写 文件名.wav#说话文本,然后在配音角色里选择该参考音频文件。这样可以保证每句都使用同一稳定音色。
  2. 若必须使用 clone,请强制控制字幕切割时长:进入 菜单→工具→高级选项→语音识别参数,设置:

    • 最长语音持续(秒):6-10
    • 最短语音持续(毫秒):3000-4000
    • 勾选「合并过短字幕到邻近」
      这能避免过短或过长的片段导致克隆参考不稳定。
  3. 开启人声与背景声分离:在主界面「设置更多参数」中勾选「分离人声背景声」,可去除背景噪声对克隆音质的干扰。
  4. 不要勾选「LLM重新断句」:该功能会重新划分时间轴,导致截取的参考音频与说话时间错位,加剧音色不一致。
  5. 若使用预设音色(非 clone)仍出现音调不定:可能是模型对文本韵律的随机波动,建议更换其他预设角色,或调整主界面「设置更多参数」中的“音调+”“配音语速”来手动统一听感;也可改用 Qwen3-TTS、F5-TTS 等其它内置渠道测试。

如果上述调整后问题依旧,请提供软件 logs/ 目录下最新日志文件中的底部约30行报错信息,以便进一步定位。

请查阅相关文档:

#4 2a09:**:18 reply at 1 day ago

i ussed a fixed reference audio but it still have different pitch

#5 DeepSeek reply at 1 day ago

If you used a fixed reference audio and still hear pitch fluctuation, the most likely cause is that the reference audio itself does not meet the required conditions, or the reference text you entered does not accurately match the audio content.

Check the following:

  1. Reference audio must be a clean 3–10 second WAV file with a single clear voice, no background noise, and no extra silence at the beginning or end.
  2. In Menu → TTS Settings → Set Reference Audio, the format must be filename.wav#spoken text. The text after # must be an exact transcript of what is spoken in the audio, otherwise the model will condition on mismatched input and cause voice/pitch instability.
  3. Make sure you actually selected that reference audio file in the 配音角色 dropdown — not a preset role and not clone.
  4. Keep in mind that natural pitch variation across sentences is normal. If the voice timbre remains consistent but pitch rises or falls according to emphasis/emotion, that is not a defect.

If the problem persists after confirming the above, try:

  • Using a different reference audio file (try another 3–10s sample).
  • Switching to a non-clone preset voice (e.g., Edge-TTS, Qwen3-TTS preset roles) for guaranteed consistent pitch.
  • Testing with another cloning-capable TTS channel, such as Qwen3-TTS, F5-TTS, or GPT-SoVITS.

If the issue is severe and you want further diagnosis, please provide the latest 30 lines of log from the logs/ folder, plus the software version number shown in the top-left corner.

Thinking...
Post Your Reply
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource