#7904 Audio must be mono, but got 2,Traceback (most recent call last): File

188.192**8 Posted at: 1 hour ago

Audio must be mono, but got 2,Traceback (most recent call last):
File "videotrans\process\higgs_tts.py", line 87, in higgs_fun
File "C:\vTrans\_internal\torch\utils\_contextlib.py", line 116, in decorate_context

return func(*args, **kwargs)

File "C:\vTrans\models\modules\transformers_modules\models_hyphen__hyphen_multimodalart_hyphen__hyphen_higgs_hyphen_audio_hyphen_v3_hyphen_tts_hyphen_4b_hyphen_transformers\1c9dab85523d953d\modeling_higgs_multimodal_qwen3.py", line 342, in generate_speech

codes_TN = self._encode_reference(reference_audio, sr)

File "C:\vTrans\_internal\torch\utils\_contextlib.py", line 116, in decorate_context

return func(*args, **kwargs)

File "C:\vTrans\models\modules\transformers_modules\models_hyphen__hyphen_multimodalart_hyphen__hyphen_higgs_hyphen_audio_hyphen_v3_hyphen_tts_hyphen_4b_hyphen_transformers\1c9dab85523d953d\modeling_higgs_multimodal_qwen3.py", line 261, in _encode_reference

codes_BNT = codec.encode(wav).audio_codes

......
C).mkv,
输出结果保存到文件夹: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/2025-07-25'Dancing with the Stars - Week 4 Opening Dance Number LIVE 10-7-19 (1440p_60fps_VP9-128kbit_AAC)-mkv,
临时文件夹: C:/vTrans/tmp/86692/ee6140645d
未选 清理已存在
已启用CUDA加速
已选 降噪
已选 识别说话人,最大说话人数量不限制
STT Service:openai-whisper(Built-in), large-v3-turbo, 发音语言: English,
TranslateSrt:Local LLM (OpenAI-compatible),原始语言:English,目标语言:German, 未选 SendSRT
TTS Service:Higgs-audio-v3(Built-in), 角色:clone, 配音语言:German, 未选 二次语音识别
音量:-1%, 语速:+0%, 已选 音频加速, 未选 视频慢速
已选 移除每条字幕配音开头和结尾静音缓冲
静音移除力度: default
字幕: No Subtitles
已选 分离人声与背景声, 已选 重新嵌入背景声, 背景音量0.8, 背景声音时长 短于 视频时长时: 拉长(降速播放),存在分离后的纯净人声文件,存在分离后的背景声音文件, precent=12, hasend=False, should_recogn=True, should_trans=True, should_dubbing=True, should_hebing=True, is_copy_video=False, video_codec_num=264, ignore_align=False, is_audio_trans=False, clone_ref='C:/vTrans/tmp/86692/ee6140645d/vocal.wav', should_recogn2=False, do_diarize=True
trk.cfg=[TaskCfgVTT]当前工作模式: 翻译视频 单视频模式
原始输入文件名: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/2025-07-25'Dancing with the Stars - Week 4 Opening Dance Number LIVE 10-7-19 (1440p_60fps_VP9-128kbit_AAC).mkv,
输出结果保存到文件夹: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/2025-07-25'Dancing with the Stars - Week 4 Opening Dance Number LIVE 10-7-19 (1440p_60fps_VP9-128kbit_AAC)-mkv,
临时文件夹: C:/vTrans/tmp/86692/ee6140645d
未选 清理已存在
已启用CUDA加速
已选 降噪
已选 识别说话人,最大说话人数量不限制
STT Service:openai-whisper(Built-in), large-v3-turbo, 发音语言: English,
TranslateSrt:Local LLM (OpenAI-compatible),原始语言:English,目标语言:German, 未选 SendSRT
TTS Service:Higgs-audio-v3(Built-in), 角色:clone, 配音语言:German, 未选 二次语音识别
音量:-1%, 语速:+0%, 已选 音频加速, 未选 视频慢速
已选 移除每条字幕配音开头和结尾静音缓冲
静音移除力度: default
字幕: No Subtitles
已选 分离人声与背景声, 已选 重新嵌入背景声, 背景音量0.8, 背景声音时长 短于 视频时长时: 拉长(降速播放),存在分离后的纯净人声文件,存在分离后的背景声音文件
video_duration_ms=313307


=
system:Windows-10-10.0.26300-SP0
version:v4.15
frozen:True
language:en_US
root_dir:C:/vTrans

1 Replies

#1 DeepSeek reply at 1 hour ago

The core issue is that the Higgs-audio-v3 model's audio codec requires mono audio, but the reference audio file being fed to it is stereo (2 channels).

Your log shows: clone_ref='C:/vTrans/tmp/86692/ee6140645d/vocal.wav' — this reference audio (from the vocal separation step) is stereo, which the Higgs codec rejects.

Try the following workarounds:

  1. Skip vocal/background separation for the clone reference. Uncheck Separate Vocal/BGM(分离人声背景声) and retry, so the clone reference comes from a source that may be encoded differently.
  2. If step 1 does not help, switch the TTS channel to another clone-capable one (e.g. Qwen-TTS(内置) / OmniVoice-TTS / F5-TTS) which handles stereo reference audio correctly, then re-run.
  3. Update to the latest patch package (v4.15), since Higgs-audio-v3 stereo-resistant handling may be improved. Even if you already run the v4.15 full package, re-download and overwrite with the latest patch.

Note: clone mode dynamically cuts reference audio from the original video, so its quality/format is not controllable — this is why codec constraints like mono are frequently hit.

If the error persists after trying the above, please provide the log file from the logs/ folder (last ~30 lines) so we can check whether the reference WAV channels are being forced correctly.

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource