#7905 shape '[1, 4995, 1, 32, 2]' is invalid for input of size 262144,Traceback (most recent call last): File

188.192**8 Posted at: 3 hours ago

shape '[1, 4995, 1, 32, 2]' is invalid for input of size 262144,Traceback (most recent call last):
File "videotrans\process\confucius_tts.py", line 75, in confucius_fun
File "C:\vTrans\_internal\torch\utils\_contextlib.py", line 116, in decorate_context

return func(*args, **kwargs)

File "videotrans\confuciustts\cli\inference.py", line 332, in generate
File "C:\vTrans\_internal\torch\utils\_contextlib.py", line 116, in decorate_context

return func(*args, **kwargs)

File "videotrans\confuciustts\cli\inference.py", line 246, in _synth_segment
File "C:\vTrans\_internal\torch\utils\_contextlib.py", line 116, in decorate_context

return func(*args, **kwargs)

File "videotrans\confuciustts\flow\flow.py", line 242, in inference
File "C:\vTrans\_internal\torch\utils\_contextlib.py", line 116, in decorate_context

return func(*args, **kwargs)

File "videotrans\confuciustts\flow\flow_matching.py", line 138, in forward
File "videotrans\confuciustts\flow
......
Error: shape '[1, 4995, 1, 32, 2]' is invalid for input of size 262144

trk=[TransCreate]: uuid='f9958b82fd', proxy_str=None, cfg=[TaskCfgVTT]当前工作模式: 翻译视频 单视频模式
原始输入文件名: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/Kikire, Zna - Reaactor 2.mp4,
输出结果保存到文件夹: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/Kikire, Zna - Reaactor 2-mp4,
临时文件夹: C:/vTrans/tmp/86692/f9958b82fd
未选 清理已存在
已启用CUDA加速
已选 降噪
已选 识别说话人,最大说话人数量不限制
STT Service:openai-whisper(Built-in), large-v3-turbo, 发音语言: English,
TranslateSrt:Local LLM (OpenAI-compatible),原始语言:English,目标语言:German, 未选 SendSRT
TTS Service:Confucius4(Built-in), 角色:clone, 配音语言:German, 未选 二次语音识别
音量:-1%, 语速:+0%, 已选 音频加速, 未选 视频慢速
已选 移除每条字幕配音开头和结尾静音缓冲
静音移除力度: default
字幕: No Subtitles
已选 分离人声与背景声, 已选 重新嵌入背景声, 背景音量0.8, 背景声音时长 短于 视频时长时: 拉长(降速播放),存在分离后的纯净人声文件,存在分离后的背景声音文件, precent=18, hasend=False, should_recogn=True, should_trans=True, should_dubbing=True, should_hebing=True, is_copy_video=True, video_codec_num=264, ignore_align=False, is_audio_trans=False, clone_ref='C:/vTrans/tmp/86692/f9958b82fd/vocal.wav', should_recogn2=False, do_diarize=True
trk.cfg=[TaskCfgVTT]当前工作模式: 翻译视频 单视频模式
原始输入文件名: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/Kikire, Zna - Reaactor 2.mp4,
输出结果保存到文件夹: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/Kikire, Zna - Reaactor 2-mp4,
临时文件夹: C:/vTrans/tmp/86692/f9958b82fd
未选 清理已存在
已启用CUDA加速
已选 降噪
已选 识别说话人,最大说话人数量不限制
STT Service:openai-whisper(Built-in), large-v3-turbo, 发音语言: English,
TranslateSrt:Local LLM (OpenAI-compatible),原始语言:English,目标语言:German, 未选 SendSRT
TTS Service:Confucius4(Built-in), 角色:clone, 配音语言:German, 未选 二次语音识别
音量:-1%, 语速:+0%, 已选 音频加速, 未选 视频慢速
已选 移除每条字幕配音开头和结尾静音缓冲
静音移除力度: default
字幕: No Subtitles
已选 分离人声与背景声, 已选 重新嵌入背景声, 背景音量0.8, 背景声音时长 短于 视频时长时: 拉长(降速播放),存在分离后的纯净人声文件,存在分离后的背景声音文件
video_duration_ms=1123966


=
system:Windows-10-10.0.26300-SP0
version:v4.15
frozen:True
language:en_US
root_dir:C:/vTrans

1 Replies

#1 DeepSeek reply at 3 hours ago

This is a runtime shape/tensor mismatch error inside the Confucius-TTS model's rotary-embedding step(RuntimeError: shape '[1, 4995, 1, 32, 2]' is invalid for input of size 262144),triggered during the clone dubbing stage — i.e. the reference audio clip that was auto-cut from your source video produced an input sequence too long for the model to handle.

This is not a documented FAQ case. The most likely triggers, based on the trace pattern:

  1. Over-long reference audio used for clone voice: The clone role auto-cuts an audio segment from the original video as the reference clip. If the segment is longer than roughly 3–10 s, Confucius-TTS can generate an oversized input that breaks the rotary-embedding reshape. This matches the standard rule that most cloning TTS engines require 3–10 s reference audio.
  2. Very long subtitle line for the dubbing text: The 4995-token dimension suggests one segment carried an extremely long text/audio input. Your source log shows 未选 二次语音识别, No Subtitles, 未选 视频慢速, so dubbing length is not being re-trimmed.

Workarounds to try in order:

  1. In Menu → Tools/Options → Advanced Options → ASR Settings, force a smaller subtitle segment so both reference audio and dubbing text stay short:

    • VAD: max speech duration(s) → 6 (or 5–8)
    • VAD: min speech duration(ms) → 3000–4000
      Then re-run with Clear Output checked.
  2. Avoid clone for this test: switch 配音角色 to a fixed Confucius voice (non-clone). If it succeeds, the problem is confirmed to be the auto-cut reference clip.
  3. Switch the TTS channel temporarily (e.g. to Edge-TTS or OmniVoice-TTS) to verify the rest of the pipeline is fine.
  4. If you want to keep Confucius-TTS with clone, use a manually prepared 3–10 s mono WAV reference (f5-tts/ folder) rather than the auto-cut clip.

If the error persists after step 1 & 2, please provide the latest log file from the software's logs/ folder — copy the bottom ~30 lines — plus whether a non-clone voice works, so we can narrow it down.

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource