The size of tensor a (8926) must match the size of tensor b (8192) at non-singleton dimension 1,Traceback (most recent call last):
File "videotrans\process\f5_tts.py", line 147, in f5tts_fun
File "f5_tts\api.py", line 124, in infer
wav, sr, spec = infer_process(File "f5_tts\infer\utils_infer.py", line 416, in infer_process
return next(File "f5_tts\infer\utils_infer.py", line 543, in infer_batch_process
result = future.result()File "concurrent\futures\_base.py", line 458, in result
File "concurrent\futures\_base.py", line 403, in __get_result
File "concurrent\futures\thread.py", line 58, in run
File "f5_tts\infer\utils_infer.py", line 523, in infer_single_process
generated_wave, generated = _infer_basic(gen_text)File "f5_tts\infer\utils_infer.py", line 497, in _infer_basic
generated, _ = model_obj.sample(File "C:\vTrans\_internal\torch\utils\_contextlib.py", line 116, in decorate_context
return func(*args, **kwargs)Fil
......
es\module.py", line 1762, in _call_impl
return forward_call(*args, **kwargs)File "C:\vTrans\_internal\torchdiffeq\_impl\misc.py", line 197, in forward
return self.base_func(t, y)File "f5_tts\model\cfm.py", line 181, in fn
pred_cfg = self.transformer(File "C:\vTrans\_internal\torch\nn\modules\module.py", line 1751, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)File "C:\vTrans\_internal\torch\nn\modules\module.py", line 1762, in _call_impl
return forward_call(*args, **kwargs)File "f5_tts\model\backbones\dit.py", line 338, in forward
x_cond = self.get_input_embed(File "f5_tts\model\backbones\dit.py", line 299, in get_input_embed
text_embed = self.text_embed(text, seq_len=seq_len, drop_text=drop_text)File "C:\vTrans\_internal\torch\nn\modules\module.py", line 1751, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)File "C:\vTrans\_internal\torch\nn\modules\module.py", line 1762, in _call_impl
return forward_call(*args, **kwargs)File "f5_tts\model\backbones\dit.py", line 120, in forward
text = text + freqsRuntimeError: The size of tensor a (8926) must match the size of tensor b (8192) at non-singleton dimension 1
cfg=[TaskCfgVTT]当前工作模式: 翻译视频 单视频模式
原始输入文件名: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/111808.wmv,
输出结果保存到文件夹: C:/Private Filez/Videoz/Zum Übertragen/Dubbing/Dubprogress/111808-wmv,
临时文件夹: C:/vTrans/tmp/42372/c8fbac555b
未选 清理已存在
已启用CUDA加速
已选 降噪
已选 识别说话人,最大说话人数量不限制
STT Service:openai-whisper(Built-in), large-v3-turbo, 发音语言: English, 断句方式:默认断句
TranslateSrt:CompatibleAI/LocalModel,原始语言:English,目标语言:German, 未选 SendSRT
TTS Service:F5-TTS(Built-in), 角色:clone, 配音语言:German, 未选 二次语音识别
音量:-1%, 语速:+0%, 已选 音频加速, 未选 视频慢速
字幕: No Subtitles
已选 分离人声与背景声, 已选 重新嵌入背景声, 背景音量0.8, 背景声音时长 短于 视频时长时: 拉长(降速播放),存在分离后的纯净人声文件,存在分离后的背景声音文件
=
system:Windows-10-10.0.26200-SP0
version:v4.12
frozen:True
language:en_US
root_dir:C:/vTrans