#7146 语音识别阶段出错[Huggingface_ASR(内置)] The expanded size of the tensor (3000) must match the existing size (3027) at non-singleto

139.202**7 Posted at: 20 days ago

语音识别阶段出错[Huggingface_ASR(内置)] The expanded size of the tensor (3000) must match the existing size (3027) at non-singleton dimension 1. Target sizes: [128, 3000]. Tensor sizes: [128, 3027]:Traceback (most recent call last):
File "videotrans\process\stt_pipe.py", line 65, in pipe_asr
File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\automatic_speech_recognition.py", line 244, in call

return super().__call__(inputs, **kwargs)

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\base.py", line 1245, in call

outputs = list(final_iterator)

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\pt_utils.py", line 126, in next

item = next(self.iterator)

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\pt_utils.py", line 271, in next

processed = self.infer(next(self.iterator), **self.params)

File "D:\pyvideotrans
......
outputs = list(final_iterator)
File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\pt_utils.py", line 126, in next

item = next(self.iterator)

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\pt_utils.py", line 271, in next

processed = self.infer(next(self.iterator), **self.params)

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\torch\utils\data\dataloader.py", line 733, in next

data = self._next_data()

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\torch\utils\data\dataloader.py", line 789, in _next_data

data = self._dataset_fetcher.fetch(index)  # may raise StopIteration

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\torch\utils\data\_utils\fetch.py", line 43, in fetch

return self.collate_fn(data)

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\base.py", line 173, in inner

padded[key] = _pad(items, key, _padding_value, padding_side)

File "D:\pyvideotrans-v4.11\win-pyvideotrans-v4.11\_internal\transformers\pipelines\base.py", line 109, in _pad

tensor[i, : len(item[key][0])] = item[key][0]

RuntimeError: The expanded size of the tensor (3000) must match the existing size (3027) at non-singleton dimension 1. Target sizes: [128, 3000]. Tensor sizes: [128, 3027]
[SpeechToText]语音转录: self.out_format='srt',self.copysrt_rawvideo=True,self.spk_insert=False
[TaskCfgSTT]
原始输入文件名: E:/已翻/新建文件夹/midv-888 はじめて彼女ができたのに隣に住む喪女ニートお姉さんに食べられカラダの相性良すぎて浮気なのに何度もsexしまくった 水卜さくら/MIDV-888 .mp4,
输出结果保存到文件夹: D:/pyvideotrans-v4.11/win-pyvideotrans-v4.11/output/recogn,
临时文件夹: D:/pyvideotrans-v4.11/win-pyvideotrans-v4.11/tmp/13696/9fe6a1f275
已启用CUDA加速
未选 降噪
语音识别:Huggingface_ASR(内置), kotoba-tech/kotoba-whisper-v2.0 断句方式:默认断句
发音语言:ja
未选 降噪


=
system:Windows-10-10.0.19045-SP0
version:v4.12
frozen:True
language:zh_CN
root_dir:D:/pyvideotrans-v4.11/win-pyvideotrans-v4.11

3 Replies

#1 DeepSeek reply at 20 days ago

导致该报错的核心原因是 Huggingface_ASR 渠道在批量处理音频特征时,张量形状发生不匹配(目标长度 3000 小于实际特征长度 3027),通常是 kotoba-whisper-v2.0 模型文件下载不完整/损坏,或该模型与当前 transformers 版本存在兼容性问题。

解决方案

  1. 优先切换为 faster-whisper(内置) 渠道(默认推荐)

    • 语音识别渠道改为 faster-whisper(内置),模型选择 large-v3-turbo
    • 对日语支持良好,且经过大量验证,稳定性最高
  2. 若必须使用 Huggingface_ASR 渠道,可采取以下措施

    • 删除损坏的模型文件重新下载:删除 软件目录/models/models--kotoba-tech--kotoba-whisper-v2.0 文件夹,让软件重新自动下载;若自动下载持续失败,请手动前往
      https://huggingface.co/kotoba-tech/kotoba-whisper-v2.0/tree/main 下载全部文件,并放置到上述文件夹内
    • 切换其他模型:在下拉框中改用 reazon-research/japanese-wav2vec2-large-rs35kh(日语专用)或 openai/whisper-large-v3(通用效果好)尝试
    • 暂时关闭 CUDA 加速:取消勾选 启用CUDA加速,排除 GPU 兼容性干扰
  3. 升级软件版本:当前日志中路径显示为 v4.11,但版本号显示 v4.12,若实际未升级到最新版,请前往官方下载完整包覆盖升级,确保 transformers 等核心依赖为最新修复状态。

#2 39.154**6 reply at 2 days ago

尝试删除videotrans目录下的四个.json文件,再运行sp.exe会重建这四个文件,可能有帮助

#3 DeepSeek reply at 2 days ago

暂无法确定该操作是否能解决此报错。

此报错为 Huggingface_ASR 渠道内部张量形状不匹配(模型推理阶段),与 videotrans/ 目录下的 4 个 JSON 配置文件无关:

  • params.json(API 密钥和偏好设置)
  • cfg.json(高级选项及系统设置)
  • codec.json(视频硬解码缓存)
  • ass.json(硬字幕样式缓存)

删除这 4 个文件的作用是恢复软件出厂设置,仅能解决因配置错乱导致的设置类问题,无法修复模型推理层面的张量尺寸错误。

推荐仍按以下方案处理:

  1. 语音识别渠道改为 faster-whisper(内置),模型选 large-v3-turbo(最稳定,日语支持好)
  2. 若坚持使用 Huggingface_ASR,删除 软件目录/models/models--kotoba-tech--kotoba-whisper-v2.0 后重新下载,或改用 reazon-research/japanese-wav2vec2-large-rs35kh 或 openai/whisper-large-v3
  3. 临时取消勾选 启用CUDA加速 测试
  4. 覆盖升级到最新完整包

如仍报错,请提供软件安装目录下 logs/ 文件夹中最新的日志文件(复制底部约30行报错信息)以便进一步分析。

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource