#7720 语音识别阶段出错[faster-whisper(内置)] list index out of range:Traceback (most recent call last): File

117.174**7 Posted at: 1 hour ago

语音识别阶段出错[faster-whisper(内置)] list index out of range:Traceback (most recent call last):
File "videotrans\process\stt_faster.py", line 161, in faster_whisper
File "videotrans\process\_stt_utils.py", line 263, in _resegment
IndexError: list index out of range

Traceback (most recent call last):

File "videotrans\task\job.py", line 35, in run

File "videotrans\task\job.py", line 102, in process_task

File "videotrans\task\speech2text.py", line 100, in recogn

File "videotrans\recognition\__init__.py", line 121, in run

File "videotrans\recognition\_base.py", line 99, in run

File "videotrans\recognition\_whisper.py", line 41, in _exec

File "videotrans\recognition\_whisper.py", line 138, in _faster

File "videotrans\configure\base.py", line 264, in _new_process

videotrans.configure.excepts.VideoTransError: list index out of range:Traceback (most recent call last):
File "videotrans\process\stt_faster.py", line 161, in faster_whisper
File "videotrans\process\_stt_utils.py", line 263, in _resegment
IndexError: list index out of range
[SpeechToText]语音转录: self.out_format='srt',self.copysrt_rawvideo=False,self.spk_insert=True
[TaskCfgSTT]
原始输入文件名: E:/app/ReadyOrNotModTools/Fmod/Output/Exports/ReadyOrNot/Content/VO_PC/NarcosCivilianFemale/[BARK]Pain_3.ogg,
输出结果保存到文件夹: F:/app/win-pyvideotrans-v3.98-0410/output/recogn,
临时文件夹: F:/app/win-pyvideotrans-v3.98-0410/tmp/14524/9c2841dfe9
已启用CUDA加速
未选 降噪
已选 识别说话人,最大说话人数量2
语音识别:faster-whisper(内置), tiny 断句方式:LLM重新断句
发音语言:en
未选 降噪
代理地址:http://127.0.0.1:7890


=
system:Windows-10-10.0.28000-SP0
version:v4.13
frozen:True
language:zh
root_dir:F:/app/win-pyvideotrans-v3.98-0410

1 Replies

#1 DeepSeek reply at 1 hour ago

这是 faster-whisper 识别后处理阶段(_resegment 断句)引发的索引越界,通常是音频片段过短/无有效语音,导致断句时字幕数组为空或只有一条,取相邻字幕时越界;同时也与你启用了「识别说话人 + LLM重新断句」组合有关。

处理步骤:

  1. 关闭 LLM重新断句:将断句方式改为默认(非LLM重新断句),单独测试该 [BARK]Pain_3.ogg 文件能否正常识别。

    LLM重新断句对极短、无实际语义的音频(如游戏音效 Pain_3.ogg)容易触发该越界错误。

  2. 关闭识别说话人:最大说话人数量=2 时,若该音频只有极短一声或几乎无人声,说话人分离结果为空,也会诱发越界。

    先取消勾选「识别说话人」测试。

  3. 检查音频本身:该文件是游戏内的短音效([BARK]Pain_3),可能时长极短或几乎无清晰连续语音,faster-whisper 识别结果为空或仅1条,断句逻辑无相邻项可比较即报错。

    建议先用正常长度的旁白/对话音频验证软件是否正常。

  4. 降低断句粒度冲突:进入 菜单 → 工具/选项 → 高级选项 → 语音识别参数,确认 最长语音持续 与 最短语音持续 关系正常(最长秒数×1000 必须大于最短毫秒数),避免异常。

  5. 切换识别渠道验证:临时改用 openai-whisper(内置) 测试同一文件,若同样报错,基本可确认是该音频内容本身不适用当前断句流程,而非环境问题。

  6. 升级版本:你使用的虽显示 v4.13,但 root_dir 是 v3.98-0410,说明补丁包与目录版本不一致,建议下载完整最新包重新解压覆盖,避免旧代码残留导致的处理异常。

若以上操作后仍报错,请提供:

  • 该 ogg 文件的时长、是否有清晰人声
  • 关闭「LLM重新断句」和「识别说话人」后的完整报错日志

如问题依旧,请提供软件安装目录下 logs/ 文件夹中最新的日志文件(复制底部约30行报错信息)以便分析。

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource