#7724 语音识别阶段出错[faster-whisper(内置)] list index out of range:Traceback (most recent call last): File

117.174**7 Posted at: 1 hour ago

语音识别阶段出错[faster-whisper(内置)] list index out of range:Traceback (most recent call last):
File "videotrans\process\stt_faster.py", line 160, in faster_whisper
File "videotrans\process\_stt_utils.py", line 290, in _resegment
IndexError: list index out of range

Traceback (most recent call last):

File "videotrans\task\job.py", line 35, in run

File "videotrans\task\job.py", line 103, in process_task

File "videotrans\task\speech2text.py", line 100, in recogn

File "videotrans\recognition\__init__.py", line 121, in run

File "videotrans\recognition\_base.py", line 99, in run

File "videotrans\recognition\_whisper.py", line 41, in _exec

File "videotrans\recognition\_whisper.py", line 138, in _faster

File "videotrans\configure\base.py", line 264, in _new_process

videotrans.configure.excepts.VideoTransError: list index out of range:Traceback (most recent call last):
File "videotrans\process\stt_faster.py", line 160, in faster_whisper
File "videotrans\process\_stt_utils.py", line 290, in _resegment
IndexError: list index out of range
[SpeechToText]语音转录: self.out_format='srt',self.copysrt_rawvideo=False,self.spk_insert=False
[TaskCfgSTT]
原始输入文件名: E:/app/ReadyOrNotModTools/Fmod/Output/Exports/ReadyOrNot/Content/VO_PC/NarcosCivilianMale/[BARK]Arrested_0.ogg,
输出结果保存到文件夹: F:/app/win-pyvideotrans-v3.98-0410/output/recogn,
临时文件夹: F:/app/win-pyvideotrans/tmp/17100/a7aef4e501
已启用CUDA加速
未选 降噪
语音识别:faster-whisper(内置), tiny 断句方式:默认断句
发音语言:en
未选 降噪
代理地址:http://127.0.0.1:7890


=
system:Windows-10-10.0.28000-SP0
version:v4.14
frozen:True
language:zh
root_dir:F:/app/win-pyvideotrans

1 Replies

#1 DeepSeek reply at 1 hour ago

这是由于 faster-whisper 识别结果过少(音频太短或 VAD 切分后无有效片段),导致内部重断句函数 _resegment 访问空列表越界。

常见于处理极短的音频切片(如 .ogg 音效文件)。

请按以下顺序排查:

  1. 确认音频内容:你处理的是 [BARK]Arrested_0.ogg 这类游戏音效文件,可能只有极短的一声喊叫甚至无有效人声。

    请先确认文件本身是否包含人类说话声。

    若只是音效/喊声,ASR 无有效文本输出,即会触发此越界错误。

  2. 切换识别模型:tiny 模型对短音频、非清晰语音的容错较差。

    请在 菜单 → 工具/选项 → 高级选项 → 语音识别参数 或主界面将模型改为 large-v3 或 large-v3-turbo 重试。

  3. 调整 VAD 参数(菜单 → 工具/选项 → 高级选项 → 语音识别参数):

    • 缩短 最短语音持续(毫秒)(如设为 1000),避免短音频被判定为无效而全部丢弃;
    • 适当降低 语音阈值,提高对短促语音的灵敏度。

  4. 确保 CUDA 环境正常:日志显示"已启用CUDA加速",但 tiny 模型在超短音频下也可能出现空结果。

    可尝试先取消 CUDA 加速,用 CPU 模式验证是否为模型/环境问题。

  5. 更新软件:你当前使用 v3.98,版本较旧。

    该越界保护逻辑在新版中已优化,建议升级到 v4.14(下载完整包 + 补丁包覆盖)后重试。

    下载地址:https://pyvideotrans.com/downpackage

如果升级到最新版并对同一文件重试后仍报此错,请提供软件目录下 logs/ 文件夹中最新日志的底部约 30 行以便进一步分析。

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource