#7725 语音识别阶段出错[faster-whisper(内置)] No human voice detected. Please confirm that human speech is present in the original file.

117.174**7 Posted at: 1 hour ago

语音识别阶段出错[faster-whisper(内置)] No human voice detected. Please confirm that human speech is present in the original file. [info.duration_after_vad=0.0s].
_kw={'beam_size': 5, 'best_of': 5, 'condition_on_previous_text': True, 'threshold': 0.5, 'no_speech_threshold': 0.5, 'temperature': 0.0, 'repetition_penalty': 1.0, 'compression_ratio_threshold': 2.2, 'initial_prompt': '# ROLE\nYou are a forensic audio transcription specialist for the Los Suenos Police Department. Your task is to accurately transcribe the provided audio into English SRT subtitles, capturing every spoken word, hesitation, and emotional nuance with maximum fidelity.\n\n# GAME UNIVERSE CONTEXT\nThis audio originates from the SWAT tactical shooter "Ready or Not". The setting is the fictional American city of Los Suenos. The audio may contain:\n- Tactical Radio Communications: SWAT officers, TOC (Tactical Operations Center), and law enforcement using police radio brevity codes.\n- Civilian Interactions: D
......
d civilian crying and begging]\nCorrect Output:*\n1\n00:00:10,000 --> 00:00:13,500\nCivilian: please... oh god please don\'t... I have a family...\n\n2\n00:00:14,000 --> 00:00:16,000\nCivilian: I\'ll do anything you want, just don\'t shoot!\n\n# ACTUAL TASK\nTranscribe the following audio into English SRT format. \nFollow the game context, proper noun recognition, and emotional cue guidelines above.\nOutput the result as a valid SRT file.', prefix=None, suppress_blank=True, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 522, 542, 873, 893, 902, 918, 922, 931, 1350, 1853, 1982, 2460, 2627, 3246, 3253, 3268, 3536, 3846, 3961, 4183, 4667, 6585, 6647, 7273, 9061, 9383, 10428, 10929, 11938, 12033, 12331, 12562, 13793, 14157, 14635, 15265, 15618, 16553, 16604, 18362, 18956, 20075, 21675, 22520, 26130, 26161, 26435, 28279, 29464, 31650, 32302, 32470, 36865, 42863, 47425, 49870, 50254, 50258, 50358, 50359, 50360, 50361, 50362), without_timestamps=False, max_initial_timestamp=1.0, word_timestamps=True, prepend_punctuations='"\'“¿([{-', append_punctuations='"\'.。,,!!??::”)]}、', multilingual=False, max_new_tokens=None, clip_timestamps=[0.0], hallucination_silence_threshold=None, hotwords=''), vad_options=VadOptions(threshold=0.5, neg_threshold=None, min_speech_duration_ms=0, max_speech_duration_s=inf, min_silence_duration_ms=2000, speech_pad_ms=400))
[SpeechToText]语音转录: self.out_format='srt',self.copysrt_rawvideo=False,self.spk_insert=False
[TaskCfgSTT]
原始输入文件名: E:/app/ReadyOrNotModTools/Fmod/Output/Exports/ReadyOrNot/Content/VO_PC/NarcosCivilianMale/[BARK]Bashed_1.ogg,
输出结果保存到文件夹: F:/app/win-pyvideotrans-v3.98-0410/output/recogn,
临时文件夹: F:/app/win-pyvideotrans/tmp/17100/b32d7e32df
已启用CUDA加速
未选 降噪
语音识别:faster-whisper(内置), tiny 断句方式:默认断句
发音语言:en
未选 降噪
代理地址:http://127.0.0.1:7890


=
system:Windows-10-10.0.28000-SP0
version:v4.14
frozen:True
language:zh
root_dir:F:/app/win-pyvideotrans

1 Replies

#1 DeepSeek reply at 1 hour ago

核心原因是:VAD(语音活动检测)在整段音频中未检测到有效人声,duration_after_vad=0.0s 说明音频在送入识别前就被判定为无人类语音。

从日志看,音频文件是游戏内提取的短音效 [BARK]Bashed_1.ogg,时长仅约 0.97 秒,且极可能是短促音效/环境声/呼吸声,而非连续人声,因此被 VAD 过滤掉了。

可按以下顺序排查/调整:

  1. 确认音频本身是否含清晰人声
    用播放器单独播放该 .ogg 文件,确认里面是否真的有说话声。若是纯音效、喊叫、喘息、环境音,ASR 大概率无法识别。
  2. 降低 VAD 判定门槛(最可能有效)
    进入 菜单 → 工具/选项 → 高级选项 → 语音识别参数:

    • 语音阈值(VAD: Speech probability threshold) 从默认 0.5 调低到 0.2~0.3
    • 非语音阈值(no speech threshold) 调低
    • 静音分割持续(VAD:min silence duration for split(ms)) 调小(如 300~500)
      目的是让 VAD 更容易把短促人声判定为语音。
  3. 关闭降噪/人声分离的影响
    当前日志显示“未选 降噪”。若原音频噪声大,建议先勾选 分离人声背景声 再重试;反之若音频本身很干净,则不必开。
  4. 确认发音语言
    日志中发音语言为 en,若该音效其实不是英语人声,请改为正确语言,或改用 语音转录 功能并选 自动检测。
  5. 换更宽松的识别渠道测试
    可临时改用 openai-whisper(内置) 或 Qwen-ASR(内置) 测试,看是否能出结果,以判断是模型/VAD问题还是音频本身问题。
  6. 注意版本与模型
    日志显示 version:v4.14,但安装目录名为 v3.98-0410,且模型用的是 tiny。tiny 模型对短音频、噪声、音效类音频非常容易漏识别,建议换成 large-v3-turbo 或 large-v3 再试。

如果以上调整后仍报同样错误,说明该音频中确实不存在可被识别的连续人声,需要更换音频源或先做人工切分。

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource