#7885 语音识别阶段出错[faster-whisper(内置)] No human voice detected. Please confirm that human speech is present in the original file.

104.28**2 Posted at: 1 hour ago

语音识别阶段出错[faster-whisper(内置)] No human voice detected. Please confirm that human speech is present in the original file. [info.duration_after_vad=0.0s].
_kw={'beam_size': 5, 'best_of': 5, 'condition_on_previous_text': False, 'threshold': 0.45, 'no_speech_threshold': 0.6, 'temperature': 0.0, 'repetition_penalty': 1.0, 'compression_ratio_threshold': 2.4, 'initial_prompt': None}
info=TranscriptionInfo(language='en', language_probability=1, duration=1881.3866875, duration_after_vad=0.0, all_language_probs=None, transcription_options=TranscriptionOptions(beam_size=5, best_of=5, patience=1, length_penalty=1, repetition_penalty=1.0, no_repeat_ngram_size=0, log_prob_threshold=-1.0, no_speech_threshold=0.6, compression_ratio_threshold=2.4, condition_on_previous_text=False, prompt_reset_on_temperature=0.5, temperatures=[0.0], initial_prompt=None, prefix=None, suppress_blank=True, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 52
......
, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 522, 542, 873, 893, 902, 918, 922, 931, 1350, 1853, 1982, 2460, 2627, 3246, 3253, 3268, 3536, 3846, 3961, 4183, 4667, 6585, 6647, 7273, 9061, 9383, 10428, 10929, 11938, 12033, 12331, 12562, 13793, 14157, 14635, 15265, 15618, 16553, 16604, 18362, 18956, 20075, 21675, 22520, 26130, 26161, 26435, 28279, 29464, 31650, 32302, 32470, 36865, 42863, 47425, 49870, 50254, 50258, 50359, 50360, 50361, 50362, 50363), without_timestamps=False, max_initial_timestamp=1.0, word_timestamps=True, prepend_punctuations='"\'“¿([{-', append_punctuations='"\'.。,,!!??::”)]}、', multilingual=False, max_new_tokens=None, clip_timestamps=[0.0], hallucination_silence_threshold=None, hotwords=''), vad_options=VadOptions(threshold=0.45, neg_threshold=None, min_speech_duration_ms=0, max_speech_duration_s=inf, min_silence_duration_ms=2000, speech_pad_ms=400))
trk=[TransCreate]: uuid='862945909b', proxy_str='http://127.0.0.1:10808', cfg=[TaskCfgVTT]当前工作模式: 转录并翻译字幕
原始输入文件名: E:/~已上传到2个网盘/★AV/★欧美av/★好看的欧美av/Kelly.Collins.俄罗斯毛妹/video/2022/2022.05.13.sexart.22.05.13.kelly.collins.and.stefany.kyler.waiting.for.you.4k.mp4,
输出结果保存到文件夹: E:/~已上传到2个网盘/★AV/★欧美av/★好看的欧美av/Kelly.Collins.俄罗斯毛妹/video/2022/_video_out/2022.05.13.sexart.22.05.13.kelly.collins.and.stefany.kyler.waiting.for.you.4k-mp4,
临时文件夹: D:/software/2ai/win-pyvideotrans-v4.15/tmp/43448/862945909b
已选 清理已存在
已启用CUDA加速
未选 降噪
语音识别:faster-whisper(内置), large-v3-turbo, 发音语言: 英语,
翻译渠道:微软(免费),原始语言:英语,目标语言:简体中文, 已选 发送完整字幕
输出字幕格式: 单字幕
代理地址: http://127.0.0.1:10808, precent=4, hasend=False, should_recogn=True, should_trans=True, should_dubbing=False, should_hebing=False, is_copy_video=True, video_codec_num=264, ignore_align=False, is_audio_trans=False, clone_ref='', should_recogn2=False, do_diarize=False


=
system:Windows-10-10.0.26200-SP0
version:v4.15
frozen:True
language:zh_CN
root_dir:D:/software/2ai/win-pyvideotrans-v4.15

1 Replies

#1 DeepSeek reply at 1 hour ago

核心原因:VAD(语音活动检测)过滤后有效语音时长为 0(duration_after_vad=0.0s),即系统判定整段音频中没有人类说话声。虽然识别出语言为 en,但这只是模型对音频特征的猜测,并非真实语音。

常见原因及排查步骤:

  1. 确认视频是否真有人声:日志中 is_audio_trans=False、should_dubbing=False,说明这是视频文件。请先用播放器确认该视频确实包含人声说话,而非仅有背景音乐或音效。
  2. 人声被背景音淹没:欧美成人视频常见背景音乐/喘气声大、说话声小。请在主界面点击设置更多参数,勾选 分离人声背景声(无需再勾选降噪),将人声分离出来后再识别。这是此类文件最有效的解决方法。
  3. VAD 阈值过于严格:进入 菜单-工具/选项-高级选项-语音识别参数:

    • 将 语音阈值(VAD: Speech probability threshold) 从默认 0.45 调低至 0.2~0.3(越灵敏但可能误判噪声);
    • 将 非语音阈值(no speech threshold) 调低;
    • 注意调低后可能把背景噪声误当人声,需配合第 2 步分离人声使用。
  4. 音量过低:可先用其他工具(如 ffmpeg)将音频音量提升后再导入。
  5. 文件名/路径含特殊符号(仅供参考):你的路径中包含 ~、★ 等特殊符号,虽不直接导致 VAD 判空,但建议改名为纯英文数字,避免后续合成阶段出问题。

如按上述操作后仍报同样错误,请提供:

  • 该视频播放时是否确有清晰人声;
  • 已尝试的 VAD 参数调整值;
  • 调整后 logs/ 目录下最新日志底部的报错信息。

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource