语音识别阶段出错[faster-whisper(内置)] No human voice detected. Please confirm that human speech is present in the original file. [info.duration_after_vad=0.0s].
_kw={'beam_size': 5, 'best_of': 5, 'condition_on_previous_text': False, 'threshold': 0.45, 'no_speech_threshold': 0.6, 'temperature': 0.0, 'repetition_penalty': 1.0, 'compression_ratio_threshold': 2.4, 'initial_prompt': None}
info=TranscriptionInfo(language='en', language_probability=0.513671875, duration=1.9971875, duration_after_vad=0.0, all_language_probs=[('en', 0.513671875), ('ja', 0.081298828125), ('ru', 0.070068359375), ('zh', 0.042510986328125), ('es', 0.0379638671875), ('ko', 0.034820556640625), ('pl', 0.031707763671875), ('de', 0.0305023193359375), ('pt', 0.021636962890625), ('fr', 0.0214691162109375), ('it', 0.0196990966796875), ('tr', 0.0157623291015625), ('id', 0.0146942138671875), ('nl', 0.00774383544921875), ('cs', 0.006381988525390625), ('vi', 0.00569915771484375), ('hu', 0.00492095947265625), ('fi', 0.0045242309570312
......
0.0), ('tt', 0.0), ('haw', 0.0), ('ln', 0.0), ('ha', 0.0), ('ba', 0.0), ('jw', 0.0), ('su', 0.0), ('yue', 0.0)], transcription_options=TranscriptionOptions(beam_size=5, best_of=5, patience=1, length_penalty=1, repetition_penalty=1.0, no_repeat_ngram_size=0, log_prob_threshold=-1.0, no_speech_threshold=0.6, compression_ratio_threshold=2.4, condition_on_previous_text=False, prompt_reset_on_temperature=0.5, temperatures=[0.0], initial_prompt=None, prefix=None, suppress_blank=True, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 522, 542, 873, 893, 902, 918, 922, 931, 1350, 1853, 1982, 2460, 2627, 3246, 3253, 3268, 3536, 3846, 3961, 4183, 4667, 6585, 6647, 7273, 9061, 9383, 10428, 10929, 11938, 12033, 12331, 12562, 13793, 14157, 14635, 15265, 15618, 16553, 16604, 18362, 18956, 20075, 21675, 22520, 26130, 26161, 26435, 28279, 29464, 31650, 32302, 32470, 36865, 42863, 47425, 49870, 50254, 50258, 50359, 50360, 50361, 50362, 50363), without_timestamps=False, max_initial_timestamp=1.0, word_timestamps=True, prepend_punctuations='"\'“¿([{-', append_punctuations='"\'.。,,!!??::”)]}、', multilingual=False, max_new_tokens=None, clip_timestamps=[0.0], hallucination_silence_threshold=None, hotwords=''), vad_options=VadOptions(threshold=0.45, neg_threshold=None, min_speech_duration_ms=0, max_speech_duration_s=inf, min_silence_duration_ms=2000, speech_pad_ms=400))
[SpeechToText]语音转录: self.out_format='srt',self.copysrt_rawvideo=False,self.spk_insert=True
[TaskCfgSTT]
原始输入文件名: E:/app/ReadyOrNotModTools/Fmod/Output/Exports/ReadyOrNot/Content/VO_PC/NarcosCivilianFemale/[BARK]Arrested_0.ogg,
输出结果保存到文件夹: F:/app/win-pyvideotrans-v4.14/output/recogn,
临时文件夹: F:/app/win-pyvideotrans-v4.14/tmp/20144/e3aebe5d43
已启用CUDA加速
未选 降噪
语音识别:faster-whisper(内置), large-v3-turbo 断句方式:LLM重新断句
发音语言:自动检测
未选 降噪
=
system:Windows-10-10.0.28000-SP0
version:v4.14
frozen:True
language:zh_CN
root_dir:F:/app/win-pyvideotrans-v4.14