语音识别阶段出错[faster-whisper(内置)] No human voice detected. Please confirm that human speech is present in the original file. [info.duration_after_vad=0.0s].
_kw={'beam_size': 5, 'best_of': 5, 'condition_on_previous_text': True, 'threshold': 0.5, 'no_speech_threshold': 0.5, 'temperature': 0.0, 'repetition_penalty': 1.0, 'compression_ratio_threshold': 2.2, 'initial_prompt': '# ROLE\nYou are a forensic audio transcription specialist for the Los Suenos Police Department. Your task is to accurately transcribe the provided audio into English SRT subtitles, capturing every spoken word, hesitation, and emotional nuance with maximum fidelity.\n\n# GAME UNIVERSE CONTEXT\nThis audio originates from the SWAT tactical shooter "Ready or Not". The setting is the fictional American city of Los Suenos. The audio may contain:\n- Tactical Radio Communications: SWAT officers, TOC (Tactical Operations Center), and law enforcement using police radio brevity codes.\n- Civilian Interactions: D
......
d civilian crying and begging]\nCorrect Output:*\n1\n00:00:10,000 --> 00:00:13,500\nCivilian: please... oh god please don\'t... I have a family...\n\n2\n00:00:14,000 --> 00:00:16,000\nCivilian: I\'ll do anything you want, just don\'t shoot!\n\n# ACTUAL TASK\nTranscribe the following audio into English SRT format. \nFollow the game context, proper noun recognition, and emotional cue guidelines above.\nOutput the result as a valid SRT file.', prefix=None, suppress_blank=True, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 522, 542, 873, 893, 902, 918, 922, 931, 1350, 1853, 1982, 2460, 2627, 3246, 3253, 3268, 3536, 3846, 3961, 4183, 4667, 6585, 6647, 7273, 9061, 9383, 10428, 10929, 11938, 12033, 12331, 12562, 13793, 14157, 14635, 15265, 15618, 16553, 16604, 18362, 18956, 20075, 21675, 22520, 26130, 26161, 26435, 28279, 29464, 31650, 32302, 32470, 36865, 42863, 47425, 49870, 50254, 50258, 50358, 50359, 50360, 50361, 50362), without_timestamps=False, max_initial_timestamp=1.0, word_timestamps=True, prepend_punctuations='"\'“¿([{-', append_punctuations='"\'.。,,!!??::”)]}、', multilingual=False, max_new_tokens=None, clip_timestamps=[0.0], hallucination_silence_threshold=None, hotwords=''), vad_options=VadOptions(threshold=0.5, neg_threshold=None, min_speech_duration_ms=0, max_speech_duration_s=inf, min_silence_duration_ms=2000, speech_pad_ms=400))
[SpeechToText]语音转录: self.out_format='srt',self.copysrt_rawvideo=False,self.spk_insert=False
[TaskCfgSTT]
原始输入文件名: E:/app/ReadyOrNotModTools/Fmod/Output/Exports/ReadyOrNot/Content/VO_PC/NarcosCivilianMale/[BARK]Bashed_1.ogg,
输出结果保存到文件夹: F:/app/win-pyvideotrans-v3.98-0410/output/recogn,
临时文件夹: F:/app/win-pyvideotrans/tmp/17100/b32d7e32df
已启用CUDA加速
未选 降噪
语音识别:faster-whisper(内置), tiny 断句方式:默认断句
发音语言:en
未选 降噪
代理地址:http://127.0.0.1:7890
=
system:Windows-10-10.0.28000-SP0
version:v4.14
frozen:True
language:zh
root_dir:F:/app/win-pyvideotrans