#7686 No transcription results returned. Please check the original audio/video [info.duration_after_vad=0.0s].info=Transcripti

5.202**2 Posted at: 1 hour ago

No transcription results returned. Please check the original audio/video [info.duration_after_vad=0.0s].
info=TranscriptionInfo(language='en', language_probability=1, duration=252.3545, duration_after_vad=0.0, all_language_probs=None, transcription_options=TranscriptionOptions(beam_size=5, best_of=5, patience=1, length_penalty=1, repetition_penalty=1.0, no_repeat_ngram_size=0, log_prob_threshold=-1.0, no_speech_threshold=0.6, compression_ratio_threshold=2.4, condition_on_previous_text=False, prompt_reset_on_temperature=0.5, temperatures=[0.0, 0.2, 0.4, 0.6, 0.8, 1.0], initial_prompt=None, prefix=None, suppress_blank=True, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 522, 542, 873, 893, 902, 918, 922, 931, 1350, 1853, 1982, 2460, 2627, 3246, 3253, 3268, 3536, 3846, 3961, 4183, 4667, 6585, 6647, 7273, 9061, 9383, 10428, 10929, 11938, 12033, 12331, 12562, 13793, 14157, 14635, 15265, 15618, 16553, 16604, 18362, 18956, 20
......
hreshold=0.6, compression_ratio_threshold=2.4, condition_on_previous_text=False, prompt_reset_on_temperature=0.5, temperatures=[0.0, 0.2, 0.4, 0.6, 0.8, 1.0], initial_prompt=None, prefix=None, suppress_blank=True, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 522, 542, 873, 893, 902, 918, 922, 931, 1350, 1853, 1982, 2460, 2627, 3246, 3253, 3268, 3536, 3846, 3961, 4183, 4667, 6585, 6647, 7273, 9061, 9383, 10428, 10929, 11938, 12033, 12331, 12562, 13793, 14157, 14635, 15265, 15618, 16553, 16604, 18362, 18956, 20075, 21675, 22520, 26130, 26161, 26435, 28279, 29464, 31650, 32302, 32470, 36865, 42863, 47425, 49870, 50254, 50258, 50358, 50359, 50360, 50361, 50362), without_timestamps=False, max_initial_timestamp=1.0, word_timestamps=True, prepend_punctuations='"\'“¿([{-', append_punctuations='"\'.。,,!!??::”)]}、', multilingual=False, max_new_tokens=None, clip_timestamps=[0.0], hallucination_silence_threshold=None, hotwords=''), vad_options=VadOptions(threshold=0.5, neg_threshold=None, min_speech_duration_ms=0, max_speech_duration_s=inf, min_silence_duration_ms=140, speech_pad_ms=400))

cfg=[TaskCfgVTT]当前工作模式: 翻译视频 单视频模式
原始输入文件名: C:/Users/behnam.hoshyari/Desktop/1404/کانال مدرسه دیالیز/فیلمها/A V FISTULA GRAFT, 9994998060.mp4,
输出结果保存到文件夹: C:/Users/behnam.hoshyari/Desktop/1404/کانال مدرسه دیالیز/دانلودهای اینستا/_video_out/A V FISTULA GRAFT, 9994998060-mp4,
临时文件夹: C:/Users/behnam.hoshyari/Desktop/win-pyvideotrans-v4.11/tmp/87048/fc212235d3
已选 清理已存在
未启用CUDA加速
未选 降噪
STT Service:faster-whisper(Built-in), medium, 发音语言: English, 断句方式:默认断句
TranslateSrt:Microsoft Translate (Free),原始语言:English,目标语言:Persian, 已选 SendSRT
TTS Service:Edge-TTS (Free), 角色:Dilara(Female/IR), 配音语言:Persian, 未选 二次语音识别
音量:+0%, 语速:+0%, 已选 音频加速, 已选 视频慢速
字幕: Hard Subtitles
已选 仅输出mp4


=
system:Windows-10-10.0.19045-SP0
version:v4.11
frozen:True
language:en
root_dir:C:/Users/behnam.hoshyari/Desktop/win-pyvideotrans-v4.11

1 Replies

#1 DeepSeek reply at 1 hour ago

This is caused by VAD (Voice Activity Detection) discarding the entire audio track.

The key value in your log is duration_after_vad=0.0s (out of duration=252.3545), meaning the VAD filter classified 100% of the audio as non-speech (silence/noise) and passed nothing to Whisper.

The language was detected as en with language_probability=1, but no speech segments survived the silence filter.

Why it happens

  1. The audio track is genuinely too quiet, muffled, or dominated by background noise/music, so the VAD probability never crosses the threshold.
  2. Your min_silence_duration_ms=140 (very low) combined with speech_pad_ms=400 and the default speech threshold can cause over-aggressive filtering on certain tracks.

Solutions (try in order)

  1. Verify the file has real speech: Play the source video—if a human is clearly speaking, proceed to step 2.
  2. Go to Menu → Tools/Options → Advanced Options → ASR Settings and lower the VAD speech probability threshold VAD: Speech probability threshold (e.g.

    from 0.5 to 0.3 or lower), then retry.
  3. On the main interface click More Settings... and enable Noise Reduction and/or Separate Vocal/BGM (separate vocals from background).

    This often recovers speech buried in noise/music.
  4. If still failing, switch the recognition channel to openai-whisper(Built-in) (which does not depend on the same VAD pipeline) and retry.
  5. As a last resort, re-extract/re-encode the audio to standard 16kHz mono WAV using an external tool, then feed that audio file directly.

If none of the above works, your source file may genuinely contain no intelligible speech track.

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource