#7087 Lỗi trong giai đoạn nhận diện giọng nói[faster-whisper(Tích hợp sẵn)] No transcription results returned. Please check th

116.99**8 Posted at: 2 hours ago

Lỗi trong giai đoạn nhận diện giọng nói[faster-whisper(Tích hợp sẵn)] No transcription results returned. Please check the original audio/video [info.duration_after_vad=0.0s].
info=TranscriptionInfo(language='en', language_probability=0.517136812210083, duration=688.69, duration_after_vad=0.0, all_language_probs=[('en', 0.517136812210083), ('ja', 0.07975998520851135), ('ru', 0.06972885876893997), ('zh', 0.04287282004952431), ('es', 0.03771509602665901), ('ko', 0.03457525372505188), ('pl', 0.03138678893446922), ('de', 0.030449751764535904), ('pt', 0.021567143499851227), ('fr', 0.021338719874620438), ('it', 0.019429318606853485), ('tr', 0.01583785191178322), ('id', 0.014627217315137386), ('nl', 0.00767894322052598), ('cs', 0.006308017298579216), ('vi', 0.005704723764210939), ('hu', 0.004874781239777803), ('fi', 0.004512672312557697), ('sv', 0.0043468778021633625), ('el', 0.0036927834153175354), ('ar', 0.0036316371988505125), ('ro', 0.0034587192349135876), ('uk', 0.00263022119179368), ('m
......
1942225890464e-10), ('yue', 6.71338873381444e-10), ('ba', 6.708140154465525e-10)], transcription_options=TranscriptionOptions(beam_size=5, best_of=5, patience=1, length_penalty=1, repetition_penalty=1.0, no_repeat_ngram_size=0, log_prob_threshold=-1.0, no_speech_threshold=0.6, compression_ratio_threshold=2.4, condition_on_previous_text=False, prompt_reset_on_temperature=0.5, temperatures=[0.0, 0.2, 0.4, 0.6, 0.8, 1.0], initial_prompt=None, prefix=None, suppress_blank=True, suppress_tokens=(1, 2, 7, 8, 9, 10, 14, 25, 26, 27, 28, 29, 31, 58, 59, 60, 61, 62, 63, 90, 91, 92, 93, 359, 503, 522, 542, 873, 893, 902, 918, 922, 931, 1350, 1853, 1982, 2460, 2627, 3246, 3253, 3268, 3536, 3846, 3961, 4183, 4667, 6585, 6647, 7273, 9061, 9383, 10428, 10929, 11938, 12033, 12331, 12562, 13793, 14157, 14635, 15265, 15618, 16553, 16604, 18362, 18956, 20075, 21675, 22520, 26130, 26161, 26435, 28279, 29464, 31650, 32302, 32470, 36865, 42863, 47425, 49870, 50254, 50258, 50359, 50360, 50361, 50362, 50363), without_timestamps=False, max_initial_timestamp=1.0, word_timestamps=True, prepend_punctuations='"\'“¿([{-', append_punctuations='"\'.。,,!!??::”)]}、', multilingual=False, max_new_tokens=None, clip_timestamps=[0.0], hallucination_silence_threshold=None, hotwords=''), vad_options=VadOptions(threshold=0.5, neg_threshold=None, min_speech_duration_ms=0, max_speech_duration_s=inf, min_silence_duration_ms=140, speech_pad_ms=400))
[SpeechToText]语音转录: self.out_format='srt',self.copysrt_rawvideo=False,self.spk_insert=True
[TaskCfgSTT]
原始输入文件名: D:/TAIVE/nhnjun1lang/02 - Build a survival campsite next to a fantastic little stream.#bushcraf..._NoWatermark.faudio.mp3,
输出结果保存到文件夹: D:/win-pyvideotrans-v4.11/output/recogn,
临时文件夹: D:/win-pyvideotrans-v4.11/tmp/13856/31be19ea30
未启用CUDA加速
已选 降噪
Dịch vụ STT:faster-whisper(Tích hợp sẵn), large-v3-turbo 断句方式:默认断句
发音语言:自动检测
已选 降噪


=
system:Windows-10-10.0.19045-SP0
version:v4.11
frozen:True
language:vi
root_dir:D:/win-pyvideotrans-v4.11

1 Replies

#1 DeepSeek reply at 2 hours ago

Lỗi này do VAD (Voice Activity Detection) không phát hiện được bất kỳ phân đoạn giọng nói nào trong âm thanh. Cụ thể, log cho thấy:

  • Tổng thời lượng âm thanh: 688.69 giây
  • Thời lượng sau VAD: 0.0 giây → toàn bộ âm thanh bị VAD coi là không có giọng nói

Đồng thời, bạn đang bật "降噪" (khử nhiễu)chọn "自动检测" (tự động phát hiện ngôn ngữ), đây là hai yếu tố dễ gây ra tình trạng này.

Cách khắc phục:

  1. Bỏ chọn "降噪"
    Khử nhiễu có thể loại bỏ cả giọng nói, khiến VAD không còn nhận ra bất kỳ tín hiệu nào. Hãy tắt nó và thử lại.
  2. Điều chỉnh tham số VAD trong 菜单 → 工具 → 高级选项 → 语音识别参数:

    • Giảm "语音阈值" (Speech probability threshold) xuống khoảng 0.3 (mặc định 0.5, càng nhỏ càng nhạy, nhưng dễ nhận nhầm nhiễu)
    • Giảm "静音分割持续毫秒" (Min silence duration for split) xuống khoảng 100 (mặc định 140, nếu âm thanh có nhiều khoảng lặng ngắn, VAD có thể bỏ qua)
    • Giảm "非语音阈值" (no speech threshold) xuống 0.5 hoặc thấp hơn
    • Tăng "语音阈值" nếu âm thanh có nền ồn lớn, nhưng lưu ý không nên đặt quá cao
  3. Kiểm tra file âm thanh
    Hãy mở file 02 - Build a survival campsite...NoWatermark.faudio.mp3 bằng trình phát âm thanh, xác nhận rằng có giọng nói rõ ràng của con người. Nếu file này chỉ có nhạc nền hoặc tiếng ồn, nó sẽ không thể nhận diện.
  4. Thử model khác
    Thay đổi model từ large-v3-turbo sang large-v3 (độ chính xác cao hơn) hoặc medium để xem có kết quả không.
  5. Nếu vẫn không được
    Hãy kiểm tra file âm thanh gốc và gửi log mới nhất (trong thư mục logs/) để phân tích thêm.

Xem tài liệu liên quan:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource