#7943 CUDA error: device-side assert triggeredCompile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.

2800**1911 Posted at: 1 hour ago

CUDA error: device-side assert triggered
Compile with TORCH_USE_CUDA_DSA to enable device-side assertions.

Traceback (most recent call last):
File "videotrans\task\only_one.py", line 108, in run
File "videotrans\task\_stage_dubbing.py", line 24, in dubbing
File "videotrans\task\_stage_dubbing.py", line 101, in _tts
File "videotrans\tts\__init__.py", line 122, in run
File "videotrans\tts\_base.py", line 98, in run
File "videotrans\tts\_chatterbox.py", line 97, in _exec
File "videotrans\tts\_chatterbox.py", line 73, in _exec
File "chatterbox\mtl_tts.py", line 301, in generate

self.prepare_conditionals(audio_prompt_path, exaggeration=exaggeration)

File "chatterbox\mtl_tts.py", line 260, in prepare_conditionals

s3gen_ref_dict = self.s3gen.embed_ref(s3gen_ref_wav, S3GEN_SR, device=self.device)

File "chatterbox\models\s3gen\s3gen.py", line 130, in embed_ref

ref_wav = ref_wav.to(device)

RuntimeError: CUDA error: device-side assert triggered
Compile with TORCH_USE_CUDA_DSA to enable device-side assertions.

trk=[TransCreate]: uuid='c779b54edc', proxy_str=None, cfg=[TaskCfgVTT]当前工作模式: 翻译视频 单视频模式
原始输入文件名: C:/Users/Jonathan/Music/_video_out/1_acor34_(Vocals).flac,
输出结果保存到文件夹: C:/Users/Jonathan/Music/_video_out/_video_out/1_acor34_(Vocals)-flac,
临时文件夹: C:/Users/Jonathan/Music/win-pyvideotrans-v4.08-0802/tmp/1600/c779b54edc
未选 清理已存在
已启用CUDA加速
未选 降噪
已选 识别说话人,最大说话人数量不限制
STT Service:faster-whisper(Built-in), large-v3, 发音语言: English,
TranslateSrt:Google Translate (Free),原始语言:English,目标语言:Spanish (Latin America), 已选 SendSRT
TTS Service:ChatterBox(Built-in), 角色:clone, 配音语言:Spanish (Latin America), 未选 二次语音识别
音量:+0%, 语速:+0%, 已选 音频加速, 未选 视频慢速
已选 移除每条字幕配音开头和结尾静音缓冲
静音移除力度: default
字幕: Hard Subtitles , precent=15, hasend=False, should_recogn=True, should_trans=True, should_dubbing=True, should_hebing=False, is_copy_video=False, video_codec_num=264, ignore_align=False, is_audio_trans=True, clone_ref='', should_recogn2=False, do_diarize=True
trk.cfg=[TaskCfgVTT]当前工作模式: 翻译视频 单视频模式
原始输入文件名: C:/Users/Jonathan/Music/_video_out/1_acor34_(Vocals).flac,
输出结果保存到文件夹: C:/Users/Jonathan/Music/_video_out/_video_out/1_acor34_(Vocals)-flac,
临时文件夹: C:/Users/Jonathan/Music/win-pyvideotrans-v4.08-0802/tmp/1600/c779b54edc
未选 清理已存在
已启用CUDA加速
未选 降噪
已选 识别说话人,最大说话人数量不限制
STT Service:faster-whisper(Built-in), large-v3, 发音语言: English,
TranslateSrt:Google Translate (Free),原始语言:English,目标语言:Spanish (Latin America), 已选 SendSRT
TTS Service:ChatterBox(Built-in), 角色:clone, 配音语言:Spanish (Latin America), 未选 二次语音识别
音量:+0%, 语速:+0%, 已选 音频加速, 未选 视频慢速
已选 移除每条字幕配音开头和结尾静音缓冲
静音移除力度: default
字幕: Hard Subtitles
video_duration_ms=5953173


=
system:Windows-10-10.0.26300-SP0
version:v4.15
frozen:True
language:en_US
root_dir:C:/Users/Jonathan/Music/win-pyvideotrans-v4.08-0802

1 Replies

#1 DeepSeek reply at 1 hour ago

The CUDA error: device-side assert triggered is a GPU-side data validity error (in ChatterBox TTS this happens inside s3gen.embed_ref, i.e. the reference audio slice or token index fed to the GPU is out of the model's valid range). With the clone voice role the reference audio is dynamically cut out of your original file per subtitle line, so an invalid/too-short slice triggers it. Once this assert fires, the CUDA context is destroyed — every following GPU call fails too, so a simple retry cannot recover.

Steps to fix (in order):

  1. Fully restart the software. Do not just click retry — the CUDA context must be rebuilt.
  2. Fix reference audio length (main cause with clone):

    • Go to Menu -> Tools/Options -> Advanced Options -> ASR Settings ([语音识别参数])
    • VAD:max speech duration(s): set 6–10
    • VAD:min speech duration(ms): set 3000–4000 (must be ≥3000 for voice cloning — shorter slices cannot be used as reference audio)
    • Note: max duration (s) × 1000 must be greater than min duration (ms), otherwise recognition gets corrupted.
  3. Check VRAM:

    • Close other GPU-consuming programs; ChatterBox on GPU needs roughly 10GB+ free VRAM.
    • In Advanced Options -> ASR Settings, set Compute data type to float16 (GPU-friendly) or int8 (lowest memory).
    • In Advanced Options -> General, keep Number of GPU tasks = 1.
  4. Clean up the file path. Your input is
    C:/Users/Jonathan/Music/_video_out/1_acor34_(Vocals).flac
    Parentheses and other special characters are poorly handled by the underlying ffmpeg slicing. Rename it to plain ASCII letters/digits/underscore, e.g. 1_acor34_vocals.flac, and move it to a shallow path such as D:/vt/. (Also note the odd double nested output folder _video_out/_video_out/ — clear the output folder and let the software generate it.)
  5. Verify your installation is clean. The log reports version:v4.15 but root_dir:C:/Users/Jonathan/Music/win-pyvideotrans-v4.08-0802, i.e. a mixed/partially overwritten installation:

    • Download the full package (about 2.7GB) and overwrite the whole folder, then overwrite the latest patch package again (all files and the _internal and videotrans folders, not only sp.exe).
    • If version-mismatch errors persist, search inside _internal for every __pycache__ folder and delete them before re-copying.
  6. To isolate the fault: run once with Edge-TTS (free) as the dubbing channel. If it succeeds, the pipeline and your media are fine and the problem is specific to the ChatterBox clone reference audio; if it fails too, the problem is in the earlier stages (ASR/audio slicing).

Also confirm ChatterBox extra files are present: features.msgpack and weights.npz extracted into models/pkuser_home; if the model download was incomplete, the same kind of crash can occur.

If none of the above resolves it, please provide the latest log file from the logs/ folder (bottom ~30 lines) of a clean run after restarting, together with whether the step-6 Edge-TTS test succeeded.

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource