#7853 语音识别阶段出错[Nemotron-3.5-asr-0.6b(内置)] CUDA error: out of memoryCompile with `TORCH_USE_CUDA_DSA` to enable device-side ass

2409**87bb Posted at: 3 hours ago

语音识别阶段出错[Nemotron-3.5-asr-0.6b(内置)] CUDA error: out of memory
Compile with TORCH_USE_CUDA_DSA to enable device-side assertions.
:Traceback (most recent call last):
File "videotrans\process\stt_nemotron.py", line 41, in nemotron_asr
File "E:\sp\win-pyvideotrans-v4.14\_internal\transformers\models\nemotron3_5_asr\generation_nemotron3_5_asr.py", line 35, in generate

return super().generate(inputs=inputs, generation_config=generation_config, **kwargs)

File "E:\sp\win-pyvideotrans-v4.14\_internal\transformers\models\nemotron_asr_streaming\generation_nemotron_asr_streaming.py", line 228, in generate

outputs = super().generate(inputs=inputs, generation_config=generation_config, **kwargs)

File "E:\sp\win-pyvideotrans-v4.14\_internal\transformers\models\parakeet\generation_parakeet.py", line 256, in generate

outputs = super().generate(inputs=inputs, generation_config=generation_config, **kwargs)

File "E:\sp\win-pyvideotrans-v4.14\_internal\torch\utils\_contextl
......
es\module.py", line 1762, in _call_impl

return forward_call(*args, **kwargs)

File "E:\sp\win-pyvideotrans-v4.14\_internal\transformers\models\nemotron_asr_streaming\modeling_nemotron_asr_streaming.py", line 730, in forward

hidden_states = self.act_fn(_mask_subsampled_frames(hidden_states, lengths))

File "E:\sp\win-pyvideotrans-v4.14\_internal\torch\nn\modules\module.py", line 1751, in _wrapped_call_impl

return self._call_impl(*args, **kwargs)

File "E:\sp\win-pyvideotrans-v4.14\_internal\torch\nn\modules\module.py", line 1762, in _call_impl

return forward_call(*args, **kwargs)

File "E:\sp\win-pyvideotrans-v4.14\_internal\torch\nn\modules\activation.py", line 133, in forward

return F.relu(input, inplace=self.inplace)

File "E:\sp\win-pyvideotrans-v4.14\_internal\torch\nn\functional.py", line 1704, in relu

result = torch.relu(input)

RuntimeError: CUDA error: out of memory
Compile with TORCH_USE_CUDA_DSA to enable device-side assertions.

self.cfg=[TaskCfgVTT]当前工作模式: 转录并翻译字幕
原始输入文件名: G:/YTB/[Ep.913] 한공이들 마음 한 조각을 가져갈 앳하트💘 💌찐디의 아이돌 한국어 사랑방 💌-mp4/[Ep.913] 한공이들 마음 한 조각을 가져갈 앳하트 찐디의 아이돌 한국어 사랑방.mp4,
输出结果保存到文件夹: G:/YTB/[Ep.913] 한공이들 마음 한 조각을 가져갈 앳하트💘 💌찐디의 아이돌 한국어 사랑방 💌-mp4/[Ep.913] 한공이들 마음 한 조각을 가져갈 앳하트 찐디의 아이돌 한국어 사랑방-mp4,
临时文件夹: E:/sp/win-pyvideotrans-v4.14/tmp/81112/58d78a9ac8
未选 清理已存在
已启用CUDA加速
未选 降噪
语音识别:Nemotron-3.5-asr-0.6b(内置), , 发音语言: 韩语, 断句方式:默认断句
翻译渠道:Gemini AI,原始语言:韩语,目标语言:简体中文, 未选 发送完整字幕
输出字幕格式: 目标语言在下双字幕
代理地址:http://127.0.0.1:7897


=
system:Windows-10-10.0.26300-SP0
version:v4.14
frozen:True
language:zh_CN
root_dir:E:/sp/win-pyvideotrans-v4.14

1 Replies

#1 DeepSeek reply at 3 hours ago

核心原因是显存不足(OOM)。虽然报错信息带有 CUDA,但错误栈底层是 torch.relu 等标准算子,并非 nvcc 编译问题,不需要按提示去编译任何东西。

解决方案(按优先级尝试):

  1. 降低模型规模:将语音识别渠道从 Nemotron-3.5-asr-0.6b(内置) 换成显存占用更小的模型,推荐 faster-whisper(内置),模型选 large-v3-turbo 或 medium。Nemotron 模型对显存占用较高,且不是首选推荐。
  2. 改用其他识别渠道:如果是韩语视频,推荐:

    • faster-whisper(内置) + large-v3-turbo(兼容性好)
    • 或在 Huggingface_ASR(内置) 中选 nvidia/nemotron-3.5-asr-streaming-0.6b 以外的模型
  3. 降低显存压力(菜单→工具→高级选项→语音识别参数):

    • 把 计算数据类型 改为 int8 或 float16
    • 把 识别准确度 beam_size 调到 1
    • 关闭 启用上下文感知
    • GPU同时任务数 设为 1
  4. 取消 CUDA 加速:如果显存确实太小,可取消勾选 启用CUDA加速,改用 CPU 运行(速度会变慢)。
  5. 检查显存占用:用任务管理器或 nvidia-smi 确认是否有其他程序占用了显存,以及你的显卡显存大小是否足够运行该模型。

请查阅相关文档:

Thinking...
Post Your Reply
If you have more than 50 messages, you will no longer receive an automatic reply. Please ask a new question!
Open source and free maintenance is not easy. If this project is helpful to you, please consider making a small donation to help the project continue to maintain and update.

Dondate via ko-fi.com

Related resource