SageAttention 是一种注意力加速算法,用近似计算替代标准注意力以节省显存,ComfyUI 以内置参数 --use-sage-attention 支持它。本地跑 MiniMax H3 视频生成时 5 秒的片子要等好几分钟,这篇记的是我在 Windows + RTX 3060 12G 上安装它的流程。
我的环境是 RTX 3060 12G + 32G 内存,Windows 系统,ComfyUI 便携版(ComfyUI_windows_portable)。
关键前提:这是内置能力,不是插件
ComfyUI 官方在 comfy/cli_args.py 里定义了一个互斥的 attention 选项组,下面这几个参数是平级的:
--use-pytorch-cross-attention
--use-sage-attention
--use-flash-attention
--use-ck-attention
也就是说 --use-sage-attention 是官方内置参数,装好 SageAttention 这个 Python 包之后,ComfyUI 启动时加这个参数就能用,不需要装任何自定义节点。
顺带说明:社区还有一个叫「Patch Sage Attention KJ」的自定义节点,那是另一条路线。它在不少环境里会卡在 triton.experimental 的导入错误上,所以我没走那条,直接用内置参数。
安装步骤
整个过程分六步,难点不在命令,而在于每一步的版本必须互相对得上。下面是完整的流程。
1. 下载并解压 ComfyUI 便携版
建议装在干净的目录里,比如 C:\comfyui\sage_clean_setup,避免和已有环境互相污染。
curl.exe -L "https://github.com/Comfy-Org/ComfyUI/releases/latest/download/ComfyUI_windows_portable_nvidia.7z" -o ".\ComfyUI_windows_portable_nvidia.7z"
tar.exe -xf ".\ComfyUI_windows_portable_nvidia.7z"
解压后先确认运行环境版本:
cd "C:\comfyui\sage_clean_setup\ComfyUI_windows_portable"
.\python_embeded\python.exe -c "import sys, torch; print('Python:', sys.version); print('Torch:', torch.__version__); print('CUDA runtime:', torch.version.cuda); print('CUDA available:', torch.cuda.is_available()); print('GPU:', torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'None'); print('Capability:', torch.cuda.get_device_capability(0) if torch.cuda.is_available() else 'None')"
这一步的输出要记下来,后面选 wheel 全靠它。我的环境是 Python 3.13、Torch 2.13.0+cu130、CUDA runtime 13.0、GPU capability (8, 9)。
2. 补齐 Python 的 include 与 libs
这是最容易卡住的一步。SageAttention 需要编译头文件和库文件,而官方便携版默认不带。
先检查:
Test-Path ".\python_embeded\include\Python.h"
Test-Path ".\python_embeded\libs\python313.lib"
两个都是 False 就是缺。下载对应 Python 版本的补齐包并解压进 python_embeded:
curl.exe -L "https://github.com/woct0rdho/triton-windows/releases/download/v3.0.0-windows.post1/python_3.13.2_include_libs.zip" -o ".\python_3.13.2_include_libs.zip"
tar.exe -xf ".\python_3.13.2_include_libs.zip" -C ".\python_embeded"
再跑一次 Test-Path,应该都变成 True。
注意:包名里的
python_3.13.2要和你的python_embeded实际版本匹配。便携版更新很勤,Python 小版本可能已经变了。
3. 安装 triton-windows
.\python_embeded\python.exe -m pip install --no-deps triton-windows==3.7.1.post27
验证:
.\python_embeded\python.exe -c "import triton; import triton.experimental; print('triton:', triton.__version__); print('experimental: OK')"
看到 triton: 3.7.1 和 experimental: OK 就成了。--no-deps 不能省,否则 pip 会试图拉一堆不匹配的依赖。
4. 安装 SageAttention 的 Windows wheel
先按你的 Torch/CUDA 版本去 Releases 找对应文件。文件名里其实已经把匹配条件写全了:
sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64.whl
└─cu13 ─┘└─ torch 2.10+ ─┘└─ post6 ─┘└ abi3 跨版本 ┘
abi3 表示这个 wheel 不锁 Python 小版本,这是它比 triton-windows 那类包省事的地方。
curl.exe -L "https://github.com/woct0rdho/SageAttention/releases/download/v2.2.0-windows.post6/sageattention-2.2.0%2Bcu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64.whl" -o ".\sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64.whl"
.\python_embeded\python.exe -m pip install --no-deps ".\sageattention-2.2.0+cu130torch2.10.0andhigher.post6-cp310-abi3-win_amd64.whl"
验证:
.\python_embeded\python.exe -c "import sageattention; print('sageattention import: OK'); print('sageattention path:', sageattention.__file__)"
5. 跑一次最小内核测试
在启动 ComfyUI 之前,先确认 SageAttention 本身能算:
.\python_embeded\python.exe -c "import torch; from sageattention import sageattn; q=torch.randn(1,1024,56,128,device='cuda',dtype=torch.bfloat16); k=torch.randn_like(q); v=torch.randn_like(q); y=sageattn(q,k,v,tensor_layout='NHD'); torch.cuda.synchronize(); print('OK', y.shape, y.dtype, torch.isfinite(y).all().item())"
输出 OK torch.Size([1, 1024, 56, 128]) torch.bfloat16 True 就算通过。这一步能在几分钟内排掉绝大多数环境问题,比等 ComfyUI 跑起来再排查省事得多。
6. 带 SageAttention 启动 ComfyUI
.\python_embeded\python.exe -s .\ComfyUI\main.py --windows-standalone-build --use-sage-attention
控制台末尾应该出现:
[INFO] Using sage attention
[INFO] Starting server
[INFO] To see the GUI go to: http://127.0.0.1:8188
看到第一行就说明 SageAttention 真的挂上了。不想用加速时,把 --use-sage-attention 去掉即可,对应的日志会变成 Using pytorch attention——这两个词可以用来确认当前跑在哪种 attention 上。
复用已有模型:extra_model_paths.yaml
如果你是重装 ComfyUI 而模型还在老目录,没必要把几十 GB 的模型复制过来(SSD 空间不够)。用配置文件让新环境指向老模型目录:
Copy-Item ".\ComfyUI\extra_model_paths.yaml.example" ".\ComfyUI\extra_model_paths.yaml"
然后编辑 extra_model_paths.yaml,把 base_path 改成你老 ComfyUI 的路径:
comfyui:
base_path: C:/comfyui/ComfyUI_windows_portable/ComfyUI/
checkpoints: models/checkpoints/
configs: models/configs/
loras: models/loras/
vae: models/vae/
text_encoders: |
models/text_encoders/
models/clip/
diffusion_models: |
models/unet/
models/diffusion_models/
clip_vision: models/clip_vision/
controlnet: |
models/controlnet/
models/t2i_adapter/
| 竖线的含义是「下面这些子目录都算」。写完直接启动 ComfyUI,没有 yaml 语法错误、模型路径不报错就说明接上了。
实际生成耗时
装好之后我用 MiniMax H3 实测了一轮:
| 项目 | 值 |
|---|---|
| 分辨率 | 672 × 1216(竖版) |
| 时长 | 10 秒 |
| 耗时 | 8 分钟以内 |
这是 RTX 3060 12G + 32G 内存环境下的单次完整生成时间(含采样与解码)。作为参考,同一环境下不用 SageAttention 会更慢。
说明:这里只记录我自己的实测值,没有做严格的开关对照实验,所以不能据此反推加速比。想要准确的加速幅度,需要在相同种子、相同参数下分别用
--use-sage-attention和不加上各跑几次取平均。
排错清单
安装过程中容易卡住的地方,按出现顺序:
| 现象 | 原因 | 处理 |
|---|---|---|
| 安装 SageAttention 时报缺头文件 | 便携版没有 include/libs | 走第 2 步补齐 |
triton.experimental 导入失败 | triton-windows 版本不匹配 | 装指定的 3.7.1.post27 |
装完 import sageattention 失败 | wheel 的 cu/torch tag 选错 | 重新按实际版本选 wheel |
| 启动报 CUDA OOM | 多卡环境下模型被分散加载 | 见下方「多卡环境」 |
多卡环境的坑
如果机器上有多张显卡,ComfyUI 会尝试把模型分散加载到不同卡上,行为不太可控,容易 OOM。显式指定用哪一张:
$env:CUDA_VISIBLE_DEVICES="0"; .\python_embeded\python.exe -s .\ComfyUI\main.py --windows-standalone-build --use-sage-attention
把 0 换成你要用的卡号。启动后确认日志里 Device: 指向的是你选中的那张卡。
版本对应关系小结
整个链条对版本的依赖很强,任何一环变了都可能要重选:
Python 3.13 → 决定用哪个 include/libs 补齐包
Torch 2.13.0+cu130 → 决定用哪个 SageAttention wheel(cu130torch2.10.0andhigher)
triton-windows 3.7.1.post27 → SageAttention 依赖它,且 experimental 必须可导入
--use-sage-attention 参数本身没有版本门槛,它只是调用已安装的 sageattention 包。
原文地址
- 参考文章:https://note.com/sepiablue/n/n1a7096eb9136 (《【MiniMax H3】ComfyUIにSageAttentionを入れてVRAM 12GBグラボで30%高速化してみた》,作者かみもと)
- ComfyUI 启动参数定义:https://github.com/Comfy-Org/ComfyUI/blob/master/comfy/cli_args.py
- ComfyUI 便携版下载:https://github.com/Comfy-Org/ComfyUI/releases/latest
- SageAttention Windows wheel:https://github.com/woct0rdho/SageAttention/releases
- triton-windows:https://github.com/woct0rdho/triton-windows/releases
说明
这篇是我的安装记录,环境是 RTX 3060 12G + 32G 内存。具体提速幅度我没有贴数字——不同显卡、不同分辨率和帧数下差异比较大,建议装好后自己在同一套参数下跑两遍(有/无 --use-sage-attention)对比,控制台会打印每次生成的耗时。
原文作者在 RTX 4070 12G 上用 MiniMax H3 5 秒视频(544×832、24fps)测得 342 秒降到 242 秒,约 29% 的提速。参考文章里还提到 MiniMax H3 的加速 LoRA 陆续有人做出来,那是另一条可以叠加的路子,本文没有覆盖。