Overview
trifuse 是面向 AI Infra 学习的精简 Triton 算子仓库,只保留三条可以用独立参考实现验证的 Transformer 推理路径:融合 RMSNorm+RoPE、标准 SwiGLU/GeGLU 的 Gated MLP,以及带在线 softmax 的 FlashAttention 前向。
项目状态 stable(v2.0.1):新功能暂停,继续维护正确性、兼容性与可复现验证。
Highlights
fused_rmsnorm_rope:融合 RMSNorm 与 RoPEfused_gated_mlp:activation(gate_proj(x)) * up_proj(x)flash_attention:在线 softmax 前向,支持 causal mask,是 cuflash 的独立参考实现import trifuse即注册进torch.ops.trifuse.*,可接入torch.compile/torch.export图- NumPy/PyTorch 参考实现、输入契约、差分测试、benchmark 与 autotuner 基础设施
- 删除了曾被误标为 “FP8 E4M3” 的 uint8 均匀量化路径,避免把 INT8 当 FP8 教材
Tech
Triton · PyTorch · torch.library · Python · 差分测试