AI-ML·중요도 7·2026. 05. 27.·r/MachineLearning

Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA [R]

── KO ──────────────────

Triton에서 CUDA 없이 NVIDIA와 AMD에서 사용 가능한 Mixture-of-Experts 커널을 소개합니다.

새로운 프리프린트에서는 OpenAI Triton으로 작성된 Mixture-of-Experts 추론 커널(TritonMoE)을 소개합니다. 이 커널은 NVIDIA와 AMD 간에 포터블하며, 공급업체 특정 코드 없이 작동합니다. 주요 성과로는 SwiGLU 프로젝션을 위한 게이트와 GEMM의 융합 계산으로 35%의 전역 메모리 트래픽을 줄였습니다. 하지만 2048 토큰 이상에서 성능이 저하됩니다.


── EN ──────────────────

Introducing a Mixture-of-Experts kernel in Triton for cross-platform use without CUDA.

A new preprint presents the Mixture-of-Experts inference kernel (TritonMoE) written entirely in OpenAI Triton, targeting portability across NVIDIA and AMD. The major highlight is the fused gate+up GEMM that reduces global memory traffic by 35%. While it achieves impressive throughput for batch sizes up to 512 tokens, it has limitations when exceeding 2048 tokens.

원문 보기 →목록으로