AI-ML·중요도 7·2026. 09. 10.·Hacker News
Getting 50 GB/S Back from the Apple Neural Engine
── KO ──────────────────
애플 M3 신경망 엔진의 성능 저하 이슈와 Llama 3.2의 속도 개선 방법을 설명합니다.
애플 M3 Neural Engine에서 발생한 RTL 성능 오류로 인해 DRAM 무게 스트리밍 처리량이 45-60 GB/s에서 17–19 GB/s로 감소했습니다. 커널 DMA 엔진의 문제 경로를 피함으로써 Llama 3.2 1B 토큰 처리량이 10.0에서 24.3 토큰/s로 증가했습니다. 이 글은 성능 향상의 구체적인 방법을 다룹니다.
── EN ──────────────────
Discusses performance issues in Apple's M3 Neural Engine and how to improve Llama 3.2's speed.
The article addresses an RTL performance error in the Apple M3 Neural Engine that reduced DRAM weight streaming throughput from a nominal 45-60 GB/s to 17-19 GB/s. By avoiding the problematic path in the kernel DMA engine's speculative prefetch ring, throughput for Llama 3.2 1B tokens was increased from 10.0 to 24.3 tokens/s. It provides insights into the methods for achieving this performance improvement.