Netflix Details Its In-House LLM Serving Platform with Triton and vLLM
넷플릭스가 내부 LLM 추론 플랫폼의 제작 과정을 설명합니다.
Netflix shares lessons on building its in-house LLM serving platform.
AI가 선별한 아티클
넷플릭스가 내부 LLM 추론 플랫폼의 제작 과정을 설명합니다.
Netflix shares lessons on building its in-house LLM serving platform.
넷플릭스는 LLM을 기존 ML 인프라와 통합해 서빙 시스템을 구축했다.
Netflix integrates LLM into existing ML infrastructure for serving systems.
구형 GPU에서도 실행 가능한 LLM 훈련 프레임워크 Picotron을 소개합니다.
Introducing Picotron, an LLM training framework that runs on older GPUs without crashing.
소프트맥스 없는 주의 모델이 공개되었습니다.
A softmax-free attention model has been released.
Triton에서 CUDA 없이 NVIDIA와 AMD에서 사용 가능한 Mixture-of-Experts 커널을 소개합니다.
Introducing a Mixture-of-Experts kernel in Triton for cross-platform use without CUDA.
Kubernetes에서 KEDA를 사용하여 GPU 자동 스케일링을 설정하는 방법에 대한 글입니다.
An article about setting up GPU autoscaling using KEDA on Kubernetes.
Triton 1.0 출시, CUDA 경험 없이도 효율적인 GPU 코드 작성 가능.
Triton 1.0 released, enabling efficient GPU coding without CUDA experience.