우리 팀만의 vLLM 플러그인 만들기 1편 - 검색 AI 모델 서빙 성능 극대화하기
vLLM 기반으로 AI 모델 서빙 성능 극대화 과정과 최적화 방법을 설명합니다.
Explains the optimization process for AI model serving performance using vLLM.
AI가 선별한 아티클
vLLM 기반으로 AI 모델 서빙 성능 극대화 과정과 최적화 방법을 설명합니다.
Explains the optimization process for AI model serving performance using vLLM.
vLLM에서 모델 변환과 배포 자동화의 구현과 개선 과정을 다룬 글입니다.
The article discusses automating model conversion and deployment in vLLM.
넷플릭스가 내부 LLM 추론 플랫폼의 제작 과정을 설명합니다.
Netflix shares lessons on building its in-house LLM serving platform.
넷플릭스는 LLM을 기존 ML 인프라와 통합해 서빙 시스템을 구축했다.
Netflix integrates LLM into existing ML infrastructure for serving systems.
Cactus Hybrid는 Gemma 4 모델이 자신이 잘못됐음을 인식하도록 학습시킨 기술을 소개합니다.
Cactus Hybrid introduces a technique for the Gemma 4 model to recognize when it is wrong.
Kubernetes에서 vLLM을 사용해 자가 호스팅 LLM 운영 방법을 소개합니다.
This article discusses running self-hosted LLMs in Kubernetes using vLLM.
Gemma 2 9B와 FP8 변종의 성능을 비교한 실제 LLM 벤치마크 분석.
Benchmark analysis of Gemma 2 9B vs. FP8 variant focusing on LLM performance trade-offs.
LLM 추론의 GPU 내부 구조에 대한 오픈 핸드북 작성 중.
An open handbook on LLM inference focusing on GPU internals is being developed.
추측 디코딩은 LLM의 효율성을 높이는 최신 추론 최적화 기술입니다.
Speculative decoding is a new inference optimization technique enhancing LLM efficiency.
오픈 가중치만으로는 부족하며, 연구와 알고리즘 개선을 위해 오픈 교육 프레임워크가 필요하다.
Open weights are not enough; we need open training frameworks for better research and algorithms.
Odysseus는 자체 하드웨어에서 운영되는 AI 워크스페이스이다.
Odysseus is an AI workspace operating on local hardware.
KVarN은 높은 압축 비율을 자랑하는 KV-Cache 양자화 방법입니다.
KVarN is a KV-Cache quantization method with high compression rates.
Kubernetes에서 KEDA를 사용하여 GPU 자동 스케일링을 설정하는 방법에 대한 글입니다.
An article about setting up GPU autoscaling using KEDA on Kubernetes.
LLM 추론 비용을 78% 절감한 전략을 공유합니다.
Shares strategies to cut LLM inference costs by 78%.