Virtio-nvgpu - KVM 가상 머신에서 NVIDIA GPU를 네이티브에 가까운 성능으로 공유
GPU 성능을 거의 네이티브처럼 여러 KVM 가상 머신에서 공유할 수 있는 프로젝트 소개.
Introducing a project that enables multiple KVM VMs to share a single GPU almost natively.
AI가 선별한 아티클
GPU 성능을 거의 네이티브처럼 여러 KVM 가상 머신에서 공유할 수 있는 프로젝트 소개.
Introducing a project that enables multiple KVM VMs to share a single GPU almost natively.
Canvas UI가 35개의 HTML-in-Canvas 컴포넌트를 출시하여 GPU 효과를 실시간으로 구현합니다.
Canvas UI releases 35 HTML-in-Canvas components that render GPU effects on live content.
AI 토큰 사용 비용이 낮아지며 품질과 인프라의 중요성이 커지고 있다.
As AI token costs drop, quality and infrastructure become increasingly important.
Three.js로 구현한 FBO 파티클 데모에 대한 설명.
Description of an FBO particle demo implemented with Three.js.
dlab은 로컬 AI 추론 및 연구 시스템을 공개하며, 적은 자원으로도 경쟁할 수 있는 목표를 세웠다.
dlab plans to unveil a local AI inference and research system to compete with advanced labs using minimal resources.
Mini-AGI는 8GB VRAM GPU에서 학습할 수 있는 동적 지속 학습 모델이다.
Mini-AGI is a dynamic continual learning model that operates on an 8GB VRAM GPU.
Three.js를 활용해 브라우저에서 LLM을 실행하는 방법을 소개합니다.
Introducing how to run LLMs in the browser using Three.js.
Claude 모델이 생체분자 모델을 4배 가속화했다.
Claude model accelerates over 30 biomolecular models by 4 times.
Bend는 AI 구현의 정확성을 증명하는 새로운 프로그래밍 언어입니다.
Bend is a new programming language that proves AI implementation accuracy.
Bend는 CPU와 GPU에서 AI 실수를 방지하는 언어입니다.
Bend is a language that prevents AI mistakes via proof on CPU and GPU.
마이크로소프트가 AI 워크로드 관리를 위해 TauGrid를 오픈 소스화했습니다.
Microsoft open-sources TauGrid for managing AI workloads in Kubernetes.
Nvidia가 Rust를 통한 GPU 프로그래밍 지원을 발표했습니다.
Nvidia announced support for GPU programming using Rust.
M4 Mac Mini와 MacBook Neo용 OpenGL ES 3.0 GPU 드라이버가 개발됐다.
An OpenGL ES 3.0 compliant GPU driver for M4 Mac Mini and MacBook Neo has been developed.
M4 Mac Mini를 위한 Linux GPU 드라이버 개발기를 보여줍니다.
This article chronicles the development of a Linux GPU driver for the M4 Mac Mini.
Perplexity의 새로운 에이전트는 GPU에서 완전 작동하지만 고비용이 따른다.
Perplexity's new agent runs entirely on your GPU but comes with a high cost.
Windows에서 AMD를 위한 CUDA 프로젝트 소개.
Introduction to a CUDA project for AMD on Windows.
Nvidia가 AI의 중앙은행 역할을 하고 있다는 주제의 기사입니다.
The article discusses Nvidia's role as the central bank of AI.
네이티브 IDE Rune이 오픈 소스로 공개되었습니다.
Native IDE Rune has been released as open source.
Cognition이 GPU 최적화를 통해 RSA-260 소인수분해에 성공했습니다.
Cognition optimized a GPU factoring pipeline to successfully factor RSA-260.
NVIDIA Personal AI Router는 로컬 컴퓨터들의 AI 작업을 자동으로 분배합니다.
NVIDIA Personal AI Router automates the distribution of AI tasks across local computers.
분산 AI 학습을 위한 신뢰할 수 있는 클라우드 네이티브 기반 구축
Building a reliable cloud native foundation for distributed AI training.
레드햇이 소프트웨어 엔지니어링 팀을 위한 AI 3.5를 출시했습니다.
Red Hat has released AI 3.5 to help software engineering teams run AI smoothly.
Kubernetes에서 다중 테넌트를 위한 보안 자기 서비스 GPU 메트릭을 다루는 글입니다.
Discusses secure, self-service metrics for multi-tenant GPU usage in Kubernetes.
Inception이 Mercury 2.5를 출시하며 품질을 개선하고 지능을 40% 향상시키다.
Inception releases Mercury 2.5, improving quality and intelligence by 40%.
Qwen3.8 27B의 4비트 양자화는 성능 저하 없이 용량을 줄일 수 있음을 보여줍니다.
Qwen3.8 27B shows that 4-bit quantization can significantly reduce size without performance loss.
AMD GPU에서 vLLM의 팩터링 디코딩 방법 소개.
Introduction to speculative decoding method in vLLM on AMD GPUs.
AI 플랫폼 엔지니어링은 GPU 이상의 이기종 인프라 문제다.
AI platform engineering is a heterogeneous infrastructure problem beyond GPUs.
Vortex를 통해 S3에서 GPU로 데이터 로딩 방식을 혁신합니다.
Vortex revolutionizes data loading from S3 to GPU.
GPU 노드에서 추론 시작 시간을 8분에서 1분 이하로 단축하는 방법에 대한 분석
Analysis of reducing GPU inference cold start time from 8 minutes to under a minute.
GPU World는 2040년까지 모든 사람이 GPU와 LLM을 사용하게 될 미래를 다룹니다.
GPU World envisions a future where everyone uses GPUs and LLMs by 2040.