Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
Cactus에서 출시한 Needle2는 14MB LLM으로 저전력 모바일 기기에서 높은 성능을 발휘합니다.
Needle2, released by Cactus, is a 14MB LLM optimized for performance on low-power mobile devices.
직접 따라 하며 실력을 키울 기술 글
Cactus에서 출시한 Needle2는 14MB LLM으로 저전력 모바일 기기에서 높은 성능을 발휘합니다.
Needle2, released by Cactus, is a 14MB LLM optimized for performance on low-power mobile devices.
Canva는 1억 개의 세션을 지원하는 세션 무효화 아키텍처를 공유했습니다.
Canva shared its session revocation architecture to support 100M sessions.
vLLM 기반으로 AI 모델 서빙 성능 극대화 과정과 최적화 방법을 설명합니다.
Explains the optimization process for AI model serving performance using vLLM.
vLLM에서 모델 변환과 배포 자동화의 구현과 개선 과정을 다룬 글입니다.
The article discusses automating model conversion and deployment in vLLM.
Shopify가 재고 예약 시스템을 Redis에서 MySQL로 이전한 사례.
Shopify transitioned its inventory reservation system from Redis to MySQL.
Postgres에서 분석 성능을 300배 향상시키는 방법에 대해 설명합니다.
Describes how to make Postgres 300x faster for analytics.
Spotify가 AI 코딩 에이전트 'Honk'를 통해 코드 베이스 마이그레이션을 다루는 방법을 설명합니다.
Spotify created 'Honk', an AI coding agent for complex codebase migrations.
AI 기능 스토어를 위한 저지연 데이터 레이어 최적화 방법을 소개합니다.
Optimizing data layers for low-latency workloads in AI feature stores is discussed.
오픈 모델이 비용 효율성에서 GPT-5.6 Sol을 초월하는 방법을 설명합니다.
Explains how open models outperform GPT-5.6 Sol in cost efficiency.
카카오는 Kanana-o 음성 생성 모델의 고도화 과정을 소개합니다.
Kakao introduces the enhancement process of the Kanana-o speech generation model.
CyberPanel의 SSL 자동 갱신 실패 문제와 해결 방법에 대해 설명합니다.
Discusses the issue of SSL auto-renewal failure in CyberPanel and how to fix it.
KEDA를 이용하여 Amazon SQS 큐 깊이에 따라 Kubernetes 파드를 스케일링하는 방법을 설명합니다.
This article explains how to scale Kubernetes pods with KEDA based on Amazon SQS queue depth.
DeepSeek를 기반으로 한 GPT-OSS의 자가 증류가 검열 특성을 전이하지 않는다는 사실을 보여주었습니다.
The self-distillation of GPT-OSS based on DeepSeek does not transfer censorship characteristics.
M1/M2 맥에서 2GB RAM으로 Gemma 4 26B 모델을 실행할 수 있는 오픈소스 엔진 TurboFieldfare를 소개합니다.
Introducing TurboFieldfare, an open-source engine running Gemma 4 26B model on M-series Macs with only 2GB RAM.
HR 소프트웨어 팀이 세 개의 모놀리스를 120개 도메인 마이크로서비스로 재구축한 사례입니다.
An HR software team transformed three monoliths into 120 domain microservices over five years.
Bun이 AI를 활용해 Zig에서 Rust로의 코드 이식을 11일에 완료한 사례.
Bun used AI to rewrite Zig code to Rust in just 11 days.
ESP32-S3에서 2,890만 매개변수 LLM을 실행하는 방법에 대한 기술적 설명.
Technical explanation of running a 28.9M parameter LLM on an ESP32-S3 microcontroller.
Claude Code를 이용해 대규모 코드 마이그레이션을 성공적으로 수행한 사례.
Successful large-scale code migration using Claude Code.
슬랙은 불변 AMI 기반의 간편한 EC2 플랫폼으로 현대적인 배포 방식을 도입하였다.
Slack has adopted an easier EC2 platform based on immutable AMIs for modern deployment.
잘란도는 초당 백만 요청을 처리하는 클라이언트 사이드 로드 밸런서를 구축했다.
Zalando built a client-side load balancer handling one million requests per second.
Claude 애플리케이션 구현을 위한 실전 가이드와 예제를 제공하는 문서입니다.
This document provides practical guides and examples for implementing Claude applications.
AI 기반의 오픈소스 macOS 비디오 편집기 Palmier Pro 소개.
Introducing Palmier Pro, an open-source macOS video editor powered by AI.
Learn OpenGL은 현대 OpenGL을 배우기 위한 광범위한 튜토리얼 자료를 제공합니다.
Learn OpenGL offers extensive tutorial resources for learning Modern OpenGL.
Kubeflow와 Cilium을 결합하여 Kubernetes에서 GPU 비율 문제를 분석하는 사례.
Debugging idle GPU resources in Kubernetes using Kubeflow and Cilium.
Cactus Hybrid는 Gemma 4 모델이 자신이 잘못됐음을 인식하도록 학습시킨 기술을 소개합니다.
Cactus Hybrid introduces a technique for the Gemma 4 model to recognize when it is wrong.
Bento는 모든 슬라이드를 하나의 HTML 파일로 편집하고 협업할 수 있는 도구입니다.
Bento is a tool to edit and collaborate on slides in a single HTML file.
Kubernetes에서 다중 클러스터 데이터베이스를 구축하는 방법을 설명합니다.
Explains how to build multi-cluster databases on Kubernetes.
DoorDash는 Envoy와 Valkey를 사용하여 1.5M RPS의 프로토타입 캐시를 구축했습니다.
DoorDash built a transparent proxy cache using Envoy and Valkey, handling 1.5M RPS with 99.99999% availability.
EKS에서 GPU 노드를 자동 복구하는 방법에 대한 사례 연구입니다.
A case study on automatically healing GPU nodes in EKS.
AI 뮤직비디오 제작 대결에서 Claude Fable 5와 GPT-5.6이 성공적으로 영상을 완성했다.
AI music video challenge shows Claude Fable 5 and GPT-5.6 successfully produced videos.