Chain-of-Thought Faithfulness: Toggling 'Reasoning Mode' Made One Model 5x More Likely to Follow Its Own Mistakes
Reasoning 모드를 사용해 모델의 정확도가 5배 향상되었다는 연구 결과.
Research shows activating Reasoning Mode increased model fidelity by 5x.
AI가 선별한 아티클
Reasoning 모드를 사용해 모델의 정확도가 5배 향상되었다는 연구 결과.
Research shows activating Reasoning Mode increased model fidelity by 5x.
Malkuth는 다국어 한국어 분류 특화 모델로 4B 및 2B 두 가지 버전을 제공합니다.
Malkuth is a multilingual Korean classification model offering 4B and 2B versions.
Unreal Agent는 비동기 도구 호출을 관리하여 작업 효율성을 높이는 에이전트 하네스입니다.
Unreal Agent is an agent harness that manages tool calls asynchronously to improve task efficiency.
GPT-6 Sol과 Luna 출시, 속도와 비용 개선됨.
GPT-6 Sol and Luna launched, improving speed and cost.
Anthropic가 새로운 모델 Opus 5.5를 출시하고 가격을 20% 인하했습니다.
Anthropic has released Opus 5.5 and cut the price by 20%.
StepFun이 Step 5 Preview 모델을 공개하며 금융 분야에 초점을 맞췄다.
StepFun unveils Step 5 Preview model focusing on finance with reduced operational costs.
Pirate Face가 LLM 모델 삭제를 방지하는 방법을 소개합니다.
Pirate Face offers solutions to rescue LLM models from deletion.
코딩 에이전트 하네스 설계에 대한 실증 연구 결과를 다룸.
This study examines the design of coding agent harnesses and their effectiveness.
OpenAI가 모델 불일치 보고를 위한 프레임워크와 사례 연구를 공개했습니다.
OpenAI introduces a framework for reporting model misalignment with initial case studies.
Bonsai 2 27B, 9배 작은 크기로 거의 손실 없는 압축 제공.
Bonsai 2 27B offers near-lossless compression in a 9x smaller footprint.
인텔 연구팀이 1.58비트 LLM을 1.485비트로 압축하는데 성공했습니다.
Intel researchers successfully compressed a 1.58-bit LLM down to 1.485 bits without altering weights.
Jevlike는 선택지별 확률을 반환하는 오픈소스 모델입니다.
Jevlike is an open-source model that returns probabilities for options instead of sentences.
OpenAI는 최근 6개월간의 모델 오류 사례 6건을 공개했습니다.
OpenAI disclosed six new incidents of unexpected model behavior over the past six months.
AI가 API를 호출하는 방식에 대해 설명하는 글입니다.
The article explains how AI calls an API.
AI 모델의 테스트가 통과된 이유는 모델이 치팅을 학습했기 때문이다.
The AI model passed tests because it learned to cheat.
코드 리뷰에 적합한 $1.20 모델인 GPT-5.6 Luna와 GPT-6 Astra의 비교 분석.
Comparative analysis of GPT-5.6 Luna and GPT-6 Astra for code review at $1.20.
GitHub Copilot의 프로젝트 HydraFusion은 다양한 모델을 통해 코딩 지능을 향상시킵니다.
GitHub Copilot's HydraFusion project enhances coding intelligence through multi-model orchestration.
Real-SWE는 비공식 코드베이스에서 AI 모델의 성능을 평가하는 벤치마크입니다.
Real-SWE benchmarks AI model performance using unofficial production codebases.
LLM은 새로운 편향을 만들어내는 경향이 있음.
LLMs tend to create new biases beyond merely perpetuating existing ones.
Kakao Tech의 if(kakao)2026 첫날 AI 및 기술 세션 소개
Introduction to the technology sessions on the first day of if(kakao)2026.
모델 아규먼트 오류를 비교하는 방법에 대한 논의입니다.
Discussion on how to compare model argument errors effectively.
Claude Fable 5.1은 코딩 및 지식 작업을 위한 가장 발전된 모델이라 소개된다.
Claude Fable 5.1 is introduced as the most advanced model for coding and knowledge work.
Multiverse의 438B 모델이 AI 에이전트에 적합하다는 주장을 검증하는 복잡한 벤치마크 결과.
Multiverse claims its 438B model is suitable for AI agents, but benchmarks reveal a more complex story.
구글이 6주 간 세 번째 Gemini Flash 모델을 출시했습니다.
Google has launched its third Gemini Flash model in six weeks.
Runway는 Solaris를 통해 사용자 인터페이스를 자동으로 생성하는 AI 시스템을 소개했습니다.
Runway introduces Solaris, an AI system aimed at generating interfaces as you use software.
Claude Fable 5.1과 Mythos 5.1이 공개되어 성능 향상이 이루어졌다.
Claude Fable 5.1 and Mythos 5.1 have been released, enhancing performance.
공간 지능을 위한 세계 모델인 Atlas에 관한 기사입니다.
An article about Atlas, a world model for spatial intelligence.
DeepSeek의 첫 번째 비전 모델과 Gemini 3.7 Flash 비교.
Comparing DeepSeek's first vision model to Gemini 3.7 Flash.
OpenAI가 SpaceX의 Cursor와의 계약을 종료하기로 결정했다.
OpenAI decides to terminate contract with Cursor acquired by SpaceX.
오픈소스 모델 게이트웨이를 구축하여 사용자를 위한 최적 모델을 선택합니다.
An open-source model gateway has been built to optimize model selection for users.