AI 코딩 에이전트가 60%의 실패율에도 불구하고 새로운 벤치마크를 기록했다.
Claude Fable 5.1은 60% 이상 실패하는 상황에서도 새로운 코딩 벤치마크에서 승리했다. 이 모델의 성공률은 38.8%에 그치며, AI의 신뢰성에 대한 논의가 필요하다는 점을 시사한다. 데이터 분석을 통해 AI의 한계가 드러나고 있으며, 개발자들은 이러한 결과를 주의 깊게 살펴봐야 한다.
AI coding agent fails 60% of the time but won a new coding benchmark.
Claude Fable 5.1 has achieved a new coding benchmark despite failing over 60% of the time. Its success rate stands at 38.8%, highlighting the need for discussion about the reliability of AI. The analysis of this data reveals the limitations of AI, which developers should consider carefully.