Opus 5의 장기 코딩 벤치마크 성능에 대한 분석.
SlopCodeBench를 사용한 장기 코딩 벤치마크에서 Opus 5는 17개 체크포인트 중 4개만 통과하며 지속적인 코드 발전에는 한계가 있다는 결과를 보였습니다. 이는 Opus 5가 신뢰할 수 있는 수준에 도달하지 못했다는 것을 나타냅니다. 테스트는 새로운 요구사항을 추가하는 방식으로 진행되었습니다.
Analysis of Opus 5's performance in long-term coding benchmarks.
In the long-term coding benchmark using SlopCodeBench, Opus 5 only passed 4 out of 17 checkpoints, indicating it is not yet reliable for continuous codebase development. This assessment shows limitations in Opus 5's ability to progress without intervention. The tests involved introducing new requirements at each checkpoint.