AI 명령 승인 과정에서 인간의 위협 탐지 정확도가 낮다는 분석 결과.
AI 코딩 에이전트의 명령 승인 과정을 분석한 결과, 평균 위협 탐지 정확도가 66.3%에 불과했다. 이는 인간 승인이 최후의 보안 경계로서 기능하기 어렵다는 문제를 시사한다. 특히 명백한 파괴 명령의 누락률이 11.7%에 달하며, 자격 증명 접근과 같은 범위 위반도 포함되어 있어 심각하다.
Analysis shows low human threat detection accuracy in AI command approval.
An analysis of the command approval process for AI coding agents revealed that the average threat detection accuracy is only 66.3%. This raises concerns about the effectiveness of human approval as the last line of defense in security. Notably, the missed rate for explicit destruction commands was 11.7%, and there were significant scope violations like credential access, indicating a serious issue.