안소프는 Claude의 안전성을 위해 한계점을 설정하는 아키텍처를 설명했다.
Anthropic은 제품 전반에 걸쳐 Claude를 안전하게 유지하기 위한 아키텍처를 자세히 설명하였다. 에이전트의 안전성은 권한 프롬프트나 안전장치가 아닌 파일 시스템, 네트워크 및 실행 환경에 대한 결정론적 한계 설정에 달려 있다고 주장한다. 특히 신뢰 경계에서의 실패와 허용된 이그레스 경로를 통하여 이러한 설계를 수정하게 된 배경을 다룬다.
Anthropic discusses containment architectures for Claude to ensure agent safety.
Anthropic detailed the containment architectures it uses for Claude in its products. It argues that agent safety depends on placing deterministic limits on an agent’s filesystem, network, and execution environment rather than relying on permission prompts or safeguards. Notably, it examines failures at trust boundaries and permitted egress paths that led to revisions in their designs.