Real-SWE는 비공식 코드베이스에서 AI 모델의 성능을 평가하는 벤치마크입니다.
Real-SWE는 기업에서 사용되는 비공식 프로덕션 코드베이스를 이용하여 AI 모델이 소프트웨어 엔지니어의 작업을 수행할 수 있는지를 평가하는 새로운 벤치마크입니다. 이 평가는 모델 단독이 아닌 모델과 기본 하네스의 조합을 통해 이루어집니다. 각 과제마다 독립적으로 8회의 실험이 진행되며, 그 결과를 평균하여 모델의 성능을 분석합니다.
Real-SWE benchmarks AI model performance using unofficial production codebases.
Real-SWE is a new benchmark that evaluates AI models' ability to perform tasks of software engineers using unofficial production codebases permitted by companies. This evaluation is based not only on the model alone but on the combination of the model and its underlying harness. For each task, eight independent runs are executed, and the average results are analyzed to assess the model's performance.