AI-ML·중요도 7·2026. 07. 16.·GeekNews

Databricks가 자체 코딩 AI 벤치마크를 만든 방법과 결과

── KO ──────────────────

Databricks가 코딩 AI 벤치마크를 개발한 과정과 결과를 소개합니다.

Databricks는 회사별 맞춤형 코딩 AI 벤치마크를 개발했습니다. 공개된 코딩 벤치마크는 치팅이 가능해 문제였고, 이를 개선하기 위해 실제 수행한 과제를 기반으로 평가 방식을 만들었습니다. GLM 5.2는 Opus 4.8과 비슷한 성능을 보이며 가격은 더 저렴합니다.


── EN ──────────────────

Databricks introduces its coding AI benchmark development process and results.

Databricks has developed a tailored coding AI benchmark that addresses the issues of publicly available benchmarks, which can be cheated. To improve this, they created an evaluation method based on actual tasks performed. GLM 5.2 demonstrated performance comparable to Opus 4.8 while being priced at only 66% of Opus's costs.

원문 보기 →목록으로