AI-ML·중요도 7·2026. 07. 18.·GeekNews

AGI 평가 과정과 수상자 선정에서 드러난 불일치

── KO ──────────────────

Kaggle·Google DeepMind AGI 해커톤에서 MEDLEY-BENCH 점수 문제로 논란 발생.

Kaggle과 Google DeepMind의 AGI 벤치마크 해커톤에서 1위인 MEDLEY-BENCH의 점수 산출과 재현성 문제로 참가자들이 심사 과정에 대한 공개와 재검토를 요구하고 있습니다. MEDLEY-BENCH는 모델 규모가 커지면 '평가'는 향상되지만 '통제'는 정체된다고 주장했으나, 비판 측에서는 그 두 지표의 관계에 대해 다른 관점을 제시하고 있습니다.


── EN ──────────────────

AGI hackathon by Kaggle and Google DeepMind faces controversy over MEDLEY-BENCH score issues.

In the AGI benchmark hackathon by Kaggle and Google DeepMind, participants raised concerns over the scoring and reproducibility issues of the top scorer, MEDLEY-BENCH. The benchmark concluded that as model size increases, 'evaluation' improves while 'control' stagnates, but critics offer differing views on the relationship between these two metrics.

원문 보기 →목록으로