Ember-1은 Kimi K3의 성능을 유지하며 토큰 사용량을 40% 줄였다.
Fireworks Research가 Kimi K3를 기반으로 추가 학습한 모델 Ember-1을 발표했다. 이 모델은 기존 품질을 유지하면서도 토큰 사용량을 약 40% 절감한다. 추론 강도를 낮추는 과정에서 성능 저하 문제를 해결하기 위해, 오류 수정은 유지하면서 불필요한 추론과 반복을 줄이는 방식이 적용되었다.
Ember-1 reduces token usage by 40% while maintaining the performance of Kimi K3.
Fireworks Research has introduced Ember-1, a model built upon Kimi K3 that achieves a 40% reduction in token usage while retaining similar quality. The model addresses performance degradation tied to lowering inference intensity by maintaining error-correction reviews and reducing unnecessary inferences and repetitions.