← October 8, 2026 briefing · added in the 20:00 KST update

Kakao presents a method at COLM 2026 that predicts the optimal learning rate for MoE pretraining at about 1% of total training cost (report)

Announced October 8, 2026

What happened

Kakao said on the 8th that it presented research at the main conference of COLM 2026, an international conference on language modeling, and at the tokenization workshop 'TokShop'. The main conference paper presents a method for predicting the optimal learning rate for pretraining large MoE (Mixture-of-Experts) models at about 1% of total training cost. At the workshop, Kakao presented a model compression technique that reduces model size while improving performance.

Provider claims

Kakao said that there is little public research on learning rates for MoE models, which makes finding the optimal value costly, and that its method can reduce this cost.

Why it matters

The method sharply cuts the cost of hyperparameter search, one of the most expensive parts of MoE pretraining. It bears directly on how cost-efficiently Korean companies can train their own large models.

Confidence medium · official source pending

Sources