Sources
Loading...
Additional media
Loading...

Researchers from DeepSeek-AI and Peking University have introduced a novel strategy called Loss-Free Balancing for Mixture-of-Experts (MoE) models. This approach eliminates the need for auxiliary loss by dynamically adjusting expert biases, ensuring optimal load distribution. The strategy is particularly effective for models with 1B-3B parameters, enhancing performance across 100B-200B tokens. Additionally, a new MoE architecture named Nexus has been developed, focusing on efficiency, specialization, and adaptability. Nexus activates only 30-40% of the model's parameters, making it run up to 1.86 times faster than similar dense models like Mistral-7B and between 1.50 and 1.71 times faster than comparable MoEs such as DeepSeekMoE-16B.




