Skip to content

moe-training

from Orchestra-Research/AI-research-SKILLs

Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. Use when training large-scale models with limited compute (5× cost reduction vs dense models), implementing sparse architectures l

v1.0.0MIT
527
Lines
1,659
Words
20
Code Blocks

Languages

bashjsonpython
19-emerging-techniques/moe-training/SKILL.md