Verified project record
THUDM/slime
slime is a lightweight, opinionated reinforcement learning post-training framework that connects Megatron training with SGLang rollout. It serves as the RL framework behind the GLM model family and supports agentic, multi-turn, and custom data generation workflows.