huggingface.co web signal

Rufus-Air Ships Open Eight-Stage Post-Training Recipe for GLM-4.5-Air Using Public Data Only

Open Source Fine-tuning ai-business

Summary

Rufus-Air documents a reproducible eight-stage post-training pipeline for GLM-4.5-Air-Base (106B total, 12B active) that moves from SFT through Reasoning RL, Coding RL, Instruction-Following RL, General/Coding/Search Agent phases and finally RLHF. The authors ordered stages by reward reliability — harder verifiable rewards first, softer judge-based signals later — and improved on the official GLM-4.5-Air release while staying competitive with comparably sized open models. The recipe relies only on open-source components and public data, aiming to serve as a transparent blueprint for teams outside frontier labs.