Can I read Polaris: Scaling RL for Reasoning Models on EtoBox?
Polaris: Scaling RL for Reasoning Models by roxeda9678 is a document available to read on EtoBox.
What is Polaris: Scaling RL for Reasoning Models about?
Polaris introduces the Polaris-4B-Preview and Polaris-7B-Preview, advanced open-recipe reasoning models that achieve high accuracy on AIME24 and AIME25 benchmarks, outperforming existing commercial models. The document outlines a post-training recipe focusing on data difficulty calibration, diversity-based rollout sampling, and inference-time length scaling to enhance reinforcement learning efficiency. The authors commit to open-sourcing their dataset, code, and training details to support further research
- Author
- roxeda9678
- Language
- EN