About this document
Controllable Test-Time Scaling for LRMs by thepakumar03 is a document available to read on EtoBox.
The paper introduces Reasoning Control Fields (RCF) to address underthinking and overthinking in long chain-of-thought reasoning for Large Reasoning Models (LRMs). It presents the Control-R-4K dataset, which contains challenging problems with detailed reasoning annotations, and proposes Conditional Distillation Finetuning (CDF) to enhance model adaptability. Experimental results demonstrate that the Control-R-32B model achieves state-of-the-art performance while allowing for controllable reasoning processes
- Author
- thepakumar03
- Language
- EN