Can I read Reinforcement Learning For Reasoning in Large Language Models With Training Example on EtoBox?
Reinforcement Learning For Reasoning in Large Language Models With Training Example by peacemakupe2010 is a document available to read on EtoBox.
What is Reinforcement Learning For Reasoning in Large Language Models With Training Example about?
This document presents a study on the effectiveness of reinforcement learning with verifiable reward using only one training example (1-shot RLVR) to enhance the mathematical reasoning capabilities of large language models (LLMs). The findings demonstrate significant performance improvements on mathematical benchmarks, achieving up to 73.6% accuracy on MATH500 with just one example, and highlight phenomena such as post-saturation generalization and cross-category generalization. The study emphasizes the imp
- Author
- peacemakupe2010
- Language
- EN