Skip to content

Opening book details…

Can I read Iterative Deepening Sampling for Large Language Models on EtoBox?

Iterative Deepening Sampling for Large Language Models by Chen, Weizhe; Koenig, Sven; Dilkina, Bistra is a scholarly article available to read on EtoBox.

What is Iterative Deepening Sampling for Large Language Models about?

The recent release of OpenAI's o1 models and other similar frameworks showcasing test-time scaling laws has demonstrated their exceptional capability to tackle complex reasoning tasks. Inspired by this, subsequent research has revealed that such test-time scaling laws hinge on the model's ability to search both within a single response (intra-response) and across multiple responses (inter-response) during training. Crucially, beyond selecting a single optimal response, the model must also develop robust self-correction capabilities within its own outputs. However, training models to achieve effective self-evaluation and self-correction remains a significant challenge, heavily dependent on the quality of self-reflection data. In this paper, we address this challenge by focusing on enhancing the quality of self-reflection data generation for complex problem-solving, which can subsequently improve the training of next-generation large language models (LLMs). Specifically, we explore how manually triggering a model's self-correction mechanisms can improve performance on challenging reasoning tasks. To this end, we propose a novel iterative deepening sampling algorithm framework designe

Author
Chen, Weizhe; Koenig, Sven; Dilkina, Bistra
Published
2025
Language
EN