Lil'Log ยท 2025-05-01

Why We Think

<p><span class=&#34;update&#34;>Special thanks to <a href=&#34;https://scholar.google.com/citations?user=itSa94cAAAAJ&amp;hl=en&#34;>John Schulman</a> for a lot of super valuable feedback and direct edits on this post.</span></p> <p>Test time compute (<a href=&#34;https://arxiv.org/abs/1603.08983&#34;>Graves et al. 2016</a>, <a href=&#34;https://arxiv.org/abs/1705.04146&#34;>Ling, et al. 2017</a>, <a href=&#34;https://arxiv.org/abs/2110.14168&#34;>Cobbe et al. 2021</a>) and Chain-of-thought (CoT) (<a href=&#34;https://arxiv.org/abs/2201.11903&#34;>Wei et al. 2022</a>, <a href=&#34;https://arxiv.org/abs/2112.00114&#34;>Nye et al. 2021</a>), have led to significant improvements in model performance, while raising many research questions. This post aims to review recent developments in how to effectively use test-time compute (i.e. &ldquo;thinking time&rdquo;) and why it helps.</p>

Open original