Greg Durrett
NYU Courant
LLM Reasoning Beyond Scaling
Greg Durrett · NYU Courant
Agentic large language models can write and debug complex code, solve competition-level math problems, and conduct in-depth literature review. These reasoning capabilities are enabled by scaling of data: pre-training data to learn vast knowledge, fine-tuning data to learn natural language reasoning, and RL environments to refine that reasoning. In this talk, I will investigate the current LLM reasoning paradigm, its boundaries, and the future of LLM reasoning beyond scaling. First, I will describe the state of reasoning models and where I think scaling will lead to additional successes. I will then shift to discussing issues which are not resolved by pure scaling. First, I will describe our work on calibrating models' decisions through better understanding of their environments. We find that explicitly telling an LLM its likelihood to succeed or fail at tasks allows it to reason about cost-benefit tradeoffs in its action space. Then, I will describe our new benchmark CREATE, which tests LLMs' capabilities for associative creativity. I will highlight limitations of LLMs applied to creative tasks like scientific ideation and where I see future work making progress in these areas.