The Illusion of the Illusion of Thinking: A Comment on Shojaee et al. (2025)
Descripción general
Resumen del artículo
The paper argues that a previous study's findings of "accuracy collapse" in Large Reasoning Models on complex planning puzzles are due to experimental design limitations, specifically output token limits and unsolvable problem instances. By using alternative representations that bypass these limitations, the authors suggest that models can solve tasks previously deemed too complex.
Explícamelo como si tuviera cinco años
Scientists found that when AI seemed to fail hard puzzles, it was often because the test was unfair or the puzzles were impossible. When the tests were made fairer, the AI could solve them after all!
Posibles conflictos de intereses
The authors are affiliated with Anthropic and Open Philanthropy, which may have interests in promoting positive views of AI capabilities. However, the critique primarily addresses methodological concerns, reducing the likelihood of significant bias.
Limitaciones identificadas
Explicación de la calificación
The paper provides a valuable critique of experimental design in AI research, highlighting the importance of considering output constraints. However, its lack of novel findings, reliance on preliminary tests, and anecdotal evidence limits its impact.
Conviene saber
Este es el análisis de Starter. Paperzilla Pro verifica cada cita, investiga los antecedentes de los autores y las fuentes de financiación, y utiliza razonamiento avanzado con IA para ofrecer información más exhaustiva.
Explorar Pro →