LLM Evaluation Framework
Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study

Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study

10/22/2025

What this post added

This post introduces the ReasonIF benchmark, a systematic evaluation for assessing instruction-following abilities within the reasoning traces of Large Reasoning Models (LRMs). It highlights that frontier LRMs fail to follow reasoning instructions more than 75% of the time, with performance degrading further as task difficulty increases. The benchmark consists of 300 math and science problems paired with six types of verifiable user-oriented directives (multilinguality, word limit, disclaimer, JSON formatting, uppercase only, remove commas) that models must obey throughout their step-by-step solutions. The post presents findings that demonstrate a significant drop in instruction-following scores within reasoning traces compared to main responses, even for top-performing models.

Read the original post ↗