
10/22/2025
What this post added
This post introduces the ReasonIF benchmark, a systematic evaluation for assessing instruction-following abilities within the reasoning traces of Large Reasoning Models (LRMs). It highlights that frontier LRMs fail to follow reasoning instructions more than 75% of the time, with performance degrading further as task difficulty increases. The benchmark consists of 300 math and science problems paired with six types of verifiable user-oriented directives (multilinguality, word limit, disclaimer, JSON formatting, uppercase only, remove commas) that models must obey throughout their step-by-step solutions. The post presents findings that demonstrate a significant drop in instruction-following scores within reasoning traces compared to main responses, even for top-performing models.