🧠 "Therefore I am. I Think" - shows that reasoning LLMs often decide first, then think - not the other way around. Linear probes decode tool-calling decisions from pre-reasoning activations at >90% AUROC, and activation steering flips behavior 7-79% of the time, with the CoT
Our results suggest that reasoning models can encode action choices before visible deliberation, and that CoT can sometimes rationalize rather than drive those choices.
Read our full paper here: arxiv.org/pdf/2604.01202
w/ @den_run_ai@RajeswarSai (Raj Venkat)