How much of a reasoning model's output is actually thinking? For three model tiers at three task difficulties, stacked bars show thinking tokens dwarfing output — then a scheme matrix shows which verification schemes can see the thinking and which attest to the answer alone.
DeepSeek-R1 emits 8,400 thinking tokens before a 341-token answer on hard math. Every existing on-chain verification scheme — TEE, optimistic, ZK, sampling — was designed when computation and output were proportional. They aren't anymore.