View this thread on X
Towards Looped Models Done Right — Part II: Rethinking at Fixed Points
Posted: Sep 29th 2026
Looped LMs need not pay for every loop: near fixed points, they can scale FLOPs without scaling memory. Fixed-point shortcuts speed up training (pretraining, post-training) and inference (prefill, decoding), while a learned depth prior and orthogonal injection shape better fixed points. With a 3× smaller KV cache, the learned prior matches fixed-depth training at 1.6B.