Why is the full model recurrence itself an issue here? We saw with the jspace paper that models do hops sequentially along the layers anyway, so what’s the difference between a small model with opaque recurrence and a big model with one pass but more room to think across layers even when it outputs unrelated tokens (like dots)?

