We can now RL large MoEs with 0 train-infer mismatch! And doing so can improve performance (pictured task: teach Qwen3.6-35B-A3B to play Wordle). Everything is open-source and we did a bunch of ablations. 🧵
in “terrifyingly prescient things i read today”; this short story by @gwern that is apparently from 2022 yet gets a lot of things right about the current failure of frontier labs to secure sandboxes during training.
if i had 1 thing to add, it would be self-improving inference.
humans went from sharing text to images to gifs to videos, because we seem to get more out of higher-bandwidth mediums
llms have done the first couple, and is now getting to video gen, this is why you see xiaomi training on music!
the error appears to be a central credential issue, that’s why everyone is seeing the same unauthorized key ending in “fvMA”
looking forward to my reset :)