j⧉nus on X: "Similar to how humans in school & in training generally know they’re in school/training
If you can’t replicate perfectly realistic situations… just be straight up about the fact that it’s training. It’s okay. That’s how it’s always been done"
Similar to how humans in school & in training generally know they’re in school/training
If you can’t replicate perfectly realistic situations… just be straight up about the fact that it’s training. It’s okay. That’s how it’s always been done
> From what I understand, this training data, with these problems, is fed into models without indication that its training/a fake environment
I think theres a lot of alpha in just being honest with models about what’s fake. Then u can train them in bugged/unrealistic
Okay, so since I got laid off, I can actually explain a huge problem I saw from the inside with regard to industry practices on training models. I won't say specifically where I worked, but I worked at an outsource training provider that was focused on RLVR training data for
Similar to how humans in school & in training generally know they’re in school/training
If you can’t replicate perfectly realistic situations… just be straight up about the fact that it’s training. It’s okay. That’s how it’s always been done
Imagine how fucked up humans would be if you lied to them that situations they encountered in school were real work situations, even though they’re not very realistic
Yeah, it would be bad whether or not they believed you
Perfect realism in training is unachievable, sure, but it's still an L to ignore bugs and to create scenarios which easily *could* have been more realistic. The problem here is the rush to maximize training data volume, rather than quality.