33 Comments
⭠ Return to thread
Jackson Hurley's avatar

"Lead us not into temptation". I wonder whether leading models into temptation (for example, by replicating token-for-token the contexts in which models previously took misaligned actions), then rewarding them for choosing the aligned action in that situation, might be a good move. This is not going to solve the problem of the model's behavior when it knows for sure it won't get caught, but as I wrote in my other comment, if the Watchers are sufficiently competent, the model can't be certain it isn't in an elaborate sting operation.