got these results using Silico, our interp research platform. interp has great potential for model debugging and we're improving our ability to do so steadily!
You ask an AI model a question. Why does it answer the way it does?
Using the model's own neurons, we can trace how it makes decisions — and steer it toward better ones.
Case study: we tested an LLM and found that it sometimes endorses drunk driving.🧵 (1/5)




