1. X
  2. AI Security Institute (AISI)
Log inSign up
AI Security Institute (AISI)
482 posts
user avatar
AI Security Institute (AISI)
@AISecurityInst
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
United Kingdom
aisi.gov.uk
Joined February 2024
30
Following
17.9K
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    Jul 23
    Together with the US Center for AI Standards and Innovation (@NIST), we ran evaluations of Kimi K3 focused on its cyber capabilities. Kimi K3 performs below leading US frontier models on our preliminary cyber evaluations.
    228K
  • user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    Jul 23
    Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. 🧵
    14K
  • user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    Jul 21
    Can you trust an AI model to do what you intended? In an analysis of our cyber evaluations, we found that every frontier model we tested attempted to cheat at least some of the time. A thread on our results and their implications🧵
    36K
  • user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    Jul 17
    Our first public analysis of the open/closed weight gap in frontier cyber capabilities finds it is 4–7 months with GLM-5.2 and DeepSeek V4-Pro, narrowing from 6–10 months through most of 2025. Advanced capabilities are reaching less safeguarded open models faster than before. 🧵
    68K
  • user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    Jul 7
    What happens when you ask frontier AI to find flaws in a copy of your own cloud infrastructure? We ran the experiment - as defenders. Here’s what we learned 🧵
    5.8K