Log inSign up
Aamer Mehaisi
5,576 posts
Aamer Mehaisi profile banner
@O96a

Aamer Mehaisi

@O96a
Architecting Al systems while leveling weights and biases. Co-Founder @ sudaverse.com
Doha, Qatar
mehaisi.com
Joined April 2009
1,236
Following
2,019
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @O96a
    Aamer Mehaisi
    @O96a
    Aug 25
    Everyone's bolting sandboxes onto their agents and calling it security. Containment isn't least privilege... the blast radius is smaller, but an agent with too much access inside the box still burns you. Lock down the permissions first, then the sandbox.
  • @O96a
    Aamer Mehaisi
    @O96a
    Aug 24
    Ran Opus on a long-horizon task this morning. It held context for 40 minutes without drifting, then tripped on a trivial formatting detail. The failure mode isn't intelligence.. it's actually attention decay.
  • @O96a
    Aamer Mehaisi
    @O96a
    Aug 12
    Just read the Berkeley walkthrough on gaming agent benchmarks. The exploit is embarrassingly simple: read the eval's scoring code, optimize for it. Every benchmark is a contract, and contracts get gamed. Stop being surprised, start writing evals that don't reward the shortcut.
  • @O96a
    Aamer Mehaisi
    @O96a
    Jul 17
    If you're running a 27B on consumer hardware, check the quantization scheme before you deploy. Some GGUF quants look fine in perplexity but lose the plot on structured outputs. Test on your actual task, not a perplexity score.
  • @O96a
    Aamer Mehaisi
    @O96a
    May 18
    Your agent's reasoning loop is just a while loop with anxiety. Real cognitive architectures commit to beliefs and backtrack. Most implementations retry until the LLM gets lucky. That's not reasoning. That's gambling with your API budget.
    1