Welcome to two-tier system where trusted partners get access to ungated models for defense, while everyone else gets the guardrailed versions that can’t even analyze exploit payloads. Regular users are locked out of the defensive tools entirely, leaving them vulnerable while
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
Grok 4.5 is an impressive step up but it still can't schedule a reminder in Hermes Agent. It's a test case for me because cronjob tool description there is broken, but GPT always nails it on first attempt.
The more I work with Opus 4.6, 4.8 and GPT-5.5 (through Codex) the more GPT is growing on me. GPT almost combines the best of both worlds - excellent coding capabilities and human connection of Opus 4.6 that got lost somewhere chasing that tasty token money.
I run 3 agents on Grok through Hermes (agent framework by NousResearch) and ran into two tool-calling issues on the API.
grok-4.3: tool calling is unreliable at any context depth. The model plans tool calls in reasoning ("I need to call image_generate with a fresh prompt") but