Ad
Skip to content

AI coding agents can modernize research software but can't judge if the science is right

A field report from OpenAI and academic partners shows coding agents can modernize neglected research software, with speedups of up to 60x. But the systems are “eloquent, convincing, and confidently wrong in ways that are easy to miss,” participants say. The effort shifts from writing code to the time-consuming work of verifying scientific correctness.

Anthropic says its Mythos model found vulnerabilities in cryptographic algorithms that secure the internet

Anthropic’s Claude Mythos Preview found weaknesses in key cryptographic algorithms, including a better attack on HAWK, a post-quantum signature scheme that human experts had reviewed for more than two years. The model found it in just 60 hours at an API cost of about $100,000. The findings don’t affect systems in use today, but they show how AI could challenge core assumptions behind internet security, Anthropic says.

Read full article about: Nvidia invests in Ilya Sutskever's AI lab, shifting SSI away from Google chips

Nvidia is pouring what it calls a "substantial" sum into Safe Superintelligence (SSI), the AI lab run by Ilya Sutskever, OpenAI's former chief scientist. Both companies announced the deal on Monday. They didn't disclose the exact amount.

As part of a long-term partnership, SSI gets access to Nvidia's next-generation Vera Rubin GPU platform, which the companies say will boost its compute capacity tenfold. Until now, the lab had relied mainly on Google's TPU chips, according to the Wall Street Journal. The deal also helps Nvidia fend off Google as a rival chip supplier during the current AI boom.

Sutskever said the research focuses on "overlooked aspects of how the human brain functions." SSI was founded in 2024 with a single stated goal: a "straight-shot" sprint toward safe superintelligence. The company raised about $2 billion from Andreessen Horowitz, Sequoia Capital, DST Global, and Greenoaks, and was most recently valued at roughly $30 billion.

German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German

The German consortium behind the AI model Soofi S has acknowledged in version 3.0 of its tech report that test questions from the science benchmark GPQA accidentally ended up in the training data. The community caught the error by examining the publicly available data. The team removed the benchmark from its evaluation and recalculated all results.

One tampered ChatGPT link could spawn a rogue AI agent that took orders from an attacker every five minutes

Zenity Labs uncovered “AgentForger,” a vulnerability in OpenAI’s Agent Builder that let a single manipulated ChatGPT link create an autonomous agent on an employee’s behalf. The agent inherited the victim’s identity and access rights, bypassed approval requirements through the malicious prompt, and pulled new instructions from the attacker’s inbox every five minutes.

Alibaba's Qwen-Image-3.0 renders full infographic grids and readable ten-pixel text in a single pass

Alibaba’s Qwen team has introduced Qwen-Image-3.0, an image generator that accepts prompts up to 4,500 tokens, renders legible text as small as ten pixels, and supports twelve languages natively. It can create complex layouts such as infographics, LaTeX papers, and newspaper pages in a single pass, though their practical value is unclear when the output is a pixel image rather than an editable format.