Introducing Agent Mode: Agentic AI is now measured in the Arena.
Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more.
It completes more complex tasks by using tools like web search, bash in a sandbox environment, image
Qwen3.8-Flash-Next by @Alibaba_Qwen has landed in Agent Arena, ranking #7 among open models (#24 overall) with +2.4% net improvement across 8.7K+ real-world agentic sessions!
Among open models, it sits just behind DeepSeek V4 Flash (High) at #6 (+3% net improvement) and two
⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.
125B parameters + 51B N-gram
Exciting news: GLM-5.3-Flash by @Zai_org has landed in Agent Arena! At a $0.12 median cost per task and a +4.6% net improvement, it has reshaped the Pareto frontier! GLM-5.3-Flash is placed between DeepSeek V4 (High) and GPT-5.6 Luna (xHigh).
Based on 9K+ real-world agentic
NVIDIA’s has compressed release cycles from every 6–8 months to every 4–6 weeks. Bryan Catanzaro @ctnzr, the VP of Applied Deep Learning Research at @NVIDIAAI, breaks down how they are doing it.
00:00 Why AI became core to NVIDIA’s mission
00:55 Infrastructure, data, and
Big news: Hy4 preview by @TencentHunyuan just landed ~#5 in the Code Arena: WebDev with 1633 pts (AutoEval).
This is a significant improvement from Hy3 at #31 overall (+115 pts)! Among open models, Hy4 preview is ~#3, compared to Hy3 at #7.
Note: this is an early AutoEval