Arena’s cover photo
Arena

Arena

Research Services

San Francisco, California 26,717 followers

Where AI meets the real world.

About us

Created by researchers from UC Berkeley, Arena (formerly LMArena) is a community-powered platform to measure and advance the frontier of AI for real-world use. Tens of millions of builders, researchers, and creative professionals come to Arena to use frontier models and give feedback on their responses, shaping a public leaderboard grounded in real-world use.

Website
https://arena.ai
Industry
Research Services
Company size
51-200 employees
Headquarters
San Francisco, California
Type
Privately Held
Founded
2025
Specialties
AI evaluation, AI research, and AI community

Locations

Employees at Arena

Updates

  • View organization page for Arena

    26,717 followers

    Qwen3.8-Flash-Next by Qwen has landed in Agent Arena, ranking #7 among open models (#24 overall) with +2.4% net improvement across 8.7K+ real-world agentic sessions! Among open models, it sits just behind DeepSeek V4 Flash (High) at #6 (+3% net improvement) and two spots behind GLM-5.3 (Max) at #5 (+3.8%). By signal, Qwen3.8-Flash-Next stands out in delivering an explicit response from the community on task completion (Confirmed Success at +12.3%), ranking #5 among open models (#7 overall). By each signal, it landed: - Confirmed Success: +12.3% - Bash Recovery: +2.9% - Praise vs. Complaint: -1.6% - Steerability: -1.1% - No issues with Tool Hallucination Qwen3.8-Flash-Next outperforms Qwen3.8-27B in overall net improvement (+2.4% vs. +1.5%) and Confirmed Success (+12.3% vs. +7.2%). Qwen3.8 Max remains ahead overall at +6%, with +10.8% Confirmed Success. Congrats to the Qwen team on this contribution to the open ecosystem!

  • View organization page for Arena

    26,717 followers

    Exciting news: GLM-5.3-Flash by Z.ai has landed in Agent Arena! At a $0.12 median cost per task and a +4.6% net improvement, it has reshaped the Pareto frontier! GLM-5.3-Flash is placed between DeepSeek V4 (High) and GPT-5.6 Luna (xHigh). Based on 9K+ real-world agentic sessions, GLM-5.3-Flash ranks #4 among open source models and #19 overall, one spot ahead of GLM-5.3 (Max) at #20 with +3.7% net improvement. By signal, GLM-5.3-Flash stands out in Confirmed Success at +15.3%, ranking #4 there. By each signal, this model scored: - Confirmed Success: +15.3% - Praise vs. Complaint: +4.9% - Steerability: +2.4% - Bash Recovery: -0.7% - No issues with Tool Hallucination Congrats to the Z.ai team on this release!

  • View organization page for Arena

    26,717 followers

    Big news: Hy4 preview by Tencent Hy just landed ~#5 in the Code Arena: WebDev with 1633 pts (AutoEval). This is a significant improvement from Hy3 at #31 overall (+115 pts)! Among open models, Hy4 preview is ~#3, compared to Hy3 at #7. Note: this is an early AutoEval score, in which a Reward Model trained on Arena's human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. Congrats to the Tencent Hy team on this release!

    • No alternative text description for this image
  • View organization page for Arena

    26,717 followers

    Grok-4.6 (xHigh) by SpaceXAI has landed in Agent Arena, with its strongest category result in Code: ranking #12 with a +7.7% net improvement, based on 4.5K+ real-world agentic code sessions. Within Code, Grok-4.6 also landed #6 in Confirmed Success, with +15%! When looking at Agent Arena overall, its clearest strength by signal is Confirmed Success (an explicit “yes, that worked” from users), with +13.2% net improvement overall (#6). This is a significant jump from Grok-4.5 with +5% (#18). With a median cost per task of $1.12, it is on par with Claude Opus 4.6 at $1.19, but doesn’t quite make the Pareto Frontier based on its performance. Grok-4.6 ranks #15 overall, with +6.1% net improvement. By signal overall: - #6 Confirmed Success (+13.2%) - #23 Praise vs. Complaint (+2.6%) - #19 Steerability (+4.9%) - #22 Bash Recovery (+8.8%) - No issues with Tool Hallucination (+1.0%) It also landed #12 in Chat (+5.0%) and #17 in Work (+6.1%). Congrats to SpaceXAI on this release!

  • View organization page for Arena

    26,717 followers

    Big news: Gemini Omni 1.1 Flash has landed #1 in the Text-to-Video Arena and #2 in the Image-to-Video Arena! For Text-to-Video the latest 1.1 model is +20pts above FLUX 3 Video at #3 (1495 pts). For Image-to-Video the release is a strong +25pt improvement from Gemini Omni Flash at #5, and is only behind MiniMax-H3 in the #1 spot (1494 pts). Congrats to the Google team on this release!

  • View organization page for Arena

    26,717 followers

    Coding in Agent Mode with GitHub is now available on Arena! We’ve overhauled our Agent Mode architecture to support deep repository integration. By integrating directly with GitHub, Agent Mode lets our users connect a repo, complete a coding task, and push changes without leaving their browser. AI coding assistants have changed how code is created, but for many developers, chat has still been a “read-only” or “copy-paste” experience. Until now. New coding features include: - GitHub OAuth connector: Securely authenticate and list your repositories directly within the workspace. - Sandbox-based cloning: We now clone your repo into a sandbox at the start of a session, giving the agent a clean, isolated environment to explore and edit your code. - Diff review panel: Forget hunting through text outputs. Our new dedicated diff panel lets you watch the agent work in real-time, toggle between "last turn" and "full branch" diffs, and inspect syntax-highlighted changes. - Git lifecycle support: From cloning to commit, push, and PR creation, the agent handles the heavy lifting, allowing you to focus on review and approval.

  • View organization page for Arena

    26,717 followers

    Exciting news: Qwen3.8-Flash-Next by Qwen ranks ~#8 in Code Arena: WebDev (#3 among open models) scoring 1617 (AutoEval). Priced at $0.16/$0.47 Mtokens, it reshapes the Code Arena: WebDev Pareto Frontier! This 125B parameter model (6B active) has been touted as a preview of the Qwen4 architecture, showing a +22pt stronger performance than Qwen3.8-27B. Note: this is an early AutoEval score, in which a Reward Model trained on Arena's human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. See thread for more info on the methodology behind AutoEval. Congrats to the Qwen team on the strong release!

Similar pages