DeepSeek V4 Flash 0731 is now open weights! DeepSeek has just released the weights for its new flash tier model, DeepSeek V4 Flash 0731. With a score of 50 on the Artificial Analysis Intelligence Index, it lands among the top 3 open weights models on the leaderboard. The weights are released under the MIT license, allowing unrestricted commercial use and modification. DeepSeek V4 Flash 0731 shares identical architecture and pricing with the earlier DeepSeek V4 Flash. At a size of 284B total parameters (13B active), released in mixed FP4/FP8 precision at ~167GB total file size, it lands on our Pareto frontier for Intelligence Index vs. Total Parameters. Among open weights models, DeepSeek V4 Flash 0731 delivers a significant leap in intelligence for its size class. DeepSeek V4 Flash 0731 is also available now through DeepSeek's first-party API. Check out Artificial Analysis to compare DeepSeek V4 Flash 0731 with other leading open weights and proprietary models: https://lnkd.in/g4bbqEre
Artificial Analysis
Technology, Information and Internet
Newark, Delaware 32,936 followers
Independent analysis of AI: Understand the AI landscape and analyze AI technologies http://artificialanalysis.com/
About us
Leading independent analysis of AI. Backed by Nat Friedman, Daniel Gross and Andrew Ng.
- Website
-
https://artificialanalysis.ai
External link for Artificial Analysis
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Headquarters
- Newark, Delaware
- Type
- Privately Held
Employees at Artificial Analysis
Locations
-
Get directions
131 Continental Dr
Suite 305
Newark, Delaware 19713, US
Updates
-
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, a 10-point jump over DeepSeek V4 Flash (released April 2026) that puts it 6 points ahead of DeepSeek V4 Pro. It shares identical architecture and pricing with the earlier DeepSeek V4 Flash, and lands on our Pareto frontier for Intelligence vs Cost per Task DeepSeek V4 Flash 0731 is one Intelligence Index point behind GPT-5.6 Luna (max, 51). Even after OpenAI’s 80% price cut on GPT-5.6 Luna today, DeepSeek V4 Flash 0731’s Cost per Task on DeepSeek’s first-party API comes in at ~60% lower than GPT-5.6 Luna (max), a model with comparable intelligence. A key driver of this is DeepSeek’s ~98% cache hit discount on its first-party API, a significantly more aggressive discount than the 90% cache hit discount offered by most of the industry The new model is a significant step up from the previous generation, DeepSeek V4 Flash (40), and places the model within 1 point of GLM-5.2 (max, 51). It remains 7 points behind the open weights frontier set by Kimi K3 (max, 57). For additional context, this places the model in line with recently released Gemini 3.6 Flash (50) and 1 point behind Muse Spark 1.1 (xhigh, 51). DeepSeek is expected to release the model’s full weights in the coming weeks DeepSeek V4 Flash 0731 retains a 1M token context window, and its size remains unchanged from DeepSeek V4 Flash at 284B total parameters and 13B active at inference time Key results: ➤ Improvements in agentic performance: DeepSeek V4 Flash 0731 achieves an Elo rating of 1559 on GDPval-AA v2, our evaluation focused on agentic real-world work tasks, up from 1189 for the previous DeepSeek V4 Flash. Once weights are released this will be the second highest open weights score, behind Kimi K3 (max, 1687) and ahead of GLM-5.2 (max, 1510). ➤ Token usage falls 12% against the predecessor: DeepSeek V4 Flash 0731 used ~206M output tokens to run the Intelligence Index, against ~234M for the previous DeepSeek V4 Flash. The new variant is more token efficient, achieving a higher Intelligence Index with a lower number of total output tokens ➤ DeepSeek V4 Flash 0731 improves over its predecessor on every evaluation in the Intelligence Index: Alongside the agentic gains, CritPt gains 9 points to 17%, SciCode 5 points to 50%, Humanity's Last Exam 5 points to 37%, AA-LCR 3 points to 66% and GPQA Diamond 1 point to 91% ➤ AA-Omniscience improvements are driven by fewer hallucinations, rather than higher accuracy: DeepSeek V4 Flash 0731 achieves an AA-Omniscience Index of -16, a +7 improvement from its predecessor. Additional model details: ➤ Context window: 1M tokens (equivalent to DeepSeek V4 Flash) ➤ Size: 284B total parameters (13B active) ➤ Input modalities: Text input and output only ➤ Accessibility: Available through DeepSeek’s first-party API ➤ Pricing: $0.14/$0.28 per 1M input/output tokens, unchanged from DeepSeek V4 Flash. Cache hit price of $0.0028 per 1M tokens
-
-
MiniMax H3 is #1 in Video Editing and ranks top 3 in both Text to Video and Image to Video on Artificial Analysis. MiniMax plans to release its weights under the MiniMax Community License, which would make it the leading open weights model by far MiniMax H3 is the latest video model from MiniMax and the successor to their Hailuo family. The model has a multimodal architecture, accepting text, images, videos, and audio clips as input to generate 5-15s clips at 24fps with native audio generation. Video input support also makes MiniMax H3 one of the few models capable of instruction-based video editing, where it takes the #1 spot in our Video Editing Leaderboard. The model is also #2 in Text to Video, and #3 in Image to Video. MiniMax H3 is available now in the Hailuo AI app and via the MiniMax API. MiniMax lists MiniMax H3 at $0.13 per second of 2K video ($7.80 per minute), with a 768p tier at $0.09 per second ($5.40 per minute) marked as coming soon in their documentation. The first five reference images are free, with each additional image at $0.04 and reference audio free. Our evaluation was conducted on the 2K tier. At $7.80 per minute with audio, MiniMax H3 at 2K ($7.80/min) undercuts Dreamina Seedance 2.0 1080p ($22.45/min), HappyHorse-1.1 ($9.90/min), and Kling 3.0 1080p ($20.16/min), but is still at a premium over Google’s Gemini Omni Flash at $6.00/min. MiniMax plans to release the weights under the MiniMax Community License, which permits commercial use for organizations under $20M in revenue, subject to prominent attribution requirements. If released, MiniMax H3 would become the strongest open weights video model, well ahead of the previous open weights leader LTX-2.3.
-
-
What does it take to run Kimi K3 locally? (it unfortunately isn’t going to run on your MacBook) Kimi K3 requires ~1.56 TB of memory for its weights alone, and the only current single-node systems that meet this bar (while supporting 4-bit precision) are NVIDIA’s B300 and AMD’s MI350X/MI355X. Unlike most earlier open-weights models, Kimi K3 is too large to run on any single Hopper or B200 node. Key takeaways when running Kimi K3 locally: ➤ It has 2.8 trillion total parameters, making it the largest open-weights model ever released ➤ Its hybrid attention mechanism reduces the amount of KV cache space required per token. For example, at 256K context on a GB300 NVL72, Kimi K3 could theoretically support over 2,000 concurrent requests using a 16-bit KV cache, compared to ~1,000 for Kimi K2.6 (when disregarding the throughput roofline) ➤ In total, running Kimi K3 locally while using its full 1M context window would require ~1.59 TB of memory for one user, growing by ~30 GB per additional concurrent user if using 16-bit KV cache ➤ The high memory requirement leads to a total system cost in the hundreds of thousands of dollars, highly dependent on specific hardware configuration and volume-based pricing ➤ The above points do not consider serving speed, which adds an additional layer of complexity - it is considerably more difficult to serve at a reasonable speed at high concurrency, vs. simply fitting the model inside available memory Realistic hardware configurations that can both fit Kimi K3’s weights and leave room for KV cache include: ➤ NVIDIA: 8x B300 or GB300; 16x B200 or GB200 (GB chips as part of an NVL72 rack) ➤ AMD: 8x MI350X or MI355X We will be publishing day-1 inference benchmarking results for Kimi K3 on a subset of capable hardware shortly, and will track how serving performance improves as inference stacks mature.
-
-
Agnes AI has launched Agnes 2.5 Pro Alpha, a low-priced reasoning model reaching 39 on the Artificial Analysis Intelligence Index at $0.45/$0.90 per 1M tokens Agnes AI is a Singapore-based AI lab that trains its own full-modality foundation models in-house across text, image, and video, and offers them through a free omni-modal API that has passed 3 million users. Agnes 2.5 Pro Alpha is a text, image, and video input reasoning model with text output. Agnes 2.5 Pro Alpha scores 39 on the Artificial Analysis Intelligence Index, placing it mid-pack among reasoning models we have tested and below the current proprietary frontier. Its defining characteristic is price: at $0.45 per 1M input tokens and $0.90 per 1M output tokens, Agnes 2.5 Pro Alpha is among the lowest-priced proprietary models at its intelligence level. Key results: ➤ Agnes 2.5 Pro Alpha debuts at 39 on the Artificial Analysis Intelligence Index, its first appearance on our leaderboard. This places it mid-pack among current reasoning models. ➤ Coding is a relative strength for Agnes 2.5 Pro Alpha, scoring 58.8 on the Coding Index, near the top for its intelligence tier. This is above similarly priced models including DeepSeek V4 Flash (56.2), GPT-5.4 mini (56.1) and Qwen3.7 Plus (55.9), and sits just behind Nex-N2-Pro (59.1). ➤ On academic reasoning, Agnes 2.5 Pro Alpha is competitive for its tier, scoring 32% on Humanity's Last Exam and 88% on GPQA Diamond. Its Humanity's Last Exam result leads similarly priced proprietary models including GPT-5.4 mini (xhigh, 27%) and GPT-5.4 nano (xhigh, 26%), and its GPQA Diamond score is in line with peers near its Intelligence Index. ➤ At $0.45/$0.90 per 1M input/output tokens, Agnes 2.5 Pro Alpha is the second cheapest model at its intelligence level. Blended at 7:2:1 (cache-input-output) it costs $0.18 per 1M tokens, behind only DeepSeek V4 Flash (Reasoning, Max Effort) at $0.06 among models scoring 38-40 on the Artificial Analysis Intelligence Index. Open weights alternatives including DeepSeek V4 Pro (Reasoning, Max Effort) and MiMo-V2.5-Pro match that $0.18 price at higher intelligence. Additional model details: ➤ Context window: 1M tokens. ➤ Pricing: $0.45 / $0.90 / $0.0038 per 1M input / output / cache hit tokens. ➤ Input modalities: Text, image, and video. ➤ Availability: Agnes AI first-party API. Read the full article at: https://lnkd.in/gXfGjvDr
-
-
SpaceXAI has released Grok Voice Think Fast 2.0 today, with the High reasoning variant debuting at #2 on the Artificial Analysis Speech to Speech Index at 82.9%, and #1 on Tau Voice for Agentic Performance at 56.5% - among the fastest models at 0.70s Time to First Audio Grok Voice Think Fast 2.0 is SpaceXAI's successor to Grok Voice Think Fast 1.0 (75.7% on the Speech to Speech Index). It is the only model in the Index's top five with an average Time to First Audio under 1 second, achieving 0.70 seconds vs. 1.14 seconds for the next fastest, GPT-Realtime-2 High. Key takeaways: ➤ Speech to Speech Index: Grok Voice Think Fast 2.0 High debuts at #2 at 82.9%, only behind Qwen Audio 3.0 Realtime Plus (84.1%) and ahead of GPT-Realtime-2.1 High (79.1%) and GPT-Realtime-2 High (77.2%). This is up 7.3 percentage points from Grok Voice Think Fast 1.0 (75.7%) ➤ Speech to Speech Index by Benchmark: On Tau Voice, Grok Voice Think Fast 2.0 High is the new leader at 56.5%, just ahead of Qwen Audio 3.0 Realtime Plus at 54.6% and its predecessor Grok Voice Think Fast 1.0 at 52.1%. On Big Bench Audio, it achieves 97.2%, behind leader Qwen Audio 3.0 Realtime Plus at 99.2%. On our Full Duplex Bench subset, it scores 95.1%, up from 77.8% for Grok Voice Think Fast 1.0, the largest driver of its index gain, behind Qwen Audio 3.0 Realtime Plus at 98.4% ➤ Speed: The model’s average Time to First Audio on Big Bench Audio is 0.70 seconds, faster than GPT-Realtime-2 High (1.14s), GPT-Realtime-2.1 High (1.21s), and Grok Voice Think Fast 1.0 (1.25s), and well ahead of Qwen Audio 3.0 Realtime Plus (4.02s) ➤ Price: Grok Voice Think Fast 2.0 is priced at $4.80 per hour of input audio, up from $3.00 for Grok Voice Think Fast 1.0 and more expensive than Qwen Audio 3.0 Realtime Plus ($4.42) and GPT-Realtime-2 High ($4.14), but ~2.2x cheaper than GPT-Realtime-2.1 High ($10.75)
-
-
OpenAI has released GPT Transcribe: a Speech to Text model scoring 3.31% on AA-WER (#9), improving 0.7 p.p. over its predecessor GPT-4o Transcribe while lowering price 25% to $4.50 per 1,000 minutes of audio GPT Transcribe is OpenAI's latest non-streaming (batch) speech transcription model, now accepting three kinds of context to improve transcription quality: a text prompt describing the recording's topic or setting, keywords for literal terms that may appear in the audio (such as product names or acronyms), and multiple language hints for multilingual and code-switching audio. The model processes audio at ~34× real-time and is available at $4.50 per 1,000 minutes of audio ($0.0045/min) via the OpenAI API Platform. OpenAI has also released GPT-Live-Transcribe, a streaming Speech to Text model. We are currently benchmarking this model and plan to share results on our Streaming Speech to Text leaderboard.
-
-
Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1 model on the Artificial Analysis Speech to Speech Index at 84.1%, ahead of GPT-Realtime-2.1 High at 79.1% Released earlier this month, Qwen Audio 3.0 Realtime is Alibaba's flagship native Speech to Speech model, available in two variants: Plus and Flash. Qwen Audio 3.0 Realtime Plus leads on all three component benchmarks comprising the Artificial Analysis Speech to Speech Index, Big Bench Audio for Speech Reasoning, Full Duplex Bench for Conversational Dynamics, and Tau Voice for Agentic Performance. We tested the China-hosted endpoints on Aliyun (Alibaba Cloud). Key takeaways: ➤ Speech to Speech Index: Qwen Audio 3.0 Realtime Plus is the new leader at 84.1%, ahead of GPT-Realtime-2.1 High (79.1%) and GPT-Realtime-2 High (77.2%). The Flash variant comes in at 4th at 76.3%. ➤ Speech to Speech Index by Benchmark: On Big Bench Audio, the Plus variant achieves 99.2%, up ~0.5 percentage points from the previous best of 98.7% (Alibaba Qwen3.5 Omni Plus Realtime), with Flash variant scoring 96.1%. On Tau Voice, the Plus variant currently leads with a score of 54.6%, ahead of Grok Voice Think Fast 1.0 at 52.1%. On Full Duplex Bench, the Plus variant leads our Full Duplex Bench subset at 98.4%, with Flash at 96.9%, both ahead of the best non-Alibaba model, GPT-Realtime-2 (Minimal) at 96.1%. ➤ Speed: The Plus variant records an average Time to First Audio of 4.02 seconds on Big Bench Audio, with Flash at 4.16 seconds, among the slowest models on our leaderboard, and well behind GPT-Realtime-2 (Minimal) at 1.10 seconds and GPT-Realtime-2 (High) at 1.14 seconds ➤ Price: Plus costs $4.42 per hour of input audio on our Big Bench Audio subset, more expensive than GPT-Realtime-2 High ($4.14) and ~2.4x cheaper than GPT-Realtime-2.1 High ($10.75). The average cost for Flash variant is $4.77, higher than the Plus variant despite lower list prices, driven by comparatively more verbose responses.
-
-
Kimi K3 weights have been released! Kimi K3 is now the leading open weights model at 57 in the Artificial Analysis Intelligence Index Moonshot has released the weights of their 2.6T parameter model under their 'Kimi K3 License' which we have labelled ‘Commercial Use Restricted’. Restrictions compared with more permissive licenses such as MIT or Apache 2.0 include requiring model-as-a-service businesses with more than $20 million in revenue to enter into a separate agreement. Commercial products with more than 100 million monthly active users or $20 million in monthly revenue also need to display “Kimi K3” in the user interface. Kimi (Moonshot AI) has also released their technical report with insights into model’s architecture and training approach. Links below to the weights on Hugging Face, the Technical Report and further benchmarks
-