Papers

Paper: AREX: Towards a Recursively Self-Improving Agent for Deep Research

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Deep research is challenging because finding potential solutions often takes significant effort, while checking whether those solutions meet all the required constraints (multiple criteria) can be broken down into smaller, more manageable steps. This “discovery-verification asymmetry” creates a bottleneck: simply searching for longer doesn’t necessarily lead to better results.

Method

The paper introduces AREX, a family of “Recursively Self-Improving” (RSI) deep research agents designed to address this challenge. AREX operates with an alternating two-loop structure:

Paper: SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Training massive, trillion-parameter Mixture of Experts (MoE) language models like DeepSeek-V4 presents significant engineering challenges when using distributed training systems. The paper highlights issues including intense memory usage, communication bottlenecks, and inefficient processing during the post-training phase—specifically, Full Parameter Post-Training (CPT) and Supervised Fine Tuning (SFT). While most existing solutions rely on GPU clusters, this research explores an alternative approach leveraging Ascend Neural Processing Units (NPUs).

Paper: VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Current open-source video understanding models face several limitations. They often struggle to generalize across different types of videos, performing well only in specific niches. These models also tend to be computationally expensive and may not be fully accessible for researchers or developers, with key training details and datasets withheld.

Method

The paper introduces VideoChat3, a “fully open” video-centric Multimodal Large Language Model (MLLM) designed to overcome these limitations. The core approach combines two key elements:

Paper: Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Den...

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Current benchmarks used to evaluate AI agents often focus on simple tasks that complete quickly and are judged solely by their final outcome. This doesn’t give a full picture of an agent’s capabilities, especially when dealing with complex, real-world scenarios requiring sustained effort and iterative problem-solving. Existing “terminal” benchmarks (which judge only the end result) provide limited insight into intermediate progress and partial solutions due to sparse reward signals.

Paper: UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Evaluating proactive AI agents—those designed to operate tools and assist users in real-world environments like personal assistants or automated workflows—is currently difficult. Existing benchmarks often use simplified, sandboxed testing grounds and evaluate agents only on single interactions. Additionally, these benchmarks categorize tasks in ways that blur the lines between different underlying capabilities of the models, making it hard to pinpoint why an agent succeeds or fails.

Paper: AlayaWorld: Long-Horizon and Playable Video World Generation

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Creating compelling game worlds and virtual environments is traditionally a resource-intensive process. Building these worlds requires significant manual effort, making customization difficult and modifications after launch costly. This paper tackles the challenge of efficiently generating interactive virtual worlds without relying solely on manual authoring.

Method

The AlayaWorld framework proposes a new approach leveraging video world models. These models work by autoregressively synthesizing future observations – essentially predicting what will happen next in the virtual environment – based on the current state and user actions. The models are trained using gameplay recordings as well as real-world videos, allowing them to learn both visual styles and realistic physics simulations. AlayaWorld itself is presented as a full-stack open-source framework encompassing data preparation, model architecture design, training, inference acceleration, and deployment – all within a modular structure.

Paper: RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Robotic manipulation in real-world environments is challenging because robots need to understand not just what things look like, but also how those objects and the environment itself will move when interacted with. Current approaches relying solely on video data (2D pixel information) often fall short of providing this necessary understanding of 3D structure and movement dynamics.

Paper: OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Selecting an optimizer for training large-scale machine learning models has become surprisingly complex. With over one hundred available methods, researchers and engineers are facing a fragmented landscape. The choice isn’t just about performance; it’s a system-level design decision that must balance computational resources, memory constraints, the effort required for tuning, and the specific requirements of different tasks.

Paper: UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Training AI agents to interact with graphical user interfaces (GUIs) across different platforms (like desktop and mobile) has proven difficult. Existing datasets often lack comprehensive coverage of various platforms, and the varying interaction conventions between platforms can lead to AI agents mixing up behaviors or forgetting how to perform tasks they previously mastered on other platforms – a phenomenon known as catastrophic forgetting.

Paper: Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Listen to this article.

Audio is available for 30 days and will be removed automatically.

Problem

Deploying embodied AI models (think robots understanding and acting on their environment) is surprisingly difficult. Current solutions are often fragmented, relying on Python code specific to each model and the hardware they’re running on. This makes it hard to move these models between different robots, simulators, or even just various edge devices with varying capabilities.