Ruishuo Chen

AI MSc Student, IIIS at Tsinghua University

Previously Statistics BSc Student, School of Mathematics at Nanjing University

About

Ruishuo Chen's ResearchStudio-Reel automates the conversion of research papers into posters, videos, and blogs. They are an MSc student at the Institute for Interdisciplinary Information Sciences at Tsinghua University supervised by Longbo Huang. Chen focuses on RL and its applications to generative modeling and LLM agents. Their other work includes PowerFlow, which utilizes principled distribution matching for LLMs, and Trajectory-Distilled Guidance for training offline GFlowNets. They have also developed efficient algorithms for adversarial imitation learning and studied optimal skill selection for LLM agents. Chen previously earned a BSc in Statistics from Nanjing University.

Experience

AI MSc Student, IIIS

Sep 2025 – Present

Tsinghua University · Beijing, China

Works under the supervision of Professor Longbo Huang.

Algorithm Intern, Research Department

Jul 2025 – Sep 2025

X-square Investment · Shanghai, China

Independently proposed and implemented an algorithm for formulaic alpha factor mining.

Statistics BSc Student, School of Mathematics

Sep 2021 – Jun 2025

Nanjing University · Nanjing, China

Papers11

Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation

As a prominent category of imitation learning methods, adversarial imitation learning (AIL) has garnered significant practical success powered by neural network approximation. However, existing theoretical studies on AIL are primarily limited to simplified scenarios such as tabular and linear function approximation and involve complex algorithmic designs that hinder practical implementation, highlighting a gap between theory and practice. In this paper, we explore the theoretical underpinnings of online AIL with general function approximation. We introduce a new method called optimization-based AIL (OPT-AIL), which centers on performing online optimization for reward functions and optimism-regularized Bellman error minimization for Q-value functions. Theoretically, we prove that OPT-AIL achieves polynomial expert sample complexity and interaction complexity for learning near-expert policies. To our best knowledge, OPT-AIL is the first provably efficient AIL method with general function approximation. Practically, OPT-AIL only requires the approximate optimization of two objectives, thereby facilitating practical implementation. Empirical studies demonstrate that OPT-AIL outperforms previous state-of-the-art deep AIL methods in several challenging tasks.

01 Nov 2024
61views5citations

Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training

Generative Flow Networks (GFlowNets) excel at sampling diverse, high-reward objects. In many practical applications where active reward queries are infeasible, these models must be trained using static offline datasets. Prevailing training methods typically rely on a proxy model to provide reward feedback for online sampled trajectories. However, constructing a reliable proxy is often challenging due to data scarcity or high evaluation costs. While existing proxy-free approaches attempt to address this, they often impose coarse constraints that limit the model's ability to explore effectively. To overcome these limitations, we propose Trajectory-Distilled GFlowNet (TD-GFN), a novel proxy-free training framework. TD-GFN utilizes inverse reinforcement learning (IRL) to extract dense, transition-level edge rewards from offline trajectories, providing rich structural guidance for efficient exploration. Crucially, to ensure robustness, these rewards guide the policy indirectly through DAG pruning and prioritized backward sampling. This design ensures that gradient updates rely exclusively on ground-truth terminal rewards from the dataset, thereby preventing error propagation. Empirical results demonstrate that TD-GFN significantly outperforms a broad range of existing baselines in both convergence speed and sample quality, establishing a more robust and efficient paradigm for offline GFlowNet training.

26 May 2025
263views2citations

Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling

Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live streaming, and data center management, where efficient resource allocation is required among users with diverse delay sensitivities. In these scenarios, schedulers must make real-time decisions to satisfy both delay and resource constraints without prior knowledge of system dynamics, which are often time-varying and challenging to estimate. {Current learning-based methods typically require online interactions with actual systems during the training stage. Therefore, these approaches are often difficult or impractical, as they can significantly degrade system performance and incur substantial service costs.} To address these challenges, we propose a novel offline reinforcement learning-based algorithm, named Scheduling By Offline Learning with Critic Guidance and Diffusion Model (SOCD), to learn efficient scheduling policies purely from pre-collected offline data. SOCD innovatively employs a diffusion policy, complemented by a sampling-free critic network for policy guidance. By integrating the Lagrangian multiplier optimization into the offline reinforcement learning, SOCD efficiently trains high-quality constraint-aware policies exclusively from available datasets, eliminating the need for online interactions with the system. Experimental results demonstrate that SOCD is resilient to various system dynamics, including partially observable and large-scale environments, and delivers superior performance compared to existing methods.

22 Jan 2025
76views2citations

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multiple dissemination formats, but a practical workflow must also keep the outputs editable in native tools and bound into one navigable deliverable for revision and reuse. We present ResearchStudio-Reel, a native-editable dissemination workspace that binds its three artifacts into one interactive deliverable at the experience level, implemented as five skills executable in Claude Code and Codex: one shared extractor, three editable artifact generators, and one interactive convergence layer. A shared asset bundle feeds a PowerPoint poster and video deck, plus a bilingual Word blog; rather than re-rendering the paper into a fourth format, Paper2Reel converges these already-produced artifacts at the experience level, binding poster regions, video segments, and blog passages into one interactive viewer. Artifact-specific release checks make this delivery contract testable, and Paper2Poster additionally uses a measured-fill loop. On the Paper2Poster benchmark, our Claude Code configuration achieves the best scores among automated systems on all three aesthetic sub-criteria and the best or tied-best scores on two of three information sub-criteria. Under two VLMjudges, it exceeds the authors' posters in average aesthetics (3.56 vs. 3.03) and wins on overall quality on 74 and 95 of the 100 papers under the two judges. The full pipeline additionally packages the native-editable source artifacts and their aligned viewer. Project is available at this https URL

05 Jul 2026
435views1citations

PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching

Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theoretical optimization target and are prone to degenerative biases. In this work, we introduce PowerFlow, a principled framework that reformulates unsupervised fine-tuning as a distribution matching problem. By casting GFlowNet as an amortized variational sampler for unnormalized densities, we propose a length-aware Trajectory-Balance objective that explicitly neutralizes the structural length biases inherent in autoregressive generation. By targeting α\alpha-power distributions, PowerFlow enables the directional elicitation of the dual nature of LLMs: sharpening the distribution (\alpha > 1) to intensify logical reasoning, or flattening it (\alpha < 1) to unlock expressive creativity. Extensive experiments demonstrate that PowerFlow consistently outperforms existing RLIF methods, matching or even exceeding supervised GRPO. Furthermore, by mitigating over-sharpening in aligned models, our approach achieves simultaneous gains in diversity and quality, shifting the Pareto frontier in creative tasks.

19 Mar 2026
279views1citations

When Context Returns: Toward Robust Internalization in On-Policy Distillation

Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so that the context is no longer needed at inference time. However, we identify a counterintuitive and previously unstudied phenomenon: reintroducing the original privileged context to the distilled student often degrades its performance, even on instances it already solves correctly without context. We term this phenomenon context-induced degradation and argue that robust internalization requires not only matching the teacher's context-conditioned behavior, but also remaining stable when the privileged context is reintroduced, a desirable property we call context invariance. To promote this property, we formulate a novel view-robust internalization risk and propose No-Context Anchoring (NCA), a lightweight yet effective consistency regularizer that uses the student's stop-gradient no-context output as an anchor and aligns its context-conditioned output via forward KL divergence. Across 14 configurations spanning diverse domains and model families, NCA improves context-conditioned accuracy in most settings and reduces context harm in 12 out of 14, while preserving or improving no-context performance, demonstrating greater robustness to context reintroduction.

10 Jun 2026
149views1citations

The Router Within: Eliciting Native Skill Routing from a Frozen LLM

Ruishuo ChenRuishuo ChenXun WangYu ChenYu ChenZhuoran Li+1

An agent’s hidden states can identify useful skills without loading skill descriptions into its context or relying on a separate retrieval model.

14 Sept 2026
220views

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

An execution-trained selector chooses complementary skills rather than merely relevant ones, improving coding-agent success while using fewer context tokens on a controlled testbed.

20 Aug 2026
217views

MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and revise reward functions as monolithic programs, making it difficult to reliably preserve and reuse effective components discovered in earlier iterations, leading to unstable performance across iterations. To address this, we propose Module Level Reward Evolution Framework (MLREF). At the core of MLREF is a module pool, a persistent repository of reusable reward components. MLREF treats the module pool as the primary optimization object: the pool evolves across iterations by accumulating successful modules, refining underperforming ones, and reusing proven components; while reward functions are constructed as linear combinations of modules drawn from this pool. To drive this evolution, MLREF integrates three mechanisms: reflection-based refinement, hybrid credit assignment, and a merge strategy with rollback, which together improve the effectiveness and robustness of reward optimization. Experiments on 17 tasks show that MLREF outperforms strong baselines by 25.2% in locomotion and 6.6% in manipulation, with more stable optimization dynamics.

19 Aug 2026
37views

Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms

Tian XuTian XuZhilong ZhangZexuan ChenRuishuo ChenRuishuo Chen+2

Adversarial imitation learning (AIL), a prominent approach in imitation learning, has achieved significant practical success powered by neural network approximation. However, existing theoretical analyses of AIL are primarily confined to simplified settings, such as tabular and linear function approximation, and involve complex algorithmic designs that impede practical implementation. This creates a substantial gap between theory and practice. This paper bridges this gap by exploring the theoretical underpinnings of online AIL with general function approximation. We introduce a novel framework called optimization-based AIL (OPT-AIL), which performs online optimization for reward learning coupled with optimism-regularized optimization for policy learning. Within this framework, we develop two concrete methods: model-free OPT-AIL and model-based OPT-AIL. Our theoretical analysis demonstrates that both variants achieve polynomial expert sample complexity and interaction complexity for learning near-expert policies. To the best of our knowledge, they represent the first provably efficient AIL methods under general function approximation. From a practical standpoint, OPT-AIL requires only the approximate optimization of two objectives, thereby facilitating practical implementation. Empirical studies demonstrate that OPT-AIL outperforms previous state-of-the-art deep AIL methods across several challenging tasks.

03 May 2026
16views