NR

Nived Rajaraman

Postdoctoral Researcher at Microsoft Research

Previously EECS PhD Student at UC Berkeley

About

Nived Rajaraman’s Scaling Test-Time Compute identifies that increasing inference-time computation for LLMs is suboptimal without verification or RL. Rajaraman is a Postdoctoral Researcher in the Reinforcement Learning group at Microsoft Research. Their research investigates the statistical and computational theory of adaptive decision-making, including imitation learning and RL. Rajaraman also developed FastSecAgg, a scalable protocol for secure aggregation in privacy-preserving federated learning. They completed a PhD at UC Berkeley’s BAIR lab under Jiantao Jiao and Kannan Ramchandran. Rajaraman holds a dual degree from IIT Madras and previously interned at DeepMind and Microsoft Research.

Experience

Postdoctoral Researcher

2025 – Present

Member of the Reinforcement Learning group.

Research Intern

2022 – Present

Worked with Ravishankar Krishnaswamy.

Research Intern

2021 – Present

Deepmind

Worked with Nevena Lazic and Dong Yin.

EECS PhD Student

2019 – 2025

UC Berkeley · Berkeley, California

Advised by Jiantao Jiao and Kannan Ramchandran; affiliated with the BLISS and BAIR labs.

Papers31

Minimax Optimal Online Imitation Learning via Replay Estimation

Online imitation learning is the problem of how best to mimic expert demonstrations, given access to the environment or an accurate simulator. Prior work has shown that in the infinite sample regime, exact moment matching achieves value equivalence to the expert policy. However, in the finite sample regime, even if one has no optimization error, empirical variance can lead to a performance gap that scales with H2/NH^2 / N for behavioral cloning and $H / \sqrt{N}$ for online moment matching, where $H$ is the horizon and $N$ is the size of the expert dataset. We introduce the technique of replay estimation to reduce this empirical variance: by repeatedly executing cached expert actions in a stochastic simulator, we compute a smoother expert visitation distribution estimate to match. In the presence of general function approximation, we prove a meta theorem reducing the performance gap of our approach to the parameter estimation error for offline classification (i.e. learning the expert policy). In the tabular setting or with linear function approximation, our meta theorem shows that the performance gap incurred by our approach achieves the optimal O~(min⁡(H3/2/N,H/N)\widetilde{O} \left( \min({H^{3/2}} / {N}, {H} / {\sqrt{N}} \right) dependency, under significantly weaker assumptions compared to prior work. We implement multiple instantiations of our approach on several continuous control tasks and find that we are able to significantly improve policy performance across a variety of dataset sizes.

30 May 2022
16views30citations

Provably Breaking the Quadratic Error Compounding Barrier in Imitation Learning, Optimally

We study the statistical limits of Imitation Learning (IL) in episodic Markov Decision Processes (MDPs) with a state space S\mathcal{S}. We focus on the known-transition setting where the learner is provided a dataset of NN length-HH trajectories from a deterministic expert policy and knows the MDP transition. We establish an upper bound O(∣S∣H3/2/N)O(|\mathcal{S}|H^{3/2}/N) for the suboptimality using the Mimic-MD algorithm in Rajaraman et al (2020) which we prove to be computationally efficient. In contrast, we show the minimax suboptimality grows as Ω(H3/2/N)\Omega( H^{3/2}/N) when ∣S∣≥3|\mathcal{S}|\geq 3 while the unknown-transition setting suffers from a larger sharp rate Θ(∣S∣H2/N)\Theta(|\mathcal{S}|H^2/N) (Rajaraman et al (2020)). The lower bound is established by proving a two-way reduction between IL and the value estimation problem of the unknown expert policy under any given reward function, as well as building connections with linear functional estimation with subsampled observations. We further show that under the additional assumption that the expert is optimal for the true reward function, there exists an efficient algorithm, which we term as Mimic-Mixture, that provably achieves suboptimality O(1/N)O(1/N) for arbitrary 3-state MDPs with rewards only at the terminal layer. In contrast, no algorithm can achieve suboptimality O(H/N)O(\sqrt{H}/N) with high probability if the expert is not constrained to be optimal. Our work formally establishes the benefit of the expert optimal assumption in the known transition setting, while Rajaraman et al (2020) showed it does not help when transitions are unknown.

25 Feb 2021
18views13citations

Scaling Test-Time Compute Without Verification or RL is Suboptimal

Despite substantial advances in scaling test-time compute, an ongoing debate in the community is how it should be scaled up to enable continued and efficient improvements with scaling. There are largely two approaches: first, distilling successful search or thinking traces; and second, using verification (e.g., 0/1 outcome rewards, reward models, or verifiers) to guide reinforcement learning (RL) and search algorithms. In this paper, we prove that finetuning LLMs with verifier-based (VB) methods based on RL or search is far superior to verifier-free (VF) approaches based on distilling or cloning search traces, given a fixed amount of compute/data budget. Further, we show that as we scale test-time compute (measured as the output token length) and training data, suboptimality of VF methods scales poorly compared to VB when the base pre-trained LLM presents a heterogeneous distribution over correct solution traces (e.g., different lengths, styles, etc.) and admits a non-sharp distribution over rewards on traces sampled from it. We formalize this condition using anti-concentration [Erd\H{o}s, 1945]. This implies a stronger result that VB methods scale better asymptotically, with the performance gap between VB and VF methods widening as test-time budget grows. We corroborate our theory empirically on both didactic and math reasoning problems with 3/8/32B-sized pre-trained LLMs, where we find verification is crucial for scaling test-time compute.

17 Feb 2025
3kviews

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

Verifier-guided autocurricula focus costly reasoning supervision on a model’s mistakes, sharply reducing teacher demonstrations and making reference-model coverage a burn-in cost.

18 Mar 2026
287views

What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains

In-context learning (ICL) is a hallmark capability of transformers, through which trained models learn to adapt to new tasks by leveraging information from the input context. Prior work has shown that ICL emerges in transformers due to the presence of special circuits called induction heads. Given the equivalence between induction heads and conditional k-grams, a recent line of work modeling sequential inputs as Markov processes has revealed the fundamental impact of model depth on its ICL capabilities: while a two-layer transformer can efficiently represent a conditional 1-gram model, its single-layer counterpart cannot solve the task unless it is exponentially large. However, for higher order Markov sources, the best known constructions require at least three layers (each with a single attention head) - leaving open the question: can a two-layer single-head transformer represent any kth-order Markov process? In this paper, we precisely address this and theoretically show that a two-layer transformer with one head per layer can indeed represent any conditional k-gram. Thus, our result provides the tightest known characterization of the interplay between transformer depth and Markov order for ICL. Building on this, we further analyze the learning dynamics of our two-layer construction, focusing on a simplified variant for first-order Markov chains, illustrating how effective in-context representations emerge during training. Together, these results deepen our current understanding of transformer-based ICL and illustrate how even shallow architectures can surprisingly exhibit strong ICL capabilities on structured sequence modeling tasks.

10 Aug 2025
179views

The Price of Hidden Curvature: Improved Lower Bounds for Bandit Convex Optimization

We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for 11-Lipschitz functions on the dd-dimensional Euclidean ball. For time horizons n≥d10/3n\ge d^{10/3}, we prove a lower bound of Ω(d4/3n)\Omega(d^{4/3}\sqrt{n}), the first nontrivial bound that exceeds the dnd\sqrt{n} dependence of linear bandits, showing that stochastic bandit convex optimization is fundamentally harder than linear bandits. For d2≤n≤d10/3d^2\le n\le d^{10/3}, we obtain a lower bound of Ω(dn3/4)\Omega(\sqrt{d}n^{3/4}), matching the regret of the algorithm of Flaxman et al. (2005), establishing its optimality in this regime. The hard class of convex functions we construct takes the following form in dimension 2d2d: for an action a=(a1,a2)∈B2da=(a^1,a^2)\in \mathbb{B}^{2d}, each function is the scaled soft maximum of a "tube", r−1∥W⋆a1−r8εa2∥r^{-1}\|W^\star a^1-\frac{r}{8\varepsilon}a^2 \| (hyperparameterized by ε,r\varepsilon,r), and a squared distance function, 12∥a1−u⋆∥2−12∥u⋆∥2\frac12\|a^1-u^\star\|^2-\frac12\|u^\star\|^2. Here u⋆∈Rdu^\star\in\mathbb{R}^d is the unknown target determining the minimizer, while W⋆∈Rd×dW^\star\in\mathbb{R}^{d\times d} hides the region in which the quadratic curvature is observable. Indeed, observations reveal substantial information about u⋆u^\star only when the learner acts near the hidden tube a2≈8εrW⋆a1a^2\approx \frac{8\varepsilon}{r}W^\star a^1; away from it, the tube branch masks the quadratic branch. Thus the learner must pay to uncover the geometry encoded by W⋆W^\star before it can effectively exploit the curvature that identifies u⋆u^\star. Formalizing this tradeoff yields a sample complexity lower bound of Ω(d5/2ε2∧d2ε4)\Omega(\frac{d^{5/2}}{\varepsilon^2}\wedge\frac{d^2}{\varepsilon^4}) for finding an ε\varepsilon-optimal action, and ultimately the Ω(d4/3n∧dn3/4)\Omega(d^{4/3}\sqrt{n}\wedge\sqrt{d}n^{3/4}) regret lower bound. The proof was developed by GPT-5.5 Pro and GPT-5.6 Sol Pro under the authors' guidance.

21 Jul 2026
55views

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning

Reinforcement learning has rapidly emerged as a key component in the training of reasoning and coding models, yet it remains poorly understood from a mechanistic perspective. We study how and through what underlying processes capabilities are acquired or enhanced via reinforcement learning post-training. Our analysis, based on controlled math reasoning experiments with Qwen-2.5-1.5B, reveals two core mechanisms: strategy selection and strategy improvement. Our results highlight the role of SFT data and reinforcement learning data in activating these mechanisms, in particular showing how supervising the model on diverse reasoning strategies can enable strategy selection and how increasing difficulty in reinforcement learning data can enable strategy improvement. Taken together, our results provide mechanistic insight into RL training and suggest practical interventions to continue scaling reasoning capabilities.

11 Jun 2026
86views

From Markov to Laplace: How Mamba In-Context Learns Markov Chains

While transformer-based language models have driven the AI revolution thus far, their computational complexity has spurred growing interest in viable alternatives, such as structured state space sequence models (SSMs) and Selective SSMs. Among these, Mamba (S6) and its variant Mamba-2 have shown remarkable inference speed-ups over transformers while achieving comparable or superior performance on complex language modeling tasks. However, despite these architectural innovations and empirical successes, the fundamental learning capabilities of Mamba remain poorly understood. In this paper, we address this gap by studying in-context learning (ICL) on Markov chains and uncovering an interesting phenomenon: even a single-layer Mamba efficiently learns the in-context Laplacian smoothing estimator, which is both Bayes and minimax optimal. To explain this, we theoretically characterize the representation capacity of Mamba and reveal the fundamental role of convolution in enabling it to represent the optimal Laplacian smoothing. These theoretical insights align strongly with empirical results and, to the best of our knowledge, represent the first formal connection between Mamba and optimal statistical estimators. Finally, we outline promising research directions inspired by these findings.

14 Feb 2025
133views

Toward the Fundamental Limits of Imitation Learning

Imitation learning (IL) aims to mimic the behavior of an expert policy in a sequential decision-making problem given only demonstrations. In this paper, we focus on understanding the minimax statistical limits of IL in episodic Markov Decision Processes (MDPs). We first consider the setting where the learner is provided a dataset of NN expert trajectories ahead of time, and cannot interact with the MDP. Here, we show that the policy which mimics the expert whenever possible is in expectation $\lesssim \frac{|\mathcal{S}| H^2 \log (N)}{N}$ suboptimal compared to the value of the expert, even when the expert follows an arbitrary stochastic policy. Here S\mathcal{S} is the state space, and HH is the length of the episode. Furthermore, we establish a suboptimality lower bound of ≳∣S∣H2/N\gtrsim |\mathcal{S}| H^2 / N which applies even if the expert is constrained to be deterministic, or if the learner is allowed to actively query the expert at visited states while interacting with the MDP for NN episodes. To our knowledge, this is the first algorithm with suboptimality having no dependence on the number of actions, under no additional assumptions. We then propose a novel algorithm based on minimum-distance functionals in the setting where the transition model is given and the expert is deterministic. The algorithm is suboptimal by $\lesssim \min { H \sqrt{|\mathcal{S}| / N} ,\ |\mathcal{S}| H^{3/2} / N }$, showing that knowledge of transition improves the minimax rate by at least a H\sqrt{H} factor.

13 Sept 2020
110views

Transformers on Markov Data: Constant Depth Suffices

Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities. In this paper, we study the behavior of transformers on data drawn from \kth Markov processes, where the conditional distribution of the next symbol in a sequence depends on the previous kk symbols observed. We observe a surprising phenomenon empirically which contradicts previous findings: when trained for sufficiently long, a transformer with a fixed depth and 11 head per layer is able to achieve low test loss on sequences drawn from \kth Markov sources, even as kk grows. Furthermore, this low test loss is achieved by the transformer's ability to represent and learn the in-context conditional empirical distribution. On the theoretical side, our main result is that a transformer with a single head and three layers can represent the in-context conditional empirical distribution for \kth Markov sources, concurring with our empirical observations. Along the way, we prove that attention-only transformers with O(log⁡2(k))O(\log_2(k)) layers can represent the in-context conditional empirical distribution by composing induction heads to track the previous kk symbols in the sequence. These results provide more insight into our current understanding of the mechanisms by which transformers learn to capture context, by understanding their behavior on Markov sources.

25 Jul 2024
69views