Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 216 results for author: Gokul

  1. arXiv:2609.36561  [pdf, ps, other] 

    cs.RO

    Riemannian Splat Regression Models for Learning Time Fields on Arbitrary Riemannian Manifolds

    Authors: Anastasios Manganaris, Rohit Banerjee, Gokul Prabhakaran, Ahmed H. Qureshi

    Abstract: Motion planning on arbitrary Riemannian manifolds is an important and difficult problem that frustrates typical planning methods for Euclidean spaces. In particular, motion planning methods that approximate optimal time-to-go functions with neural networks, e.g., Neural Time Fields (NTFields), cannot be directly applied without using ad-hoc coordinate projections into higher dimensions. Using thes… ▽ More

    Submitted 28 September, 2026; originally announced September 2026.

    Comments: 5 pages, 2 figures, Spotlight Paper at IEEE/RSJ IROS Workshop on Geometric Representations in Robotics

  2. arXiv:2609.30715  [pdf, ps, other] 

    cs.RO

    RoboMonitor: Label-Efficient Runtime Monitoring of Robot Task Execution via Predictive Representation Learning

    Authors: Abhiroop Ajith, Gokul Narayanan, Kyle Coelho, Tingji Zhao, Yash Shahapurkar, Brian Zhu, Melih Erdogan, Ted Krubasik, Constantinos Chamzas, Eugen Solowjow

    Abstract: Learned robot policies produce actions, but their outputs alone do not establish whether execution is progressing as intended. Robot execution monitoring requires identifying the current execution phase, detecting failures, and recognizing task completion from observations available during execution. Training such monitors requires annotations that are scarce in datasets collected for robot-policy… ▽ More

    Submitted 24 September, 2026; originally announced September 2026.

  3. arXiv:2609.17533  [pdf, ps, other] 

    cs.SC cs.DS

    GPU-Accelerated Search for Fast Matrix Multiplication over $\mathbb{F}_2$

    Authors: Zeke Medley, Anagha Gokul, Quan Luu, Panagiotis Manolios

    Abstract: We present a GPU-accelerated algorithm for searching for ways to multiply matrices with few scalar multiplications. Our algorithm searches the matrix multiplication flip graph and, on an NVIDIA H200, performs over a billion search steps per second, a $1000\times$ improvement over previous GPU-accelerated search on the tensor for $7\times7$ matrix multiplication. Our second contribution is a proof… ▽ More

    Submitted 22 June, 2026; originally announced September 2026.

  4. You Shall Not Pass into Ring-0! A User Privacy-Friendly Anti-Cheat Architecture for Personal Computers

    Authors: Santosh Gokul Narayanan, Giovanni Paladino, Chuqi Zhang, Sangho Lee, Zhenkai Liang, Adil Ahmad

    Abstract: Kernel-level anti-cheats are effective against malicious player behavior in competitive video games, but raise significant user privacy concerns regarding installing unverifiable components at privileged modes (i.e., ring-0 in x86). While existing research has focused on improving the effectiveness of anti-cheats, the user privacy concern has been largely ignored. Tirith is an anti-cheat architect… ▽ More

    Submitted 15 September, 2026; originally announced September 2026.

    Comments: 15 pages, 8 figures, 2 tables. To appear in Proceedings of the 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26), November 15-19, 2026, The Hague, Netherlands

  5. arXiv:2609.02942  [pdf, ps, other] 

    cs.CL cs.AI

    Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation

    Authors: Anshul Bagaria, Sowmya S Sundaram, Gokul S Krishnan, Balaraman Ravindran

    Abstract: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only on rubric text, without access to any evaluated response, achieve nontrivial predictive performance on judge outputs. This suggest… ▽ More

    Submitted 31 August, 2026; originally announced September 2026.

    Comments: Accepted for publication at EMNLP 2026. 5 pages, 6 figures

    ACM Class: I.2.7

  6. arXiv:2609.01579  [pdf, ps, other] 

    cs.RO

    SG-AMP: Scene-Graph-Guided Active Perception and Semantics-Aware Motion Planning for Pepper Plants

    Authors: Rohit Menon, Shiva Rudra Lolla, Niklas Mueller-Goldingen, Gokul Chenchani, Ribana Roscher, Maren Bennewitz

    Abstract: We present SG-AMP, integrating robust depth completion with input-conditioned uncertainty, persistent panoptic mapping, plant scene-graph reasoning, and semantics-aware active view-motion planning. Beyond inspecting uncertain observed regions, the scene graph explicitly hypothesizes unobserved pepper--peduncle attachments and directs close-range sensing toward them. Candidate views are selected ac… ▽ More

    Submitted 1 September, 2026; originally announced September 2026.

  7. Location-Aware Language Models via Secondary Embeddings

    Authors: Gokul Srinivasagan, Munir Georges

    Abstract: Pretrained transformer-based language models achieve strong performance across a wide range of NLP tasks but remain limited in encoding geo-locational semantics, leading to suboptimal representations of place names and spatial entities. In this work, we propose a lightweight, model-agnostic approach for injecting geo-spatial awareness into pretrained embeddings without modifying the tokenizer or r… ▽ More

    Submitted 31 August, 2026; originally announced September 2026.

    Comments: Accepted for publication at the 29th International Conference on Text, Speech and Dialogue (TSD 2026)

  8. arXiv:2608.25061  [pdf, ps, other] 

    cs.CL cs.AI cs.DB cs.LG cs.PL

    DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

    Authors: Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie, Andrea Giovannini, Katja Hose

    Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that… ▽ More

    Submitted 27 August, 2026; v1 submitted 25 August, 2026; originally announced August 2026.

    Comments: Accepted at EMNLP 2026

  9. arXiv:2608.23831  [pdf, ps, other] 

    cs.RO cs.LG

    Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

    Authors: Brian Zhu, Momen Khalil, E Harrison, Emanuele Poggi, Philipp Schmitt, Bernd Kast, Philine Meister, Pranav Atreya, Qiyang Li, Finn Ferchau, Cesar Colmenero, Yash Shahapurkar, Gokul Narayanan, Melih Erdogan, Kai Wurm, Georg von Wichert, Oier Mees, Eugen Solowjow, Andrew Wagenmaker, Sergey Levine

    Abstract: While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamics and, if not correctly accoun… ▽ More

    Submitted 26 August, 2026; v1 submitted 24 August, 2026; originally announced August 2026.

    Comments: 25 pages, 12 figures, project website: https://async-rl-intermediate-information.github.io/

  10. arXiv:2608.21348  [pdf, ps, other] 

    cs.DS cs.GT cs.LG

    Truthful Calibration Measures for Sequential Prediction

    Authors: Anagha Gokul, Jason Hartline, Lunjia Hu, Jonathan Ullman, Yifan Wu

    Abstract: Calibration requires probabilistic reports to be conditionally unbiased and reliably interpretable as probabilities. A calibration measure assigns numerical error to miscalibrated reports. Haghtalab et al. (2024) proposed an approximately truthful calibration measure for online prediction, leaving open whether exact truthfulness is compatible with completeness and soundness. We resolve this ques… ▽ More

    Submitted 21 August, 2026; originally announced August 2026.

  11. arXiv:2608.20683  [pdf, ps, other] 

    cs.DB

    Improving Join Order Optimization on Gate-Based Quantum Computers via Structured Parameter Initialization

    Authors: Divya Shekar, Ruokun Wu, Dhanvi Bharadwaj, Gokul Subramanian Ravi, Lin Ma

    Abstract: Join Order Optimization (JOO) is one of the most computationally expensive tasks in relational query optimization due to the exponential growth of possible join plans with increasing query size. Recent work has explored quantum and quantum-inspired approaches for solving JOO by reformulating the problem as a Quadratic Unconstrained Binary Optimization (QUBO) problem suitable for optimization using… ▽ More

    Submitted 24 August, 2026; v1 submitted 20 August, 2026; originally announced August 2026.

    Comments: VLDB 2026 Workshop:QC&DKM 2026 - 2nd Workshop on Quantum Computing and Data/Knowledge Management

  12. arXiv:2608.05724  [pdf, ps, other] 

    cs.CL cs.LG

    Sparse PPMI Graph Averaging for Random Indexing Embeddings

    Authors: Sriram Loganathan, Gokul Anand, Aung Bo Bo, Yourui Shao, William B. Andreopoulos

    Abstract: We study a specific sparse post-processing pipeline for Random Indexing (RI) on kinship analogies in a small fairytales corpus. The published artifacts use uniform RI context accumulation with 200 dimensions and eight nonzeros, followed by one residual graph average, $\mathbf{E}=(1-α)\mathbf{E}_0+α\mathbf{P}\mathbf{E}_0$, where $\mathbf{P}$ is a row-normalized PPMI graph and $α=0.3$. Terminal row… ▽ More

    Submitted 21 August, 2026; v1 submitted 6 August, 2026; originally announced August 2026.

  13. arXiv:2607.24730  [pdf, ps, other] 

    cs.CV cs.AI

    KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

    Authors: Krithi Shailya, Ananya Lakshmi Ravi, Venkatanathan K. V., Sowmya S. Sundaram, Gokul S. Krishnan, Aditi Anand, Balaraman Ravindran

    Abstract: Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician trust. In clinical workflows, chest X-ray classifiers are increasingly paired with Vision-Language Models (VLMs) to generate natural-language explanations. However, these systems add linguistic fluency without addressing the underlying opacity of the visual mode… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

    Comments: MICCAI 2026

  14. arXiv:2607.01077  [pdf, ps, other] 

    cs.CL cs.LG

    Message Passing Enables Efficient Reasoning

    Authors: Xuecheng Liu, Daman Arora, Gokul Swamy, Andrea Zanette

    Abstract: While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck. Thus, in contrast to sequential scaling methods like CoT, recent parallel scaling techniques instead use fork and join (FJ) primitives to divide work across multiple LLM threads. However, in the fork-join paradigm, thre… ▽ More

    Submitted 8 August, 2026; v1 submitted 1 July, 2026; originally announced July 2026.

    Comments: COLM 2026

  15. arXiv:2606.31132  [pdf, ps, other] 

    cs.RO

    ELASTIC: Efficiently Learning to Adaptively Scale Test-Time Compute for Generative Control Policies

    Authors: Andrew Zou Li, Gokul Swamy, Yonatan Bisk, Andrea Bajcsy

    Abstract: Generative control policies (GCPs), such as diffusion policies and flow-based vision-language-action models, enable test-time scaling in robot control. Test-time compute can be allocated along two axes: sequential scaling, which increases denoising steps to refine actions, and parallel scaling, which samples multiple candidate actions to search across modes of the policy distribution. However, the… ▽ More

    Submitted 30 June, 2026; originally announced June 2026.

  16. arXiv:2606.30805  [pdf, ps, other] 

    quant-ph cs.AR cs.ET

    CryoZip: An Efficient Cryogenic Compressor for Quantum Error Correction Syndromes

    Authors: Guanchen Tao, Alexander Knapen, Jacob Mack, Gokul Subramanian Ravi, Qirui Zhang, Mehdi Saligane, Dennis Sylvester

    Abstract: Scaling fault tolerant quantum computing is increasingly constrained by the limited bandwidth and power budget across the 4 K to room temperature (RT) interface. We present CryoZip, a cross stack cryogenic compression framework that cooperates with a lightweight cryogenic quantum error correction (QEC) predecoder to reduce 4 K to RT syndrome transmission under realistic, circuit level noise. CryoZ… ▽ More

    Submitted 29 June, 2026; originally announced June 2026.

  17. arXiv:2606.16972  [pdf, ps, other] 

    cs.RO eess.SY

    When Should a Robot Replan? Regret-Guided Update Scheduling in Time-Varying MDPs

    Authors: Negin Musavi, Gokul Puthumanaillam, Ruben Hernandez, William Schafer, Melkior Ornik

    Abstract: Robots operating in non-stationary environments must continually adapt their policies as the dynamics drift, but onboard energy and compute budgets cap how often a full state estimation and re-planning step can be performed. This raises a question: \emph{when}, along a horizon, should a robot spend its limited budget? We formulate this problem in time-varying Markov decision processes (TVMDPs) wit… ▽ More

    Submitted 15 June, 2026; originally announced June 2026.

  18. arXiv:2606.16078  [pdf, ps, other] 

    cs.RO

    A Deployment Case Study in Robotic Apparel Automation: Digital Twin Integration, Interoperability, and Workforce Enablement

    Authors: Gokul Narayanan, Abhiroop Ajith, Jonathan Zornow, Carlos Calle, Auralis Herrero Lugo, Jose Luis Susa Rincon, Chengtao Wen, Eugen Solowjow

    Abstract: Despite steady advances in flexible automation in sectors such as electronics and automotive manufacturing, apparel automation remains challenging because fabrics are deformable and difficult to manipulate with robots. This paper presents a deployment-oriented case study of a robotic sewing system for denim manufacturing, emphasizing the system-level integration required for practical adoption. At… ▽ More

    Submitted 14 June, 2026; originally announced June 2026.

    Comments: 4 pages, 3 figures, IEEE ICRA 2026 Workshop Paper

  19. arXiv:2606.15154  [pdf, ps, other] 

    cs.RO

    Task-Aware Environment Augmentation for Reliable Navigation via Shielded Conditional Diffusion

    Authors: Bharawee Phoompho, Gokul Puthumanaillam, Yan Miao, Ruben Hernandez, Tim Bretl, Sayan Mitra, Melkior Ornik

    Abstract: Reliable trajectory planning under partial observability depends not only on computing a feasible geometric path, but also on whether the robot receives informative observations while executing that trajectory. Existing approaches usually keep the environment fixed and adapt the robot through belief-space planning, active localization, or added sensing, often incurring costly uncertainty propagati… ▽ More

    Submitted 13 June, 2026; originally announced June 2026.

  20. arXiv:2606.13672  [pdf, ps, other] 

    cs.RO

    WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

    Authors: Arnav Kumar Jain, Yilin Wu, Jesse Farebrother, Gokul Swamy, Andrea Bajcsy

    Abstract: The potential impacts of world models (WMs, i.e., learned simulators) on robotics are far-reaching -- policy evaluation, policy improvement, and test-time planning -- all with limited real-world interaction. To unlock these downstream capabilities, a WM needs to jointly satisfy three desiderata: $\textit{(i)}$ fidelity (i.e., producing simulated trajectories that correlate with reality),… ▽ More

    Submitted 16 June, 2026; v1 submitted 11 June, 2026; originally announced June 2026.

  21. arXiv:2606.12978  [pdf, ps, other] 

    cs.RO cs.CV eess.SY

    Trajectory-Level Redirection Attacks on Vision-Language-Action Models

    Authors: Gokul Puthumanaillam, Vardhan Dongre, Pranay Thangeda, Hooshang Nayyeri, Dilek Hakkani-Tür, Melkior Ornik

    Abstract: Vision-language-action (VLA) policies bring natural language into closed-loop robot control, enabling robots to execute manipulation tasks directly from text instructions. The same interface gives text a recurring role in control because the prompt is reused at every replanning step, and each prompt-conditioned action changes the future observations on which the policy acts. Existing VLA attacks s… ▽ More

    Submitted 13 June, 2026; v1 submitted 11 June, 2026; originally announced June 2026.

  22. arXiv:2606.12299  [pdf, ps, other] 

    cs.RO cs.LG

    Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

    Authors: Hyun Joe Jeong, Gokul Swamy, Andrea Bajcsy

    Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone. As a result, both human instructions and zero-shot language models can fail to relia… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

    Comments: 22 pages, 14 tables, 14 figures

  23. arXiv:2606.02800  [pdf, ps, other] 

    cs.CV cs.AI cs.LG cs.MM cs.RO

    Cosmos 3: Omnimodal World Models for Physical AI

    Authors: NVIDIA, :, Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson , et al. (271 additional authors not shown)

    Abstract: We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, worl… ▽ More

    Submitted 23 June, 2026; v1 submitted 1 June, 2026; originally announced June 2026.

  24. arXiv:2605.29198  [pdf, ps, other] 

    cs.CV

    Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

    Authors: Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yuta Kyuragi, Aditya Grover

    Abstract: Group-advantage-based reinforcement learning methods, such as GRPO and DAPO, have demonstrated strong performance across diverse domains, including mathematical reasoning and text-to-image generation. However, their reliance on sample-level rewards introduces a key limitation as uniform credit assignment across all tokens fails to capture fine-grained, token-level contributions. To address this is… ▽ More

    Submitted 29 May, 2026; v1 submitted 27 May, 2026; originally announced May 2026.

    Comments: 21 pages, 11 figures

  25. arXiv:2605.28526  [pdf, ps, other] 

    cs.AI cs.CL

    Entropy-aware Masking for Masked Language Modeling

    Authors: Gokul Srinivasagan, Kai Hartung, Munir Georges

    Abstract: Masked language modeling has become a standard pretraining objective for training encoder-based language models. In this approach, certain tokens in the input are masked, and the model learns to predict them using the surrounding context. This process enables the model to capture both syntactic and semantic properties of language. Conventionally, the tokens selected for masking are chosen at rando… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

    Comments: accepted at starsem 2026 Conference

  26. arXiv:2605.27461  [pdf, ps, other] 

    cs.RO

    A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons

    Authors: Brian Zhu, Philipp Schmitt, Philine Meister, Lukas Gensler, Momen Khalil, Emmanuele Poggi, Johannes Hechtl, Carsten Braunroth, Kai Wurm, Gokul Narayanan, Eugen Solowjow, Georg von Wichert, Andre Scholz, Felix Albrecht, Maxmillian Metzner

    Abstract: Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability demands of real-world deployment. We present a deployment study of an industrial packaging task at Siemens Factory (GWE, Erlangen, Germany), where a robot must pick a transparent accessory bag from a cluttered pile, insert it into the remaining cavity… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

  27. arXiv:2605.26349  [pdf, ps, other] 

    cs.RO

    Closing the Loop in Teleoperation: Episode-Level Data Quality Assessment and Feedback for High-Quality Demonstration Collection

    Authors: Gokul Narayanan, Yash Shahapurkar, Melih Erdogan, Brian Zhu, Eugen Solowjow

    Abstract: Industrial automation is at a pivotal moment, as Physical AI is driving a transition from rigid, hand-engineered automation systems toward more flexible and adaptive systems. This shift has created a growing demand for large-scale, real-world robot demonstration data, making teleoperation an increasingly important mechanism for data collection. However, high-quality teleoperated demonstrations rem… ▽ More

    Submitted 25 May, 2026; originally announced May 2026.

  28. arXiv:2605.23138  [pdf, ps, other] 

    quant-ph cs.AI cs.ET cs.LG

    Classical State Preparation for Variational Quantum Algorithms via Reinforcement Learning

    Authors: Gino Kwun, Dhanvi Bharadwaj, Gokul Subramanian Ravi

    Abstract: Variational Quantum Algorithms (VQAs) potentially offer a pathway to practical quantum advantage, but their optimization is heavily hindered by barren plateaus and numerous local minima. While classically simulable Clifford circuits can warm-start VQAs to accelerate convergence, existing heuristic-based initialization methods struggle to scale within vast combinatorial search spaces. To overcome t… ▽ More

    Submitted 21 May, 2026; originally announced May 2026.

    Comments: 22 pages, 4 figures

  29. arXiv:2605.21537  [pdf, ps, other] 

    cs.SE

    Articulate but Wrong: Self-Review Failures in LLM-Based Code Modernization

    Authors: Gokul Chandra Purnachandra Reddy, Aditya Lolla, Harsha Sanku

    Abstract: Large language model (LLM) agents are increasingly used to migrate legacy code to modern stacks. We ask a deceptively simple question: when an LLM modernizes legacy code, can the same model be relied upon to recognize when its own output silently changes observable behavior? We run 1,980 real modernization calls across 11 production LLMs from 7 distinct families on a balanced 60-snippet legacy-Pyt… ▽ More

    Submitted 20 May, 2026; originally announced May 2026.

    Comments: 11 pages, 6 figures, 2 tables. Corpus, oracle, output extractor, prompts, harness, self-review probe, and all 1,980 + 1,979 raw model outputs released as supplementary material at https://zenodo.org/records/20300861

    Report number: 11 pages, 6 figures, 2 tables. Corpus, oracle, prompts, harness, self-review probe, and all 1,980 + 1,979 raw model outputs released as supplementary material

  30. arXiv:2605.09999  [pdf, ps, other] 

    cs.RO cs.PF eess.SY

    Muninn: Your Trajectory Diffusion Model But Faster

    Authors: Gokul Puthumanaillam, Hao Jiang, Ruben Hernandez, Jose Fuentes, Paulo Padrao, Leonardo Bobadilla, Melkior Ornik

    Abstract: Diffusion-based trajectory planners can synthesize rich, multimodal robot motions, but their iterative denoising makes online planning and control prohibitively slow. Existing accelerations either modify the sampler or compress the network--sacrificing plan quality or requiring retraining without accounting for downstream control risk. We address the problem of making diffusion-based trajectory pl… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

    Comments: Accepted to Robotics: Science and Systems 2026

  31. arXiv:2605.03180  [pdf, ps, other] 

    quant-ph cs.AR cs.ET

    Mitigating Classical Resource Costs in Quantum Error Correction via Generalized qLDPC Predecoding

    Authors: Alexander Knapen, Junyi Luo, Guanchen Tao, Yuxuan Wang, Tomas Bruno, Qirui Zhang, Dennis Sylvester, Mehdi Saligane, Gokul Subramanian Ravi

    Abstract: Large-scale fault-tolerant quantum computing (FTQC) will require quantum-classical interfaces (QCIs) that orchestrate real-time decoding over thousands to millions of logical qubits simultaneously. To scale FTQC systems, complex decoding resources must be shared between logical qubits, creating resource contention bottlenecks in the QCI. Mitigating this contention via optimal resource allocation r… ▽ More

    Submitted 3 August, 2026; v1 submitted 4 May, 2026; originally announced May 2026.

    Comments: 14 pages, 11 figures; Updates to main text; 1 new experiment + minor updates to existing experiments, including new qLDPC codes

  32. arXiv:2604.21011  [pdf, ps, other] 

    cs.CV q-bio.NC

    Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition

    Authors: Naga VS Raviteja Chappa, Evangelos Sariyanidi, Lisa Yankowitz, Gokul Nair, Casey J. Zampella, Robert T. Schultz, Birkan Tunç

    Abstract: Micro-actions are subtle, localized movements lasting 1-3 seconds such as scratching one's head or tapping fingers. Such subtle actions are essential for social communication, ubiquitously used in natural interactions, and thus critical for fine-grained video understanding, yet remain poorly understood by current computer vision systems. We identify a fundamental challenge: micro-actions exhibit d… ▽ More

    Submitted 22 April, 2026; originally announced April 2026.

    Comments: Accepted to International Conference on Automatic Face and Gesture Recognition (FG)

  33. arXiv:2604.13107  [pdf, ps, other] 

    cs.SE cs.AI cs.LG

    Can Coding Agents Be General Agents?

    Authors: Maksim Ivanov, Abhijay Rana, Gokul Prabhakaran

    Abstract: As coding agents have seen rapid capability and adoption gains, users are applying them to general tasks beyond software engineering. In this post, we investigate whether coding agents can successfully generalize to end-to-end business process automation. We identify gaps in current evaluations, and conduct a case study to evaluate a coding agent on practical business tasks in an open-core Enterpr… ▽ More

    Submitted 10 April, 2026; originally announced April 2026.

  34. arXiv:2604.04418  [pdf, ps, other] 

    cs.HC cs.AI

    Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality

    Authors: Xiaoyuan Zhu, Kimberly Le Truong, Riccardo Fogliato, Gokul Swamy, Weijian Zhang, Minglai Yang, Longtian Ye, Bangya Liu, Minghao Liu, Andrew Ilyas, Steven Wu

    Abstract: As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or explanations. Yet, no standard measure exists for whether these justifications help users distinguish correct answers from incorrect ones. We formalize this idea as error verifiability and propose $v_{\text{bal}}$, a bal… ▽ More

    Submitted 8 April, 2026; v1 submitted 6 April, 2026; originally announced April 2026.

  35. arXiv:2603.14056  [pdf, ps, other] 

    cs.RO eess.SY

    Amortizing Trajectory Diffusion with Keyed Drift Fields

    Authors: Gokul Puthumanaillam, Melkior Ornik

    Abstract: Diffusion-based trajectory planners can synthesize rich, multimodal action sequences for offline reinforcement learning, but their iterative denoising incurs substantial inference-time cost, making closed-loop planning slow under tight compute budgets. We study the problem of achieving diffusion-like trajectory planning behavior with one-step inference, while retaining the ability to sample divers… ▽ More

    Submitted 14 March, 2026; originally announced March 2026.

  36. arXiv:2603.06512  [pdf, ps, other] 

    cs.RO cs.CV

    SG-DOR: Learning Scene Graphs with Direction-Conditioned Occlusion Reasoning for Pepper Plants

    Authors: Rohit Menon, Niklas Mueller-Goldingen, Sicong Pan, Gokul Krishna Chenchani, Maren Bennewitz

    Abstract: Robotic harvesting in dense crop canopies requires effective interventions that depend not only on geometry, but also on explicit, direction-conditioned relations identifying which organs obstruct a target fruit. We present SG-DOR (Scene Graphs with Direction-Conditioned Occlusion Reasoning), a relational framework that, given instance-segmented organ point clouds, infers a scene graph encoding ph… ▽ More

    Submitted 6 March, 2026; originally announced March 2026.

  37. arXiv:2603.05807  [pdf, ps, other] 

    cs.CV

    EventGeM: Global-to-Local Feature Matching for Event-Based Visual Place Recognition

    Authors: Adam D. Hines, Gokul B. Nair, Nicolás Marticorena, Michael Milford, Tobias Fischer

    Abstract: Event cameras are rapidly rising in popularity for robotic and computer vision tasks because their sparse activation delivers energy-efficient, high-dynamic-range, and fast sensing. Event cameras have been used in robotic navigation and localization tasks where positioning must occur in real time with sufficient accuracy. However, current event-based localization methods suffer from poor spatial u… ▽ More

    Submitted 18 September, 2026; v1 submitted 5 March, 2026; originally announced March 2026.

    Comments: 9 pages, 5 figures, 5 tables, under review

  38. arXiv:2603.04833  [pdf, ps, other] 

    cs.MA cs.AI

    SCoUT: Scalable Communication via Utility-Guided Temporal Grouping in Multi-Agent Reinforcement Learning

    Authors: Manav Vora, Gokul Puthumanaillam, Hiroyasu Tsukamoto, Melkior Ornik

    Abstract: Communication can improve coordination in partially observed multi-agent reinforcement learning (MARL), but learning \emph{when} and \emph{who} to communicate with requires choosing among many possible sender-recipient pairs, and the effect of any single message on future reward is hard to isolate. We introduce \textbf{SCoUT} (\textbf{S}calable \textbf{Co}mmunication via \textbf{U}tility-guided \t… ▽ More

    Submitted 5 March, 2026; originally announced March 2026.

  39. arXiv:2603.04736  [pdf, ps, other] 

    cs.LG

    Distribution-Conditioned Transport

    Authors: Nic Fishman, Gokul Gowri, Paolo L. B. Fischer, Marinka Zitnik, Omar Abudayyeh, Jonathan Gootenberg

    Abstract: Learning a transport model that maps a source distribution to a target distribution is a canonical problem in machine learning, but scientific applications increasingly require models that can generalize to source and target distributions unseen during training. We introduce distribution-conditioned transport (DCT), a framework that conditions transport maps on learned embeddings of source and tar… ▽ More

    Submitted 25 September, 2026; v1 submitted 4 March, 2026; originally announced March 2026.

  40. arXiv:2603.04730  [pdf, ps, other] 

    cs.LG

    Count Bridges enable Modeling and Deconvolving Transcriptomic Data

    Authors: Nic Fishman, Gokul Gowri, Tanush Kumar, Jiaqi Lu, Valentin de Bortoli, Jonathan S. Gootenberg, Omar Abudayyeh

    Abstract: Many modern biological assays, including RNA sequencing, yield integer-valued counts that reflect the number of molecules detected. These measurements are often not at the desired resolution: while the unit of interest is typically a single cell, many measurement technologies produce counts aggregated over sets of cells. Although recent generative frameworks such as diffusion and flow matching hav… ▽ More

    Submitted 4 March, 2026; originally announced March 2026.

  41. arXiv:2602.19041  [pdf, ps, other] 

    cs.LG

    Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning

    Authors: Jiahao Zhang, Lujing Zhang, Keltin Grimes, Zhuohao Yu, Gokul Swamy, Zhiwei Steven Wu

    Abstract: A recurring challenge in preference fine-tuning (PFT) is handling $\textit{intransitive}$ (i.e., cyclic) preferences. Intransitive preferences often stem from either $\textit{(i)}$ inconsistent rankings along a single objective or $\textit{(ii)}$ scalarizing multiple objectives into a single metric. Regardless of their source, the downstream implication of intransitive preferences is the same: the… ▽ More

    Submitted 5 May, 2026; v1 submitted 21 February, 2026; originally announced February 2026.

    Comments: 24 pages, 5 figures

  42. arXiv:2602.18166  [pdf, ps, other] 

    cs.PL

    Grammar Repair with Examples and Tree Automata: Extended Version

    Authors: Yunjeong Lee, Gokul Rajiv, Ilya Sergey

    Abstract: Context-free grammars (CFGs) are the de-facto formalism for declaratively describing concrete syntax for programming languages and generating parsers. One of the major challenges in defining a desired syntax is ruling out all possible ambiguities in the CFG productions that determine scoping rules as well as operator precedence and associativity. Practical tools for parser generation typically app… ▽ More

    Submitted 20 February, 2026; originally announced February 2026.

  43. arXiv:2601.18722  [pdf, ps, other] 

    cs.CL cs.LG

    Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning

    Authors: Lintang Sutawika, Gokul Swamy, Zhiwei Steven Wu, Graham Neubig

    Abstract: When asked a question in a language less seen in its training data, current reasoning large language models (RLMs) often exhibit dramatically lower performance than when asked the same question in English. In response, we introduce \texttt{SP3F} (Self-Play with Privileged Pairwise Feedback), a two-stage framework for enhancing multilingual reasoning without \textit{any} data in the target language… ▽ More

    Submitted 26 January, 2026; originally announced January 2026.

    Comments: Code available at https://github.com/lintangsutawika/SP3F

  44. arXiv:2601.15118  [pdf, ps, other] 

    cs.SD cs.CL cs.LG

    WavLink: Compact Audio-Text Embeddings with a Global Whisper Token

    Authors: Gokul Karthik Kumar, Ludovick Lepauloux, Hakim Hacid

    Abstract: Whisper has become the de-facto encoder for extracting general-purpose audio features in large audio-language models, where a 30-second clip is typically represented by 1500 frame features projected into an LLM. In contrast, audio-text embedding models like CLAP-based models have largely relied on alternative audio encoders (e.g., HTS-AT, PaSST), and have not leveraged Whisper effectively. We pres… ▽ More

    Submitted 22 January, 2026; v1 submitted 21 January, 2026; originally announced January 2026.

    Comments: Accepted at ICASSP 2026

  45. arXiv:2512.17655  [pdf, ps, other] 

    cs.CV q-bio.NC

    Bitbox: Behavioral Imaging Toolbox for Computational Analysis of Behavior from Videos

    Authors: Evangelos Sariyanidi, Gokul Nair, Lisa Yankowitz, Casey J. Zampella, Mohan Kashyap Pargi, Aashvi Manakiwala, Maya McNealis, John D. Herrington, Jeffrey Cohn, Robert T. Schultz, Birkan Tunc

    Abstract: Computational measurement of human behavior from video has recently become feasible due to major advances in AI. These advances now enable granular and precise quantification of facial expression, head movement, body action, and other behavioral modalities and are increasingly used in psychology, psychiatry, neuroscience, and mental health research. However, mainstream adoption remains slow. Most… ▽ More

    Submitted 19 December, 2025; originally announced December 2025.

  46. arXiv:2512.14014  [pdf, ps, other] 

    cs.AI

    MobileWorldBench: Towards Semantic World Modeling For Mobile Agents

    Authors: Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Aditya Grover

    Abstract: World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face practical limitations in GUI settings, where predicting complex visual elements in future states is often difficult. In this work, we explore an alternative formulation of world modeling for GUI agents, where state transitio… ▽ More

    Submitted 15 December, 2025; originally announced December 2025.

    Comments: 21 pages, 13 figures

  47. arXiv:2512.12068  [pdf, ps, other] 

    quant-ph cs.AR cs.DC cs.ET

    TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms

    Authors: Yuewen Hou, Dhanvi Bharadwaj, Gokul Subramanian Ravi

    Abstract: Variational Quantum Algorithms (VQAs) are promising for near- and intermediate-term quantum computing, but their execution cost is substantial. Each task requires many iterations and numerous circuits per iteration, and real-world applications often involve multiple tasks, scaling with the precision needed to explore the application's energy landscape. This demands an enormous number of execution… ▽ More

    Submitted 19 December, 2025; v1 submitted 12 December, 2025; originally announced December 2025.

    Comments: To appear at 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2026)

  48. Pinball: A Cryogenic Predecoder for Surface Code Decoding Under Circuit-Level Noise

    Authors: Alexander Knapen, Guanchen Tao, Jacob Mack, Tomas Bruno, Mehdi Saligane, Dennis Sylvester, Qirui Zhang, Gokul Subramanian Ravi

    Abstract: Scaling fault tolerant quantum computers, especially cryogenic systems based on the surface code, to millions of qubits is challenging due to poorly-scaling data processing and power consumption overheads. One key hurdle is the design of real-time quantum error correction (QEC) decoders, which demands high data rates for error processing; this is particularly apparent in systems with cryogenic qub… ▽ More

    Submitted 12 December, 2025; v1 submitted 10 December, 2025; originally announced December 2025.

    Comments: Minor text/figure/title updates. 17 pages, 26 figures. To appear at the 32nd IEEE International Symposium on High-Performance Computer Architecture (HPCA 2026)

  49. arXiv:2511.05931  [pdf, ps, other] 

    cs.AI cs.SE

    Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement

    Authors: Hiroaki Hayashi, Bo Pang, Wenting Zhao, Ye Liu, Akash Gokul, Srijan Bansal, Caiming Xiong, Semih Yavuz, Yingbo Zhou

    Abstract: Large language model (LLM) based agents are increasingly used to tackle software engineering tasks that require multi-step reasoning and code modification, demonstrating promising yet limited performance. However, most existing LLM agents typically operate within static execution frameworks, lacking a principled mechanism to learn and self-improve from their own experience and past rollouts. As a… ▽ More

    Submitted 8 November, 2025; originally announced November 2025.

  50. arXiv:2510.26336  [pdf, ps, other] 

    cs.CL cs.AI

    From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning

    Authors: Nishit Neema, Srinjoy Mukherjee, Sapan Shah, Gokul Ramakrishnan, Ganesh Venkatesh

    Abstract: Large Language Models (LLMs) excel at general tasks but underperform in specialized domains like economics and psychology, which require deep, principled understanding. To address this, we introduce ACER (Automated Curriculum-Enhanced Regimen) that transforms generalist models into domain experts without sacrificing their broad capabilities. ACER first synthesizes a comprehensive, textbook-style c… ▽ More

    Submitted 30 October, 2025; originally announced October 2025.