1. X
  2. Xinyu Zhu
Log inSign up
Xinyu Zhu
179 posts
user avatar
Xinyu Zhu
@tianhongzxy
RS Intern @Meta MSL. CS Ph.D. student @UVA | Prev @Apple @MSFTResearch, master @Tsinghua_uni.
Bellevue, WA
zhuxinyu.top
Joined November 2017
531
Following
485
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Xinyu Zhu
    @tianhongzxy
    Jul 17
    Introducing our latest work, Self-Guided TTT by @AIatMeta and @UVA.🚀 LLMs can read 100k+ tokens, but reading ≠ using. Test-Time Training (TTT) helps, but only if the model trains on the RIGHT tokens.💡 Our fix: the model picks which spans to train on at test time.👇🧵 1/
    33K
  • user avatar
    Xinyu Zhu
    @tianhongzxy
    May 14
    🚀 <20% RLVR training → 100% performance via checkpoint extrapolation! RELEX predicts future checkpoints from minimal training dynamics! Check out this great project led by @weizhepei!
    user avatar
    Zhepei Wei
    @weizhepei
    May 14
    😢RLVR is powerful but expensive 🤯Imagine using <20% RLVR training while achieving 100% performance? Sounds surprising? We show that minimal RLVR training is enough to know where training is going, and predict future ckpts at no training cost! 📃tinyurl.com/minimal-rlvr 🧵[1/n]
    2.1K
  • user avatar
    Xinyu Zhu
    @tianhongzxy
    Apr 20
    Didn’t think I’d ever miss google translate until X switched to Grok.
    185
  • user avatar
    Xinyu Zhu
    @tianhongzxy
    Dec 2, 2025
    Excited to be at #NeurIPS2025 in San Diego this week! 🌞🌊 I'll be presenting our work on Wednesday 11 AM at Exhibit Hall C/D/E, #1910. Come by and chat about LLM research!🤩
    user avatar
    Xinyu Zhu
    @tianhongzxy
    Jun 2, 2025
    🔥The debate’s been wild: How does the reward in RLVR actually improve LLM reasoning?🤔 🚀Introducing our new paper👇 💡TL;DR: Just penalizing incorrect rollouts❌ — no positive reward needed — can boost LLM reasoning, and sometimes better than PPO/GRPO! 🧵[1/n]
    2.9K
  • user avatar
    Xinyu Zhu
    @tianhongzxy
    Oct 9, 2025
    Outcome-only rewards aren’t enough for search agents!🚫 We show that training LLM agents only on final answers leads to poor search behaviors. We propose a two-stage RL training framework that decouples search and answering for smarter search agents! Great work led by @blancokdb!
    user avatar
    Yiding Wang @ACL 2026
    @blancokdb
    Oct 9, 2025
    Most RL methods train search agents on final answers, assuming good search will follow. But do they? 🤔 Our research finds this assumption is flawed❗️ 🤩Meet DeSA (Decoupling Search and Answering) – a 2-stage framework improving both search quality and answer accuracy! 🧵[1/n]
    1.2K