Log inSign up
Log inSign up
Ming Zhong
134 posts
@MingZhong_

Ming Zhong

@MingZhong_
Research Scientist at @GoogleDeepmind
California, USA
maszhongming.github.io
Joined September 2022
1,131
Following
2,111
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @MingZhong_
    Ming Zhong
    @MingZhong_
    Oct 9, 2025
    Vibe coding with an LLM, but the final vibe is off? 🤔 We analyze why models fail the "vibe check" and what truly matters to users. Key insight: human preference 🧑‍💻 ≈ functional correctness ✅ + instruction following 🎯 Check out our paper: arxiv.org/abs/2510.07315
    2
  • @MingZhong_
    Ming Zhong
    @MingZhong_
    Oct 10, 2025
    What makes for a good "vibe coding" experience? Check out our recent work on digging into the factors that truly shape the user preference. It's more than just correctness! 👇
    @sunjiao123sun_
    Jiao Sun
    @sunjiao123sun_
    Oct 10, 2025
    🆕Drop from Gemini VibeCoding Team: Why does code functionality correctness NOT necessarily translate to 𝗯𝗲𝘁𝘁𝗲𝗿 𝘂𝘀𝗲𝗿 𝗽𝗿𝗲𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗳𝗼𝗿 𝗰𝗼𝗱𝗶𝗻𝗴 tasks? In 𝗩𝗶𝗯𝗲𝗖𝗵𝗲𝗰𝗸𝗲𝗿: arxiv.org/abs/2510.07315, we find that the human preference correlates the
  • @MingZhong_
    Ming Zhong
    @MingZhong_
    Sep 3, 2025
    LLMs are winning gold medals in math & coding🥇, but what if they have to learn a new language from scratch, with only a grammar book and a dictionary? 🤔 On our new language Camlang, GPT-5's reasoning plummets: 98% in English → 47%. Check out our paper for more details!
    @Yulongchen1010
    Yulong Chen
    @Yulongchen1010
    Sep 3, 2025
    Can LLMs learn a new language using only a grammar book and a dictionary like how human adult L-2 learners do? Check our in-progress paper! The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang arxiv.org/pdf/2509.00425
    1
  • @MingZhong_
    Ming Zhong
    @MingZhong_
    Jun 20, 2025
    Thrilled to share our new reasoning model, Polaris✨! The 4B version achieves a score of 79.4 on AIME 2025, surpassing Claude 4 Opus (75.5) We’re releasing the full RL recipe, data, and weights 🔓 — see all the details below
    @AnChancy46881
    Chenxin An
    @AnChancy46881
    Jun 20, 2025
    # 🚨 4B open-recipe model beats Claude-4-Opus 🔓 100% open data, recipe, model weights and code. Introducing Polaris✨--a post-training recipe for scaling RL on advanced reasoning models. 🥳 Check out how we boost open-recipe reasoning models to incredible performance levels
    1
  • @MingZhong_
    Ming Zhong
    @MingZhong_
    Apr 24, 2025
    I will be presenting our poster for the “Law of the Weakest Link” paper at ICLR today! If you're interested in this topic, feel free to stop by and chat! 📍 Location: Hall 3 + Hall 2B #257 ⏰ Time: Apr 25 | 10:00 AM – 12:30 PM SGT
    @MingZhong_
    Ming Zhong
    @MingZhong_
    Oct 1, 2024
    Excited to share our recent work! We define and benchmark cross capabilities in LLMs, revealing the "Law of the Weakest Link": collaborative performance clusters around the weakest individual capability. 📄 Paper: arxiv.org/abs/2409.19951 🌐 Website: llm-cross-capabilities.org