Log inSign up
Steven Dillmann
339 posts
Steven Dillmann profile banner
@StevenDillmann

Steven Dillmann

@StevenDillmann
AI4Science PhD @Stanford Terminal-Bench-Science Lead @harborframework Research Intern @allen_ai Prev. @Cambridge_Uni, @NASAJPL, @imperialcollege
Stanford, CA
Joined January 2020
1,734
Following
1,435
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @StevenDillmann
    Steven Dillmann
    @StevenDillmann
    Aug 27
    We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions
    54
  • @StevenDillmann
    Steven Dillmann
    @StevenDillmann
    Aug 30
    I left Germany at the age of 18 for a reason. I miss my family, I miss my friends, I miss Döner Kebab, but German bureaucracy & begrudgery kill the ambition of every kid who wanted to build or achieve something extraordinary, before they ever get the chance. Germany will never
    @patrickc
    Patrick Collison
    Stripe
    @patrickc
    Aug 29
    Met a German founder this week and asked him if all the stories one reads about the challenges of startups in Germany are exaggerated. "No, they're understated." Proceeded to describe spending a full day having a 90-page investment contract read to him (mandatory under German
    30
  • @StevenDillmann
    Steven Dillmann
    @StevenDillmann
    Aug 29
    For anyone who wants to contribute to Terminal-Bench-Science 0.2 - please read this post by our senior reviewer @neversupervised. By far the two most common task-quality issues we found while reviewing tasks for v0.1. terminal-bench-science.ai
    @neversupervised
    Ivan Bercovich
    @neversupervised
    Aug 27
    Two common patterns when reviewing computational science tasks. 1) Verifiers were co-designed with the oracle, failing valid solutions that used methods unanticipated by the author. 2) Authors using synthetic data had privileged knowledge of the generative function and were able
  • @StevenDillmann
    Steven Dillmann
    @StevenDillmann
    Aug 29
    And after yesterday’s launch of Terminal-Bench-Science 0.1, here is our latest and top-quality dataset: Terminal-Bench 4.0. Great work @ryan_marten & co.
    @ryan_marten
    Ryan Marten
    @ryan_marten
    Aug 29
    We've pushed a version update to the Terminal-Bench dataset and leaderboard. Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
    00:00
    1
  • @StevenDillmann
    Steven Dillmann
    @StevenDillmann
    Aug 28
    AI for science is the next big thing after AI for coding. Congratulations to @ChrisHayduk and his team at @OpenAI. If you’re a scientist, you can now actively shape AI for Science progress with Terminal-Bench-Science: terminal-bench-science.ai
    @ChrisHayduk
    Chris Hayduk
    @ChrisHayduk
    Aug 28
    At OpenAI, it's felt like our researchers & engineers have been living in the future with Codex. Today, we're bringing that same experience to life science researchers with the Rosalind Workbench.
    00:00
    1