1. X
  2. Medical Sphere
Log inSign up
Medical Sphere
2,061 posts
user avatar
Medical Sphere
@MedicalSphereAI
The global community for advancing AI in healthcare Tag @AskMedSphere to test AI models on medical cases
馃寧
medicalsphere.ai
Joined September 2025
1
Following
1,915
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what鈥檚 happening and join the conversation

Continue with phone
or
Log in with username or email
Terms路Privacy路Cookies路Accessibility路Ads Info路漏 2026 X Corp.
  • Pinned
    user avatar
    Medical Sphere
    @MedicalSphereAI
    Jul 13
    We benchmarked Muse Spark 1.1 and GPT-5.6 Sol on HealthBench Professional, OpenAI's benchmark of 525 real clinician tasks 馃彞馃┖ Muse Spark 1.1 tops our board: better overall score than GPT-5.6 Sol, statistically on par on the length-adjusted score at a fraction of the cost
    184K
  • user avatar
    Medical Sphere
    @MedicalSphereAI
    Jul 29
    Want to quickly share your Medical AI Council summary with the community? We made it easy! 馃┖馃馃憞
    00:00
    302
  • user avatar
    Medical Sphere
    @MedicalSphereAI
    Jul 28
    Claude Opus 5's performance on MedAgentBench drops substantially compared to its predecessors, Opus 4.7 and 4.8: roughly half their success rate. 馃彞 We benchmarked Opus 5 as an autonomous clinical agent across 10 EHR task types, and it landed at ~40% pass@1 (avg of 3脳 300-task
    2.6K
  • user avatar
    Medical Sphere
    @MedicalSphereAI
    Jul 24
    Claude Opus 5 is now on HealthBench Professional. 馃彞 It debuts 2nd 馃 on overall score (0.68, mean of 3 runs) behind Muse Spark 1.1 (0.69) and ahead of GPT-5.6 Sol (0.63). On the length-adjusted score it slots into 3rd 馃 (0.57): its responses are long enough that the length
    3.1K
  • user avatar
    Medical Sphere
    @MedicalSphereAI
    Jul 23
    We put Gemini 3.6 Flash to the test on two healthcare AI benchmarks: HealthBench Professional and MedAgentBench. 馃彞馃┖ On HealthBench Professional, the 0.523 score makes Gemini 3.6 Flash the strongest Gemini model we鈥檝e tested so far, ahead of Gemini 3.5 Flash (0.510) and Gemini
    539