Works Mentioning or Using ClawBench
Public papers and reports that cite, discuss, or build on ClawBench and related real-world agent evaluation work.
Paper • 2605.14271 • Published • 55Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
NeuroClaw Technical Report
Paper • 2604.24696 • Published • 1Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
Paper • 2605.13831 • Published • 90Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 19Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
Paper • 2605.27882 • Published • 17Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 46Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
Paper • 2607.08768 • Published • 34Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Paper • 2606.22557 • PublishedNote Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
Paper • 2605.14678 • Published • 108Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
AcademiClaw: When Students Set Challenges for AI Agents
Paper • 2605.02661 • Published • 17Note Paper/report that cites or discusses ClawBench; see the linked source for the precise relationship.
Mach-Mind-4-Flash Technical Report
Paper • 2607.09375 • PublishedNote Indexed as a citing paper for ClawBench by Semantic Scholar; inspect the paper for the exact citation context.
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
Paper • 2606.16748 • Published • 7Note Indexed as a citing paper for ClawBench by Semantic Scholar; inspect the paper for the exact citation context.
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
Paper • 2605.26086 • Published • 25Note Indexed as a citing paper for ClawBench by Semantic Scholar; inspect the paper for the exact citation context.
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
Paper • 2605.10779 • PublishedNote Indexed as a citing paper for ClawBench by Semantic Scholar; inspect the paper for the exact citation context.
On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows
Paper • 2605.06110 • Published • 17Note Indexed as a citing paper for ClawBench by Semantic Scholar; inspect the paper for the exact citation context.