Papers with Code

Recent AI research papers

Canonical AI research papers added or published in the last fourteen days.

  1. Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions
  2. DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection
  3. WebWorld: The Browser as a World Model for Self-Improving Web Code
  4. LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
  5. Verification-Aware Training for Speculative Decoding
  6. Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
  7. Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation
  8. PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
  9. Normalized Low-Rank Adaptation
  10. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
  11. CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
  12. CogEvol: Towards Efficient and Reliable Learning Environment Generation
  13. Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
  14. DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
  15. BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
  16. MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
  17. Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
  18. Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
  19. PaperGym: Rubric-Centered Evolution for Research-Plan Generation
  20. Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
  21. MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
  22. SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
  23. ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
  24. Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
  25. Cross-lingual Functional Vectors for Emotion Detection in Large Language Models
  26. NutriTrack: Multimodal AI Food Recognition with USDA RAG Nutrition Lookup
  27. EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
  28. Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
  29. SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
  30. Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models
  31. Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase
  32. Dynamic Important Example Mining for Reinforcement Finetuning
  33. GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
  34. Evaluating the Hidden Costs of Personalization in Large Language Models
  35. Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
  36. EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
  37. Sliding-window beats linear attention
  38. Hy4 preview
  39. LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
  40. Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
  41. Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
  42. LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
  43. OpenStamp: A Watermark for Open-Source Language Models
  44. Video Generative Models as Geometry Learner
  45. Rubric-to-Code Credit Assignment for Reinforcement Learning
  46. Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
  47. ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
  48. Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
  49. Generative Semantic Scene Completion
  50. Deterministic multi-hash routing supports long-horizon training in a compact language model
  51. Fast Weight Attention for Continual Learning
  52. Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
  53. Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
  54. J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
  55. Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
  56. Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
  57. Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
  58. LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
  59. CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
  60. EditaLive! Unified Character Video Editing for Live Streaming
  61. PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
  62. Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
  63. WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
  64. Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
  65. Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
  66. Magpie: Real-Time World Renderer for Interactive Games
  67. Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
  68. UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
  69. TTPO: Test-Time Policy Optimization
  70. PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
  71. What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
  72. GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models
  73. Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
  74. LMSM: LLM Security Framework Inspired by Linux Security Modules
  75. TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
  76. CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
  77. Procedura: Agentic 3D Modeling with Procedural Control
  78. Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
  79. Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
  80. Prefix Sliding for efficient test-time scaling
  81. Skill Issue: Are Skills Language-Invariant in LLMs?
  82. D'Agnese DIF-FNO: Diffeomorphic Implicit Fourier Neural Operators with Topological Guarantees on Non-Convex Domains
  83. SureRoute: Toward a Hallucination-Free Self-Improving Platform for Retrosynthesis
  84. RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval
  85. MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
  86. Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
  87. VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
  88. A Programming Paradigm for Spatiotemporal Composability
  89. StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
  90. V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning
  91. Code World Model: Coding Agent as World Brain
  92. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
  93. JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
  94. GLM-5.3-Flash: Frontier Intelligence, Flash Cost
  95. Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation
  96. Granite 4.2 LLMs: How They're Built
  97. StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
  98. StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
  99. PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
  100. Luce: Relightable Gaussians for 3D Asset Generation