Skip to content
View Aaryan-Kapoor's full-sized avatar
👀
Waiting for AGI...
👀
Waiting for AGI...

Highlights

  • Pro

Block or report Aaryan-Kapoor

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Aaryan-Kapoor/README.md

who is Aaryan? - a self-playing agent session

AI Engineer & Researcher - agents, inference, evals

I make models do cool stuff for you at scale :)

100,000+ model downloads · 5.2M+ views of my technical writing · 1,000+ public GPU-template hours/month · Multi hackathon winner

Research Highlights

  • First author, ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval - EMNLP 2026 Main Conference. Code retrieval that measures whether the code runs, not whether it looks right. Paper: arXiv:2609.01865, Dataset + harness

Agent systems Highlights

  • Surface (previously public) - a universal AI agent interaction primitive.
  • 24hr-research-agent - decomposes a question into 200+ tasks and compiles book-length, fully cited reports from 1,000+ sources
  • SOTA-Coder - autonomous coding-agent harness built from scratch: event-driven runtime, WebSocket streaming, dynamic MCP tool loading, no agent-SDK dependencies
  • video-production-skill, d2l-cli / degreeworks-cli

Inference & models

  • AK quant lines - custom per-tensor bit allocations, benchmarked head-to-head on KL divergence with 95% bootstrap CIs:
    • Qwen3.5-9B: 31 wins / 3 ties / 0 losses across 34 comparisons vs Unsloth, bartowski, lmstudio, mradermacher, byteshape, AtomicChat - up to 49% lower KLD than Unsloth at the same size
    • Meta Muse-Glimmer-30B: 26 wins / 6 ties / 0 losses vs Meta's own, Unsloth and bartowski
  • Auto-Quant (private) - my end-to-end quantize-test-release pipeline behind huggingface.co/AaryanK: Day-0 GGUFs for major launches, 100k+ downloads
  • TurboQuant KV-cache quantization in llama.cpp - 4.4× smaller KV cache, no throughput loss; ported to ik_llama.cpp by the community
  • One of the first toggleable-reasoning LLMs (Feb 2025, DOI) - trained with GRPO on a single consumer GPU
  • Also: upstream llama.cpp contributor (MiMo-V2-Flash, Kimi Linear)

Will work for GPUs · The Why: waiting for RSI so I can take a nap

Hugging Face • LinkedIn • OpenReview • Email

Pinned Loading

  1. z-image-turbo z-image-turbo Public

    A professional web interface for the Tongyi-MAI Z-Image-Turbo model — lightning-fast text-to-image generation with 6B parameters.

    JavaScript 140 19

  2. ModelGate-Hackathon ModelGate-Hackathon Public

    🏆Winning Project | ModelGate is a contract-aware AI control plane that ingests customer contracts, extracts SLA/privacy/routing constraints, and generates an OpenAI-compatible endpoint that automat…

    TypeScript 52 8

  3. 24hr-research-agent 24hr-research-agent Public

    An experimental autonomous research system that conducts comprehensive, multi-hour research sessions and produces book-length reports with full citations on any topic.

    HTML 51 3

  4. video-production-skill video-production-skill Public

    Turn an AI agent into a full-stack educational video producer: research, storyboard, narration, animated visuals, render, QC, archive, and delivery.

    Python 16 3

  5. Reflect-KSU-Hackathon Reflect-KSU-Hackathon Public

    🏆Winning Project | Reflect is a cutting-edge mental health counseling platform designed to bridge the gap between students and counselors. By integrating real-time video therapy with advanced AI an…

    JavaScript 7 1