-
National University of Singapore
- Singapore
-
15:24
(UTC +08:00)
Stars
Accelerating MoE with IO and Tile-aware Optimizations
VideoSys: An easy and efficient system for video generation
高性价比人生指南: 长寿防病、急救、省钱理财、法律红线、失业与工伤、医保社保、恋爱婚育、怀孕育儿、创业与做平台合规、出国与技能。每条写明成本、收益、证据等级和原始出处,只引期刊论文与官方文件。
HieraSparse: Hierarchical Semi-Structured KV-Cache Attention on Sparse Tensor Core
Code for the paper “SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference”
Code for the paper "Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning" https://arxiv.org/abs/2609.03430
My Python scripts to make high-quality figures for publications in top AI conferences and journals.
Defend sound research against LLM reviewer bias through meaning-preserving rewrites.
Benchmarking Knowledge Transfer in Lifelong Robot Learning
Official repository of LIBERO-plus, a generalized benchmark for in-depth robustness analysis of vision-language-action models.
Measuring and evolving with the frontier of agent work
A Vulkan renderer for side-by-side quality and performance comparison of upscaling algorithms.
Memory is a long-term memory module for AI agents, providing capabilities for memory extraction, storage, retrieval, and migration for agents running on the openJiuwen framework.
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
mars-compute-ai / xgrammar-xserve
Forked from mlc-ai/xgrammarFast, Flexible and Portable Structured Generation
FlashInfer: Kernel Library for LLM Serving
Benchmarking physical understanding in generative video models
[NeurIPS'24] HippoRAG is a novel RAG framework inspired by human long-term memory that enables LLMs to continuously integrate knowledge across external documents. RAG + Knowledge Graphs + Personali…
PhyAI is a high-performance framework for running Physical AI models (VLA, WAM, and beyond), supporting both cloud-based serving and on-device deployment.
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
PonderTTT: Adaptive Budget-Aware Test-Time Training
TokenSpeed is a speed-of-light LLM inference engine.
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
提供一个人人会用的的路由、NAS系统 (目前活跃的分支是 istoreos-24.10,main或master分支不维护请勿使用)
Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"




