Stars
A privacy-first app that strips AI watermarks from content you own.
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
14MB foundation model for tiny devices; phones, wearables, smart home, and robots.
Universal LLM Deployment Engine with ML Compilation
Tile primitives for speedy kernels
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
DeepSeek Harness: Everything is a Plugin.
Automatic verification of LLVM optimizations
A list of awesome compiler projects and papers for tensor computation and deep learning.
The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-request decode with support of FP8 weight ( Join Discord :https://…
A simple C++11 Thread Pool implementation
📚 Modern C++ Tutorial: C++11 to C++26 On the Fly | https://changkun.de/modern-cpp/
Empowering everyone to build reliable and efficient software.
egg is a flexible, high-performance e-graph library
Open-source high-performance RISC-V processor
The CORE-V CVA6 is a highly configurable, 6-stage RISC-V core for both application and embedded applications. Application class configurations are capable of booting Linux.
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
A tutorial on modern GPU programming for machine learning systems
🍎 One kernel a day keeps high latency away. A hands-on CUDA learning path featuring a rich collection of kernels, from the basics to peak performance, seamlessly integrated as PyTorch C++ extensions.
PyTorch Tutorial for Deep Learning Researchers
Tensors and Dynamic neural networks in Python with strong GPU acceleration




