Skip to content
View kzos's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report kzos

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. vllm-project/vllm vllm-project/vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 93k 22.9k

  2. Dao-AILab/flash-attention Dao-AILab/flash-attention Public

    Fast and memory-efficient exact attention

    Python 25k 3.1k

  3. flashinfer-ai/flashinfer flashinfer-ai/flashinfer Public

    FlashInfer: Kernel Library for LLM Serving

    Cuda 6.5k 1.5k

  4. vllm-project/semantic-router vllm-project/semantic-router Public

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    Go 6k 993

  5. VectorSpaceLab/general-agentic-memory VectorSpaceLab/general-agentic-memory Public

    A general memory system for agents, powered by deep-research

    Python 862 87