Looking forward to attending #NeurIPS2024 this week. Come meet @itsnamgyu who will present on block-transformer, which speeds up decoding by up to 20x.
Do you know your LLM uses less than 1% of your GPU at inference? Too much time is wasted on KV cache memory access ➡️ We tackle this with the 🎁 Block Transformer: a global-to-local architecture that speeds up decoding up to 20x 🚀
@KAIST_AI @LG_AI_Research w/ @GoogleDeepMind 🧵





