Hi team,
great job on FreeToken! The adaptive CPU-GPU co-execution strategy for MoE offloading on a single node is impressive.
I was wondering: is multi-node / distributed inference in your current roadmap? Enabling FreeToken to scale across a distributed cluster of consumer/commodity hardware (sharding experts across multiple networked servers) would completely disrupt enterprise AI infrastructure costs, bypassing single-node PCIe/VRAM limitations.
Thanks!
Hi team,
great job on FreeToken! The adaptive CPU-GPU co-execution strategy for MoE offloading on a single node is impressive.
I was wondering: is multi-node / distributed inference in your current roadmap? Enabling FreeToken to scale across a distributed cluster of consumer/commodity hardware (sharding experts across multiple networked servers) would completely disrupt enterprise AI infrastructure costs, bypassing single-node PCIe/VRAM limitations.
Thanks!