Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
An end-to-end engineering practice for the MiMo-V2.5 series inference system, covering KVCache management, tiered caching, SWA-aware prefix cache trees, scheduling, Prefill/Decode pipelines, and multimodal optimizations.