Introducing LongCat-2.0: Trillion-Parameter MoE Large Language Model
1.6 trillion parameter MoE architecture with 1M long context, trained on domestic superpod accelerators, excelling in coding and agent tasks.
1.6 trillion parameter MoE architecture with 1M long context, trained on domestic superpod accelerators, excelling in coding and agent tasks.
Proposes GUI-CIDER mid-training method that explicitly internalizes GUI world knowledge into models through causal internalization and density-aware sample reselection.
Identifies user noise and tool noise as two major sources of interaction noise, proposing training in noisy environments to enhance LLM agent robustness.