Tencent Hy News · 2027-08-05
Understanding pretraining dynamics is crucial for designing effective training hyperparameters for large language models (LLMs), particularly the learning rate (LR). However, LR does not always reflect how much the function represented by the model actually changes, and the optimal LR often shifts with model and data scale. In this post, we identify that the effective learning rate (ELR), which controls the directional changes of model weights, is a more intrinsic heuristic than LR as a tunable hyperparameter. Specifically, ELR delivers more accurate loss prediction under the multi-power law (MPL) model and transfers more reliably across model scales. More broadly, the ELR perspective guides the design of better schedules, which outperform conventional LR-schedule baselines.
Open original
Tencent Hy News · 2026-08-27
Hy LLM
Open original
Tencent Hy News · 2026-07-21
Today, we are introducing Hyra-1.0, the first release of Hyra (Hunyuan Research Agent) - an agent capable of recursive self-improvement, purpose-built for performance-driven research and engineering tasks.
Open original
Tencent Hy News · 2026-07-07
Hy LLM
Open original
Tencent Hy News · 2026-04-30
Previously, we built CL-Bench to test whether AI could learn complex new knowledge, such as rules, and use it effectively. But real life doesn't come with a rulebook. To be a helpful everyday assistant, an AI must do more than just understand complex rules; it should also understand complex human relationships and piece together clues from messy, scattered context.
Open original
Tencent Hy News · 2026-04-22
Hy LLM
Open original
Tencent Hy News · 2026-02-13
Guanhua Huang;Tingqiang Xu;Jinbo Wang
Open original
Tencent Hy News · 2026-02-03
Shihan Dou and Pluto Zhou
Open original