Tencent Open-Sources Hy4 Preview: The Cost and Evidence Gap Behind 770B Parameters and a Million-Token Context
TL;DR
Tencent Hy4 preview enters the open-model market with 770B total parameters, 49B active parameters, and a context exceeding one million tokens; its internal evaluation, inference cost, and 31.8% throughput gain still need external reproduction.
Tencent Hy4 preview comes with a cost claim that can be tested over the next three to six months. If developers repeat the work with the same million-token input, concurrency, and hardware but cannot approach Tencent’s reported throughput gain, or if the output tokens needed to finish the same engineering tasks erase the advantage in list price, the release will establish availability without establishing cheaper production work. Tencent released and open-sourced the model on August 28, 2026. Both Tencent and TechNode report 770B total parameters, 49B active parameters, and a context window exceeding one million tokens.
The model is available through WorkBuddy, CodeBuddy, Yuanbao, and ima. Developers can also call it through Tencent Cloud TokenHub and OpenRouter. WorkBuddy and CodeBuddy provide free access for two weeks. The API costs USD 0.834 per million input tokens, USD 2.501 per million output tokens, and USD 0.042 per million tokens served from cache. Those prices make a token bill calculable, but they do not reveal the cost of completing a job. Total spending also depends on retries after errors, output length, and latency when the prompt approaches one million tokens.
Tencent lists software engineering, office work, data analysis, game development, and scientific research among the target uses. Its internal blind evaluation involved 163 experts and 203 engineering tasks. Hy4 preview received an average score of 2.99 out of 4.00, compared with 2.92 for GLM-5.3 and 2.94 for Kimi K3. The difference is only 0.05 to 0.07 points. Tencent did not publish inter-rater agreement, the distribution of tasks, model settings, or statistical intervals. The figures describe an internal result and should not be treated as a substitute for an independent evaluation.
The announcement makes a second measurable claim about the inference system. Tencent says the model analyzed bottlenecks and iterated on operator fusion and communication optimization, raising end-to-end throughput by 31.8% from a baseline. The company says gains held across different context lengths and concurrency levels. It does not identify the baseline hardware, absolute token rate, latency percentiles, or power use. Without those denominators, a buyer cannot separate improvements caused by the model, the software stack, or a particular cluster configuration, and cannot compare output per dollar with another service.
Tencent also used Hy4 preview to assist with training methods, data strategies, evaluation frameworks, and low-level operators. According to the company, the model proposed approaches, ran experiments, and fed code, logs, and feedback into another round of exploration. Tencent calls this an early-stage recursive self-improvement loop. The published material does not show the model choosing its own training objective, acquiring compute, or deploying changes without approval. Engineers still define the experimental boundary. The evidence therefore supports a research-automation workflow, not an unattended system upgrading itself.
Three results would make the launch easier to evaluate: whether third parties reproduce the internal ranking on public tasks, the full cost per completed million-token job, and how much of the 31.8% throughput gain remains on different hardware. TechNode independently confirmed the prices, evaluation size, and distribution channels, but its performance figures also originate with Tencent. External tests that disclose hardware, token consumption, latency, and failed retries can show whether the model’s parameter scale converts into production efficiency that customers can buy.
Sources:
Related Articles
Google Sends TPUs Into Orbit: A 15-Minute Test Is Not Continuous Compute
Google and Planet launch a Suncatcher prototype to test Gemma on four TPUs. Orbital hardware enables thermal and radiation testing, but cooling intervals, sustained compute and commercial costs remain unresolved.
Akamai Wins Anthropic Cloud Deal: CPU Capacity Requires Spending First
Anthropic commits to $11.6 billion of cloud services over seven years, with Akamai estimating $5.5 billion in capital spending. CPU demand has a named buyer, but delivery and revenue still lie ahead.