Navigating China Internet:
DeepSeek V4 launch —
implications for China AI
models/cloud/data centers
领航中国互联⽹:
DeepSeek V4 发布 —— 对
中国 AI 模型/云/数据中⼼
的影响
Compared with the DeepSeek de1ning moment last
year, we view DeepSeek's latest V4 model launch as a
continuation of its focus in computing e?ciency and
open-source approach and highlight DeepSeek V4's
architectural upgrades (vs. V3) that enable: 1)
signi'cantly less memory needed for long context
window, where V4 supports 1M context window that
is comparable to US state-of-the-art (SOTA) models
but at a fraction of memory/KV cache needed (at 7-
10% of V3 model, for its Pro and 284bn
parameter Flash models), thus could facilitate further
proliferation of agentic applications given highly
aXordable costs, and 2) adoption of domestic chips,
where DeepSeek expects signi1cant price cuts for its
Pro models from 2H 2026, following its anticipated
mass availability of Huawei's Ascend 950 super nodes.
Implications for China AI models/cloud/data
centers: We note the rise in competition amongst
Chinese AI models, given the recent accelerated pace
of new model launches against DeepSeek's open-
source V4 (Kimi , Alibaba's -Max,
Tencent's Hy3 preview, Xiaomi , and potential for
MiniMax's M3/Hailuo launches in May, GSe), where
coding/task completion success rates and multi-modal
will likely be key diXerentiating factors for model
pricing power ahead. We continue to rank Cloud &
Data Centers as our #1 preferred sub-sector (with key
ideas: GDS, VNET, Alibaba and Kingsoft Cloud, see
report) on proliferation of AI token demand with
improving cloud pricing driven by enterprise/AI agent
growth and consumer AI assistants. We believe
improvement in computing cost e?ciencies could drive
further room for wider
adoption/exploration/proliferation of AI applications.
与去年 DeepSeek 的⾼光时刻相⽐,我们认为
DeepSeek 最新发布的 V4 模型延续了其对计算效
率和开源路径的专注。我们重点关注 DeepSeek
V4 相⽐ V3 的架构升级,这些升级实现了:1) ⻓
上下⽂窗⼝所需的内存显著减少,V4 ⽀持 100 万
(1M)上下⽂窗⼝,与美国最先进(SOTA)模型
相当,但所需的内存/KV 缓存仅为极⼩部分(其
Pro 模型和 284bn 参数 Flash 模型仅需 V3 模
型的 7-10%),鉴于极具性价⽐的成本,这可能
促进智能体(agentic)应⽤的进⼀步普及;2) 采
⽤国产芯⽚,DeepSeek 预计在华为昇腾 950 超级
节点⼤规模可⽤后,其 Pro 模型将从 2026 年下半
年开始⼤幅降价。对中国 AI 模型/云/数据中⼼的
影响:我们注意到中国 AI 模型之间的竞争⽇益加
剧,近期新模型发布节奏加快,以应对 DeepSeek
的开源 V4(包括 Kimi 、阿⾥巴巴的通义千问
-Max、腾讯混元 Hy3 预览版、⼩⽶
,以及 MiniMax 可能在 5 ⽉发布的 M3/海
螺,⾼盛预测),其中代码编写/任务完成成功率
和多模态能⼒可能是未来模型定价权的关键差异化
因素。我们继续将云与数据中⼼列为⾸选⼦⾏业
(核⼼标的:万国数据、世纪互联、阿⾥巴巴和⾦
⼭云,详⻅报告),原因在于 AI Token 需求激
增,且在企业/AI 智能体增⻓和消费者 AI 助⼿的推
动下,云定价环境有所改善。我们认为计算成本效
率的提升将为 AI 应⽤的更⼴泛采⽤、探索和普及
提供进⼀步空间。
Key highlights of DeepSeek V4:
DeepSeek V4 的核⼼亮点:
DeepSeek released its latest open-sourcing
DeepSeek-V4 Preview on Apr 24, 2026, in two
versions (Pro and Flash), where the Pro version is a
dagship-scale model featuring parameters
(with 49B activated), while the Flash version is
relatively small with 284 billion parameters (and 13B
activated).
DeepSeek 于 2026 年 4 ⽉ 24 ⽇发布了其最新
的开源预览版 DeepSeek-V4,包含两个版本
(Pro 和 Flash)。其中 Pro 版本是旗舰级模
型,拥有 万亿参数(激活参数为 490 亿);
⽽ Flash 版本规模相对较⼩,拥有 2840 亿参数
(激活参数为 130 亿)。
We note DeepSeek continues to be foundation text
focused, while internet mega-caps and other
independent model players like MiniMax are more
multi/full-modal focus in their model approach, in
their ongoing pursuit of AGI.
我们注意到,DeepSeek 继续专注于基础⽂本模
型,⽽互联⽹巨头以及 MiniMax 等其他独⽴模
型⼚商在追求通⽤⼈⼯智能(AGI)的过程中,
其模型路径更倾向于多模态或全模态。
1M context window for both Pro and Flash
models, via new architectural breakthroughs:
Both Pro and Flash models support an ultra-long
context window of 1M tokens, with substantially
lower inference FLOPS and KV-cache footprint vs.
, helped by 3 key upgrades. Besides
DeepSeek MoE and Multi-Token Prediction,
DeepSeek V4 incorporates 3 key upgrades in
architecture and optimizations:
通过全新的架构突破,Pro 和 Flash 模型均⽀
持 100 万(1M)上下⽂窗⼝:得益于三项关
键升级,Pro 和 Flash 模型均⽀持 100 万
token 的超⻓上下⽂窗⼝,且与 DeepSeek-
相⽐,推理算⼒(FLOPS)需求和 KV 缓
存占⽤⼤幅降低。除了 DeepSeek MoE 和多
token 预测(Multi-Token Prediction)之
外,DeepSeek V4 在架构和优化⽅⾯还引⼊了
三项核⼼升级:
Hybrid attention architecture combines CSA
and HCA. Compressed Sparse Attention (CSA)
1rst compresses the KV cache along the sequence
dimension and then applies sparse attention,
while Heavily Compressed Attention (HCA)
applies more aggressive KV compression but
keeps dense attention. As a result, V4 reduces the
amount of temporary memory the model needs
to keep for long inputs.
混合注意⼒架构结合了 CSA 和 HCA。压缩
稀疏注意⼒(CSA)⾸先沿序列维度压缩 KV
缓存,然后应⽤稀疏注意⼒;⽽重度压缩注
意⼒(HCA)则采⽤更激进的 KV 压缩⽅
式,但保持稠密注意⼒。因此,V4 减少了模
型在处理⻓输⼊时需要保留的临时内存量。
mHC improves training reliability: DeepSeek-V4
adds mHC to make the model more stable as
information passes through many layers.
mHC 提升了训练可靠性:DeepSeek-V4 加
⼊了 mHC,使得模型在信息经过多层传递时
更加稳定。
Muon optimizer: DeepSeek-V4 uses Muon as its
main training optimizer, while still using AdamW
for some smaller/special modules. In simple
terms, Muon is the training method that helps the
model learn in a more stable way as V4 has a more
complex architecture than V3.
Muon 优化器:DeepSeek-V4 使⽤ Muon 作
为主要训练优化器,同时在⼀些较⼩或特殊
的模块中仍沿⽤ AdamW。简单来说,由于
V4 的架构⽐ V3 更复杂,Muon 这种训练⽅
法能帮助模型以更稳定的⽅式进⾏学习。
These updates result in meaningful eRciency
gains versus . At 1M context, V4-
Pro reportedly uses 27% of DeepSeek ’s
single-token inference FLOPs and 10% of its KV
cache; V4-Flash is reported at 10% FLOPs and
7% KV cache. This suggests e?ciency gains for
long-context workloads, where the likely use-case
relevance is long-horizon tasks rather than
ordinary short prompts.
这些更新带来了相较于 显
著的效率提升。据报道,在 1M 上下⽂⻓度
下,V4-Pro 的单 token 推理 FLOPs 仅为
DeepSeek 的 27%,KV cache 仅为
10%;V4-Flash 的 FLOPs 更是低⾄ 10%,
KV cache 仅为 7%。这表明⻓上下⽂⼯作负
载的效率得到了⼤幅提升,其应⽤场景更倾
向于⻓程任务,⽽⾮普通的短提示词。
Implications for China AI models/cloud/data
centers:
对中国 AI 模型/云服务/数据中⼼的影响:
With DeepSeek-V4's competitive model capabilities,
we continue to rank Cloud & Data Centers as our
preferred sub-sector (#1) (with key ideas: GDS,
VNET, Alibaba and Kingsoft Cloud) on continued
proliferation of AI token demand with improving
cloud/token pricing power driven by enterprise/AI
agent growth and consumer AI assistants. We
believe further improvement in computing cost
e?ciencies could drive further room for wider
adoption/exploration/proliferation of AI
applications.
凭借 DeepSeek-V4 极具竞争⼒的模型能⼒,
我们继续将云与数据中⼼列为⾸选⼦⾏业(排
名第⼀)(核⼼标的包括:万国数据、世纪互
联、阿⾥巴巴和⾦⼭云)。这主要基于企业级/
AI 智能体增⻓以及消费者 AI 助⼿推动下,AI
Token 需求持续激增,且云/Token 定价能⼒不
断提升。我们认为,计算成本效率的进⼀步提
⾼,将为 AI 应⽤的更⼴泛采⽤、探索和普及提
供更⼤空间。
For AI models, given the recent accelerated pace of
new model launches against DeepSeek's open-
source V4 (Kimi , Alibaba's -Max,
Tencent's Hy3 preview, Xiaomi , and potential
for MiniMax's M3/Hailuo launches in May, GSe),
where coding (noting Knowledge Atlas/Zhipu's GLM
model is highly ranked), task completion success
rates and multi-modal (ByteDance, Alibaba and
MiniMax) will likely be key diXerentiating factors for
model pricing power ahead.
在 AI 模型⽅⾯,鉴于近期新模型发布节奏加
快,正⾯对标 DeepSeek 的开源 V4(包括
Kimi 、阿⾥巴巴的 -Max、腾讯
混元 预览版、⼩⽶ ,以及 MiniMax
可能在 5 ⽉发布的 M3/海螺模型,⾼盛预
测),代码能⼒(注:智谱 GLM 模型排名靠
前)、任务完成成功率以及多模态能⼒(字节
跳动、阿⾥巴巴和 MiniMax)可能成为未来模
型定价权的关键差异化因素。
We believe independent players' competitive
advantages vs. internet mega-caps rest on their high
organization e?ciency and decision-making
processes in identifying next key AI model
developments, ., MiniMax's highly e?cient model
design and inference enables the company to
achieve 40% GPM (GSe) for its foundation text API
channel, even at highly competitive API pricing
levels.
我们认为,独⽴⼚商相对于互联⽹巨头的竞争
优势在于其极⾼的组织效率和决策流程,能够
快速识别下⼀个关键的 AI 模型发展⽅向。例
如,MiniMax ⾼效的模型设计和推理能⼒使其
基础⽂本 API 渠道即使在极具竞争⼒的 API 定
价⽔平下,仍能实现 40% 的⽑利率(⾼盛预
测)。
For internet mega-caps, we believe they are best
positioned to capture the AI infrastructure/cloud
opportunity given strong underlying operating cash
dows from their core business, while separate
standalone incentive schemes for AI chip/model
teams (., Doubao AI team already has separately
incentive programs) will likely be needed to
incentivise/retain top AI talent vs. independent AI
native players.
对于互联⽹巨头⽽⾔,我们认为凭借其核⼼业
务产⽣的强⼤底层经营现⾦流,他们最适合捕
捉 AI 基础设施/云市场的机遇。同时,为了与
独⽴ AI 原⽣⼚商竞争并吸引/留住顶尖 AI ⼈
才,巨头们可能需要为 AI 芯⽚/模型团队设⽴
独⽴的激励⽅案(例如,⾖包 AI 团队已经拥有
独⽴的激励计划)。
We also notenews reports that Tencent and Alibaba
are potentially in talks to invest in DeepSeek at over
US$20bn valuation, compared with last close market
caps for Knowledge Atlas (Zhipu)/MiniMax at
US$53bn/31bn.
我们还注意到有新闻报道称,腾讯和阿⾥巴巴
可能正洽谈以超过 200 亿美元的估值投资
DeepSeek,⽽智谱 AI (Knowledge Atlas) 和
MiniMax 的最近⼀轮市场估值分别为 53 亿美
元和 31 亿美元。
Related research: 相关研究:
Navigating China Internet: What to do from here & key
focuses post results season; addressing key AI debates, April
1, 2026
中国互联⽹⾏业导航:后续操作建议及业绩季后
的核⼼关注点;应对 AI 关键争论,2026 年 4 ⽉ 1
⽇
Tencent Holdings (): Hy3 preview marks a key
development in Tencent's AI revamp; Buy, April 23, 2026
腾讯控股 ():2026 年上半年预览标志着
腾讯 AI 架构重组的关键进展;买⼊,2026 年 4
⽉ 23 ⽇
Xiaomi Corp. (): series release;
Accelerating model update pace paves way for
commercialization and expanding AI applications; Buy,
April 24, 2026
⼩⽶集团 (): 系列发布;加
速模型更新步伐,为商业化和扩⼤ AI 应⽤铺平道
路;买⼊,2026 年 4 ⽉ 24 ⽇
MiniMax Group (): Global full-modal AI company
reaching hypergrowth stage; initiate at Neutral on
valuation, Feb 23, 2026
MiniMax 集团 ():全球全模态 AI 公司进
⼊⾼速增⻓阶段;基于估值给予中性评级,2026
年 2 ⽉ 23 ⽇
The authors would like to thank Iris Xiao for her
contribution to this report.
作者感谢 Iris Xiao 对本报告的贡献。
Exhibit 1: Our quarterly updated sub-sector preference
within China Internet (Apr 2026 update)
图表 1:我们每季度更新的中国互联⽹⼦⾏业偏好
(2026 年 4 ⽉更新)
Source: Goldman Sachs Global Investment Research
来源:⾼盛全球投资研究部
DeepSeek-V4 pricing/performance comparison
with counterparts
DeepSeek-V4 与同类产品的价格/性能对⽐
Exhibit 2: DeepSeek V4 continues to show pricing
competitiveness with V4 Pro price expected to further
decline with Huawei-backed domestic compute
capacity ramps in 2H26
图表 2:DeepSeek V4 持续展现价格竞争⼒,随着
华为⽀持的国内算⼒在 2026 年下半年逐步释放,预
计 V4 Pro 的价格将进⼀步下降
Source: Arti1cial Analysis, Company data, Data compiled by Goldman
Sachs Global Investment Research
Exhibit 3: DeepSeek V4 pushes 1M context with better
cost eOciency amongst open-source models
Mimo Pro to be open-sourced soon
Source: Company data, OpenRouter, Data compiled by Goldman Sachs
Global Investment Research
Exhibit 4: DeepSeek V4 Pro and V4 Flash leads amongst
open-sourced models on GDPval-AA
Source: Arti1cial Analysis
Exhibit 5: Key players across diTerent levels: Alibaba
leads on external AI cloud revenue scale serving To-B
enterprises, while ByteDance is currently largest in AI
To-C chatbot/total daily token usage
Source: Respective company data, Data compiled by Goldman Sachs
Global Investment Research
DeepSeek-V4 technical details
Exhibit 6: DeepSeek V4 marks a step-change in long-
text eOciency
Source: Company data, Data compiled by Goldman Sachs Global
Investment Research
Exhibit 7: Overall architecture of DeepSeek-V4 series,
where DeepSeek used hybrid CSA (Compressed Sparse
Attention) and HCA (Heavily Compressed Attention) for
attention layers, DeepSeekMoE for feed-forward
layers, and strengthen conventional residual
connections with mHC
Source: Company data
Exhibit 8: Inference FLOPs and KV cache size of
DeepSeek-V4 series and
Source: Company data
Exhibit 9: Benchmark performance of DeepSeek-V4-
Pro-Max and its US counterparts
Source: Deepseek
Latest API tokens tracker: Chinese players
continued to gain market share
Exhibit 10: Chinese AI models have moved up the ranks
amongst Top 15 on OpenRouter in terms of token usage
this month (Red denotes Chinese models)
Source: OpenRouter
Exhibit 11: Chinese players have continued to gain
market share through 2026
Source: OpenRouter, Data compiled by Goldman Sachs Global Investment
Research
To-C AI applications: Doubao remains the
dominant leader
Exhibit 12: AIGC to-C / Chatbot engagement increased
+36% mom in Mar
Source: QuestMobile
Exhibit 13: Chinese AIGC apps DAU trends
PDF Share Bookmark More
24 April 2026 | 9:01PM HKT | Research | Equity| By Ronald
Keung, CFA and others
2026 年 4 ⽉ 24 ⽇ | 晚上 9:01 HKT | 研究 | 股票 | 作者:
Ronald Keung, CFA 等
Goldman Sachs Research ⾼盛研究部 Market Insights 市场洞察
2026/4/26 17:20
⻚码# 1/1