Qwen3.8-Flash-Next 架构笔记 · REFERENCES
参考文献与资料边界
统一列出专题使用的官方资料、原始论文与实现证据;正文编号在所有章节中保持不变。
参考文献采用专题内固定编号。架构数值与消融结论优先引用 Qwen 官方报告、模型卡和配置;历史演变尽量回到原始论文。笔记中的“我的理解”只用于解释不同部件之间的关系,不替代原文结论。2026-09-05 补入模块发展主线的文献 [40]—[59],并核对已有引用中的 Engram 命名与 batch-size warmup 比较方向。教学例子的数值不来自模型实验。
官方资料
原始论文
实现 / 配置
Qwen3.8 官方资料
- Qwen Team. “Qwen3.8-Flash-Next.” Official Blog, 2026. 原文 官方资料
- Qwen Team. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability. 2026. arXiv v1 全文 · 项目 PDF 官方资料
- Qwen Team. “Qwen3.8-Flash-Next Model Card.” Hugging Face, 2026. 模型卡 官方资料
- Qwen Team. “Qwen3.8-Flash-Next
config.json.” Hugging Face, 2026. 配置 实现 / 配置
Transformer 基线、归一化与前馈网络
- Vaswani et al. “Attention Is All You Need.” NeurIPS, 2017. arXiv:1706.03762
- He et al. “Deep Residual Learning for Image Recognition.” CVPR, 2016. arXiv:1512.03385
- Xiong et al. “On Layer Normalization in the Transformer Architecture.” ICML, 2020. arXiv:2002.04745
- Zhang and Sennrich. “Root Mean Square Layer Normalization.” NeurIPS, 2019. arXiv:1910.07467
- Shazeer. “GLU Variants Improve Transformer.” 2020. arXiv:2002.05202
Attention 与序列混合
- Schlag, Irie, and Schmidhuber. “Linear Transformers Are Secretly Fast Weight Programmers.” ICML, 2021. arXiv:2102.11174
- Yang, Kautz, and Hatamizadeh. “Gated Delta Networks: Improving Mamba2 with Delta Rule.” ICLR, 2025. arXiv:2412.06464
- Su et al. “RoFormer: Enhanced Transformer with Rotary Position Embedding.” Neurocomputing, 2024. arXiv:2104.09864
- DeepSeek-AI et al. “DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.” 2025. arXiv:2512.02556
Embedding 与条件记忆
- Roy et al. “N-Grammer: Augmenting Transformers with Latent N-grams.” 2022. arXiv:2207.06366
- Huang et al. “Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling.” ICML, 2025. arXiv:2501.16975
- Yu et al. “Scaling Embedding Layers in Language Models.” NeurIPS, 2025. arXiv:2502.01637
- Cheng et al. “Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.” 2026. arXiv:2601.07372(论文中的模块名为 Engram。)
MoE
- Shazeer et al. “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.” ICLR, 2017. arXiv:1701.06538
- Fedus, Zoph, and Shazeer. “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.” JMLR, 2022. arXiv:2101.03961
- Dai et al. “DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.” ACL, 2024. arXiv:2401.06066
- Qiu et al. “Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models.” 2025. arXiv:2501.11873
Residual 路径
- Baykal et al. “Alternating Updates for Efficient Transformers.” 2023. arXiv:2301.13310
- Zhu et al. “Hyper-Connections.” 2024. arXiv:2409.19606
- Xie et al. “mHC: Manifold-Constrained Hyper-Connections.” 2025. arXiv:2512.24880
- Qiu et al. “A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling Is Essential for Transformer Training.” 2026. arXiv:2601.22966
MTP 与投机解码
- Gloeckle et al. “Better & Faster Large Language Models via Multi-Token Prediction.” ICML, 2024. arXiv:2404.19737
- Leviathan, Kalman, and Matias. “Fast Inference from Transformers via Speculative Decoding.” ICML, 2023. arXiv:2211.17192
- DeepSeek-AI et al. “DeepSeek-V3 Technical Report.” 2024. arXiv:2412.19437
Optimizer 与分布式实现
- Kingma and Ba. “Adam: A Method for Stochastic Optimization.” ICLR, 2015. arXiv:1412.6980
- Loshchilov and Hutter. “Decoupled Weight Decay Regularization.” ICLR, 2019. arXiv:1711.05101
- Jordan et al. “Muon: An Optimizer for Hidden Layers in Neural Networks.” 2024. 原文
- Liu et al. “Muon Is Scalable for LLM Training.” 2025. arXiv:2502.16982
- Amsel et al. “The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm.” 2025. arXiv:2505.16932
- Wang et al. “Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-Based Optimizers.” 2026. arXiv:2602.06079
Qwen 前代资料
- Qwen Team. “Qwen3-Next: Towards Ultimate Training & Inference Efficiency.” Official Blog, 2025. 原文 官方资料
架构参照模型
- Touvron et al. “LLaMA: Open and Efficient Foundation Language Models.” 2023. arXiv:2302.13971
- Chowdhery et al. “PaLM: Scaling Language Modeling with Pathways.” 2022. arXiv:2204.02311
- Gemma Team. “Gemma 2: Improving Open Language Models at a Practical Size.” 2024. arXiv:2408.00118
实现补充
- Hugging Face Transformers. “Qwen4-Exp Model Implementation.” 2026. 源码 实现 / 配置
发展主线补充:表示、注意力与位置
- Mikolov et al. “Efficient Estimation of Word Representations in Vector Space.” 2013. arXiv:1301.3781
- Sennrich, Haddow, and Birch. “Neural Machine Translation of Rare Words with Subword Units.” 2015 预印本 / ACL 2016. arXiv:1508.07909
- Devlin et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” 2018 预印本 / NAACL 2019. arXiv:1810.04805
- Shazeer. “Fast Transformer Decoding: One Write-Head is All You Need.” 2019. arXiv:1911.02150
- Ainslie et al. “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” 2023. arXiv:2305.13245
- Dao et al. “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.” 2022. arXiv:2205.14135
- Katharopoulos et al. “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.” 2020. arXiv:2006.16236
- Gu and Dao. “Mamba: Linear-Time Sequence Modeling with Selective State Spaces.” 2023. arXiv:2312.00752
- Dao and Gu. “Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.” 2024. arXiv:2405.21060
- Press, Smith, and Lewis. “Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.” 2021. arXiv:2108.12409
- Peng et al. “YaRN: Efficient Context Window Extension of Large Language Models.” 2023. arXiv:2309.00071
发展主线补充:稳定性、条件计算与优化
- Ba, Kiros, and Hinton. “Layer Normalization.” 2016. arXiv:1607.06450
- Wang et al. “DeepNet: Scaling Transformers to 1,000 Layers.” 2022. arXiv:2203.00555
- Lepikhin et al. “GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.” 2020. arXiv:2006.16668
- Cai et al. “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.” 2024. arXiv:2401.10774
- Li et al. “EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.” 2024. arXiv:2401.15077
- DeepSeek-AI. “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.” 2024. arXiv:2405.04434
- Yuan et al. “Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.” 2025. arXiv:2502.11089
- Shazeer and Stern. “Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.” 2018. arXiv:1804.04235
- Gupta, Koren, and Singer. “Shampoo: Preconditioned Stochastic Tensor Optimization.” 2018. arXiv:1802.09568
