全部笔记

Qwen3.8-Flash-Next 架构笔记 · REFERENCES

参考文献与资料边界

统一列出专题使用的官方资料、原始论文与实现证据;正文编号在所有章节中保持不变。

中文草稿 本笔记由 GPT-5.6-Sol 和 GPT-6-Astra 混合撰写。 官方博客 技术报告

参考文献采用专题内固定编号。架构数值与消融结论优先引用 Qwen 官方报告、模型卡和配置;历史演变尽量回到原始论文。笔记中的“我的理解”只用于解释不同部件之间的关系,不替代原文结论。2026-09-05 补入模块发展主线的文献 [40]—[59],并核对已有引用中的 Engram 命名与 batch-size warmup 比较方向。教学例子的数值不来自模型实验。

官方资料 原始论文 实现 / 配置

Qwen3.8 官方资料

  1. Qwen Team. “Qwen3.8-Flash-Next.” Official Blog, 2026. 原文 官方资料
  2. Qwen Team. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability. 2026. arXiv v1 全文 · 项目 PDF 官方资料
  3. Qwen Team. “Qwen3.8-Flash-Next Model Card.” Hugging Face, 2026. 模型卡 官方资料
  4. Qwen Team. “Qwen3.8-Flash-Next config.json.” Hugging Face, 2026. 配置 实现 / 配置

Transformer 基线、归一化与前馈网络

  1. Vaswani et al. “Attention Is All You Need.” NeurIPS, 2017. arXiv:1706.03762
  2. He et al. “Deep Residual Learning for Image Recognition.” CVPR, 2016. arXiv:1512.03385
  3. Xiong et al. “On Layer Normalization in the Transformer Architecture.” ICML, 2020. arXiv:2002.04745
  4. Zhang and Sennrich. “Root Mean Square Layer Normalization.” NeurIPS, 2019. arXiv:1910.07467
  5. Shazeer. “GLU Variants Improve Transformer.” 2020. arXiv:2002.05202

Attention 与序列混合

  1. Schlag, Irie, and Schmidhuber. “Linear Transformers Are Secretly Fast Weight Programmers.” ICML, 2021. arXiv:2102.11174
  2. Yang, Kautz, and Hatamizadeh. “Gated Delta Networks: Improving Mamba2 with Delta Rule.” ICLR, 2025. arXiv:2412.06464
  3. Su et al. “RoFormer: Enhanced Transformer with Rotary Position Embedding.” Neurocomputing, 2024. arXiv:2104.09864
  4. DeepSeek-AI et al. “DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.” 2025. arXiv:2512.02556

Embedding 与条件记忆

  1. Roy et al. “N-Grammer: Augmenting Transformers with Latent N-grams.” 2022. arXiv:2207.06366
  2. Huang et al. “Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling.” ICML, 2025. arXiv:2501.16975
  3. Yu et al. “Scaling Embedding Layers in Language Models.” NeurIPS, 2025. arXiv:2502.01637
  4. Cheng et al. “Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.” 2026. arXiv:2601.07372(论文中的模块名为 Engram。)

MoE

  1. Shazeer et al. “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.” ICLR, 2017. arXiv:1701.06538
  2. Fedus, Zoph, and Shazeer. “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.” JMLR, 2022. arXiv:2101.03961
  3. Dai et al. “DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.” ACL, 2024. arXiv:2401.06066
  4. Qiu et al. “Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models.” 2025. arXiv:2501.11873

Residual 路径

  1. Baykal et al. “Alternating Updates for Efficient Transformers.” 2023. arXiv:2301.13310
  2. Zhu et al. “Hyper-Connections.” 2024. arXiv:2409.19606
  3. Xie et al. “mHC: Manifold-Constrained Hyper-Connections.” 2025. arXiv:2512.24880
  4. Qiu et al. “A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling Is Essential for Transformer Training.” 2026. arXiv:2601.22966

MTP 与投机解码

  1. Gloeckle et al. “Better & Faster Large Language Models via Multi-Token Prediction.” ICML, 2024. arXiv:2404.19737
  2. Leviathan, Kalman, and Matias. “Fast Inference from Transformers via Speculative Decoding.” ICML, 2023. arXiv:2211.17192
  3. DeepSeek-AI et al. “DeepSeek-V3 Technical Report.” 2024. arXiv:2412.19437

Optimizer 与分布式实现

  1. Kingma and Ba. “Adam: A Method for Stochastic Optimization.” ICLR, 2015. arXiv:1412.6980
  2. Loshchilov and Hutter. “Decoupled Weight Decay Regularization.” ICLR, 2019. arXiv:1711.05101
  3. Jordan et al. “Muon: An Optimizer for Hidden Layers in Neural Networks.” 2024. 原文
  4. Liu et al. “Muon Is Scalable for LLM Training.” 2025. arXiv:2502.16982
  5. Amsel et al. “The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm.” 2025. arXiv:2505.16932
  6. Wang et al. “Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-Based Optimizers.” 2026. arXiv:2602.06079

Qwen 前代资料

  1. Qwen Team. “Qwen3-Next: Towards Ultimate Training & Inference Efficiency.” Official Blog, 2025. 原文 官方资料

架构参照模型

  1. Touvron et al. “LLaMA: Open and Efficient Foundation Language Models.” 2023. arXiv:2302.13971
  2. Chowdhery et al. “PaLM: Scaling Language Modeling with Pathways.” 2022. arXiv:2204.02311
  3. Gemma Team. “Gemma 2: Improving Open Language Models at a Practical Size.” 2024. arXiv:2408.00118

实现补充

  1. Hugging Face Transformers. “Qwen4-Exp Model Implementation.” 2026. 源码 实现 / 配置

发展主线补充:表示、注意力与位置

  1. Mikolov et al. “Efficient Estimation of Word Representations in Vector Space.” 2013. arXiv:1301.3781
  2. Sennrich, Haddow, and Birch. “Neural Machine Translation of Rare Words with Subword Units.” 2015 预印本 / ACL 2016. arXiv:1508.07909
  3. Devlin et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” 2018 预印本 / NAACL 2019. arXiv:1810.04805
  4. Shazeer. “Fast Transformer Decoding: One Write-Head is All You Need.” 2019. arXiv:1911.02150
  5. Ainslie et al. “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” 2023. arXiv:2305.13245
  6. Dao et al. “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.” 2022. arXiv:2205.14135
  7. Katharopoulos et al. “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.” 2020. arXiv:2006.16236
  8. Gu and Dao. “Mamba: Linear-Time Sequence Modeling with Selective State Spaces.” 2023. arXiv:2312.00752
  9. Dao and Gu. “Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.” 2024. arXiv:2405.21060
  10. Press, Smith, and Lewis. “Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.” 2021. arXiv:2108.12409
  11. Peng et al. “YaRN: Efficient Context Window Extension of Large Language Models.” 2023. arXiv:2309.00071

发展主线补充:稳定性、条件计算与优化

  1. Ba, Kiros, and Hinton. “Layer Normalization.” 2016. arXiv:1607.06450
  2. Wang et al. “DeepNet: Scaling Transformers to 1,000 Layers.” 2022. arXiv:2203.00555
  3. Lepikhin et al. “GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.” 2020. arXiv:2006.16668
  4. Cai et al. “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.” 2024. arXiv:2401.10774
  5. Li et al. “EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.” 2024. arXiv:2401.15077
  6. DeepSeek-AI. “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.” 2024. arXiv:2405.04434
  7. Yuan et al. “Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.” 2025. arXiv:2502.11089
  8. Shazeer and Stern. “Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.” 2018. arXiv:1804.04235
  9. Gupta, Koren, and Singer. “Shampoo: Preconditioned Stochastic Tensor Optimization.” 2018. arXiv:1802.09568