AI Cognitive Map · 2026 修订版Revised 2026

从 Token 到线缆认知引擎与物理基础设施:一张 AI 技术栈全景地图 From Token to WireCognitive engines and physical infrastructure: a map of the AI stack

屏幕上蹦出的每一个字,背后都串着一整条技术链:分词、模型、检索、GPU、互联网络。本文把这条链拆成七层,先从最小的单位 Token 讲起,再逐层下沉到算力与网络。读完后,你应该能判断一个 AI 系统“答错了、变慢了、太贵了”时,问题大概出在哪一层。 Every word an AI types on screen rides on a whole chain of technology: tokenization, models, retrieval, GPUs and interconnect. This map splits that chain into seven layers, starting from the smallest unit, the token, and descending to compute and networking. By the end you should be able to tell which layer is likely at fault when an AI system is wrong, slow or expensive.

初版写于 2026 年,2026 年 10 月修订First written 2026 · revised Oct 2026 硬件数据为厂商公开口径Hardware figures are vendor-published 来源见文末Sources at the end

🗺️ AI 技术栈全景架构图AI Technology Stack — Full Architecture

上层依赖下层;点击节点跳到对应章节Upper layers depend on lower ones; click a node to jump

每层只列代表性技术,并非完整清单。正文按“理解顺序”而不是层序展开:先 Token(L4),再模型(L5)、应用(L6–L7),最后下沉到算力与网络(L3–L1)。Each layer lists representative technologies only. The text follows the order of understanding rather than layer order: tokens (L4), then models (L5), applications (L6–L7), then down to compute and network (L3–L1).

L4 · 数据层Data

Token 与向量表示Tokens and vector representations

The atomic unit of AI language

要理解后面的一切,先要知道模型“看到”的是什么:不是文字,而是一串编号;编号再被换成向量,含义才变成可以计算的几何关系。 Before anything else, it helps to know what a model actually "sees": not words but a sequence of IDs. Those IDs become vectors, and only then does meaning turn into geometry that can be computed.

01 · Tokenization什么是 Token?What is a token?

Token 是大语言模型处理文本的基本单位。文本输入模型前要先经过 Token 化(Tokenization):分词器按自己的词表把文本切成片段——可能是一个词、一个词根(subword)、一个汉字或几个字节——再把每个片段映射成整数 ID。

A token is the basic unit an LLM processes. Before text enters the model it is tokenized: a tokenizer splits it according to its vocabulary—into a word, a subword, a single Chinese character or a few bytes—and maps each piece to an integer ID.

例如英文 unbelievable 可能被切成 un + believ + able。同一句话在不同模型里的 Token 数并不相同,这也是 API 按 Token 计费、上下文窗口按 Token 计算的原因。

For example, unbelievable might become un + believ + able. The same sentence yields different token counts in different models, which is why APIs bill—and context windows are sized—in tokens.

BPEWordPieceSentencePiece

02 · Embeddings向量嵌入Embeddings

整数 ID 本身没有含义。模型的嵌入层把每个 ID 查表换成一个高维向量(常见为数千维),语义相近的 Token 在训练后会落在空间中相近的位置。这些向量是注意力计算的输入。

An integer ID carries no meaning by itself. The model's embedding layer looks each ID up and returns a high-dimensional vector (often thousands of dimensions); after training, semantically similar tokens sit close together. These vectors are the input to attention.

注意区分两类“嵌入”:LLM 内部的 Token 嵌入,与 RAG 检索使用的文本嵌入——后者通常由专门的嵌入模型把一整段文字压成一个向量。

Keep two kinds apart: the token embeddings inside an LLM, and the text embeddings used for RAG retrieval, which a dedicated embedding model produces by compressing a whole passage into one vector.

03 · Vector database向量数据库Vector databases

专门存储和查询文本嵌入的数据库。传统数据库做关键字的精确匹配,向量数据库做语义相似性搜索:用余弦相似度或内积找出与查询向量最接近的文档片段。

A database built to store and query text embeddings. Where a traditional database matches keywords exactly, a vector database performs semantic similarity search, using cosine similarity or inner product to find the passages closest to a query vector.

为了在百万、亿级向量中毫秒级返回,它们通常使用 HNSW、IVF 等近似最近邻(ANN)索引,用极少的召回损失换取速度。它是多数 RAG 系统的检索基座,也常与关键词检索混合使用。

To answer in milliseconds over millions or billions of vectors, they use approximate nearest neighbour (ANN) indexes such as HNSW or IVF, trading a little recall for speed. They are the retrieval foundation of most RAG systems and are often combined with keyword search.

L5 · 模型层Model

大语言模型(LLM)Large language models

From next-token prediction to general capability

规模带来的“涌现能力”"Emergent abilities" from scale

随着参数量和训练数据扩大,LLM 不只会翻译、摘要,还能做多步推理,并能通过上下文学习(In-context Learning)完成从未专门训练过的任务——只需在提示里给几个示例。Wei 等人(2022)把这种随规模“突然出现”的能力称为涌现能力。

As parameters and training data grow, LLMs go beyond translation and summarisation to multi-step reasoning and in-context learning: solving tasks they were never explicitly trained for from a few examples in the prompt. Wei et al. (2022) called abilities that seem to appear suddenly with scale emergent abilities.

不过“突然”二字存在争议:Schaeffer 等人(2023)指出,部分跃变来自评测指标的选择,换成连续指标后能力提升更平滑。可以确定的是能力随规模持续增强;它是否等同于“类人认知”,目前没有定论。

"Sudden" is contested, though: Schaeffer et al. (2023) showed that some jumps stem from the choice of metric and look smooth under continuous metrics. What is clear is that capability keeps rising with scale; whether that amounts to human-like cognition remains open.

Transformer 架构与注意力机制Transformer architecture and attention

今天的主流 LLM(GPT、Gemini、Llama 等系列)都基于 Transformer 架构,核心是自注意力(Self-Attention):处理某个 Token 时,模型为序列中其他 Token 计算相关性权重,据此汇聚上下文信息,从而捕捉长距离依赖。

Today's mainstream LLMs (the GPT, Gemini and Llama families, among others) are built on the Transformer, whose core is self-attention: when processing a token, the model computes relevance weights over the other tokens in the sequence and aggregates context accordingly, capturing long-range dependencies.

生成式模型大多采用仅解码器(Decoder-only)结构,并使用因果注意力:每个 Token 只能看到它之前的 Token,看不到“未来”。模型据此输出整个词表上的概率分布,选出下一个 Token,把它接回输入,再预测下一个。

Generative models are mostly decoder-only and use causal attention: each token can see only the tokens before it, never the "future". The model outputs a probability distribution over the whole vocabulary, picks the next token, appends it to the input and predicts again.

  • 多头注意力(Multi-Head Attention):多组注意力并行计算,各自捕捉不同类型的关系。
  • Multi-head attention: several attention heads run in parallel, each capturing a different kind of relationship.
  • KV Cache(键值缓存):缓存已处理 Token 的 Key/Value 向量,生成新 Token 时不必重算前文;代价是显存占用随上下文长度增长。
  • KV cache: stores the key/value vectors of tokens already processed so earlier context is not recomputed for each new token; the cost is memory that grows with context length.
  • 自回归生成:每一步只产出一个 Token,循环直到结束符或长度上限。
  • Autoregressive generation: one token per step, looping until an end token or a length limit.
模型的生命周期:从预训练到偏好对齐Model lifecycle: from pre-training to alignment

数据准备Data preparation

收集海量文本与代码;去重、过滤低质与有害内容、处理隐私信息。

Collect large volumes of text and code; deduplicate, filter low-quality and harmful content, handle personal data.

预训练Pre-training

以“预测下一个 Token”为目标的自监督学习,获得广泛的语言与世界知识。

Self-supervised learning with a next-token objective, acquiring broad linguistic and world knowledge.

→ 基座模型→ base model

监督微调(SFT)Supervised fine-tuning

用高质量“指令—回答”样本教模型遵循指令、以对话方式回应。

High-quality instruction–response pairs teach the model to follow instructions and respond conversationally.

偏好对齐Preference alignment

RLHF(基于人类反馈的强化学习)或 DPO 等方法,使输出更有用、诚实、无害。

RLHF (reinforcement learning from human feedback) or methods such as DPO make outputs more helpful, honest and harmless.

→ 对话模型→ chat model

推理型模型通常还会追加一段强化学习,用可自动验证的奖励(如数学答案、代码测试)训练长链推理。Reasoning models typically add a further reinforcement-learning stage with automatically verifiable rewards (maths answers, code tests) to train long chains of reasoning.

推理流水线:从提示词到逐字输出Inference pipeline: from prompt to streamed output

屏幕上的“打字机效果”并非刻意放慢:模型确实是一个 Token 一个 Token 算出来的,服务端算出一个就流式推送一个。一次请求会经历两个性质不同的阶段,分别决定两个用户能感知的延迟指标。

The "typewriter effect" is not an artificial delay: the model really does compute one token at a time, and the server streams each one as soon as it exists. A request passes through two phases of very different character, each governing a latency metric users can feel.

Token 化Tokenize

分词器把提示词切成 Token 并转换为 ID。

The tokenizer splits the prompt into tokens and IDs.

预填充(Prefill)Prefill

并行处理全部输入 Token,建立 KV Cache;以计算为瓶颈。

Processes all input tokens in parallel and builds the KV cache; compute-bound.

决定首 Token 延迟 TTFTsets time-to-first-token (TTFT)

解码(Decode)Decode

自回归循环,每步生成 1 个 Token,需反复读取权重与 KV Cache;以显存带宽为瓶颈。

Autoregressive loop, one token per step, re-reading weights and KV cache each time; memory-bandwidth-bound.

决定 Token 间延迟 ITL / TPOTsets inter-token latency (ITL / TPOT)

去 Token 化Detokenize

把生成的 ID 转回文字,流式返回给用户。

Turns generated IDs back into text and streams it to the user.

每一步在硬件上发生了什么What happens in hardware at each step

Kubernetes 等调度器在部署时把推理服务放到合适的 GPU 上(不参与逐 Token 的计算);每生成一个 Token,CUDA 内核在 GPU 上执行大量矩阵乘法;若模型被切分到多张 GPU(张量并行),NCCL 还要在每一层通过 NVLink 做 AllReduce / AllGather 来同步中间结果。这就是为什么后文的显存带宽和互联带宽会直接体现在“字蹦得快不快”上。

A scheduler such as Kubernetes places the inference service on suitable GPUs at deployment time (it is not involved per token). For every token, CUDA kernels run large matrix multiplications on the GPU; if the model is split across GPUs (tensor parallelism), NCCL also performs AllReduce / AllGather over NVLink at every layer to synchronise intermediate results. That is why memory and interconnect bandwidth, covered later, show up directly in how fast words appear.

如何评价 LLM:主流基准与指标Evaluating LLMs: common benchmarks and metrics
类别Category 基准 / 指标Benchmark / metric 测什么What it measures
知识与综合Knowledge MMLU / MMLU-Pro / GPQA MMLU 覆盖 57 个学科的选择题;头部模型已接近饱和,因此出现了更难的 MMLU-Pro、GPQAMMLU spans 57 subjects of multiple-choice questions; top models are near saturation, hence harder successors such as MMLU-Pro and GPQA
语言理解Language understanding GLUE / SuperGLUE 经典自然语言理解任务集;对当今 LLM 已基本饱和,多用于历史对比Classic NLU task suites; largely saturated by current LLMs and mostly used for historical comparison
数学推理Maths GSM8K / MATH 小学应用题到竞赛级数学题Grade-school word problems to competition-level maths
代码生成Code HumanEval / MBPP 根据描述写函数并通过单元测试(Pass@k)Write functions from descriptions that pass unit tests (Pass@k)
人类偏好Human preference Chatbot Arena 真实用户对两个匿名模型的回答做盲评,汇总为排名Real users blind-compare two anonymous models; votes aggregate into a ranking
文本相似度Text overlap BLEU / ROUGE / BERTScore 与参考答案的重合程度:BLEU、ROUGE 看 n-gram 重叠,BERTScore 看语义相似;适合翻译、摘要,不衡量事实正确性Overlap with a reference: BLEU and ROUGE count n-grams, BERTScore measures semantic similarity; suited to translation and summaries, not to factual correctness

公开基准会饱和,也可能混入训练数据(数据污染)。评估自己的场景时,最可靠的仍是用业务数据构建的私有测试集。Public benchmarks saturate and can leak into training data (contamination). For your own use case, a private test set built from real data remains the most reliable signal.

L6–L7 · 认知引擎与应用Cognitive engine & apps

检索增强生成(RAG)与 Agentic AIRetrieval-augmented generation and agentic AI

From passive retrieval to autonomous research

为什么需要 RAG?Why RAG?

单靠模型参数里的知识有三个硬伤:会编造看似合理的内容(幻觉);知识停留在训练截止日期;看不到企业内部的私有数据。RAG 在生成前先从外部知识库检索相关片段,连同问题一起交给模型,让回答“有据可查”。它能显著缓解这些问题,但不能根除——检索错了,模型照样会答错。

Relying only on knowledge stored in parameters has three hard limits: models invent plausible content (hallucination), their knowledge stops at the training cutoff, and they cannot see private enterprise data. RAG retrieves relevant passages from an external knowledge base before generating and hands them to the model with the question, so answers can be traced to evidence. It mitigates these problems substantially but does not eliminate them: if retrieval is wrong, the answer will be too.

Stage 1Naive RAG

经典的单轮“检索—阅读”管线:文档切块 → 向量化 → 存入向量库;查询向量化 → 召回相似度最高的 Top-K 片段 → 拼进提示词 → LLM 生成。对单点事实问答有效,但遇到跨文档、多跳推理时,容易召回噪音或漏掉关键证据。

The classic single-pass retrieve-then-read pipeline: chunk documents → embed → store; embed the query → retrieve the top-K most similar chunks → add them to the prompt → generate. It works for single-fact questions but, for cross-document or multi-hop questions, tends to retrieve noise or miss key evidence.

IndexingRetrievalGeneration

Stage 2Advanced RAG

在检索前后加优化环节:查询重写消除歧义 → 关键词与向量的混合检索 → 重排序模型(Reranker)精排 → 上下文合并与裁剪。复杂问题下的检索精度明显提高。

Adds optimisation around retrieval: query rewriting to remove ambiguity → hybrid search combining keywords and vectors → a reranker for fine ordering → context merging and pruning. Retrieval precision on complex questions improves markedly.

Query RewriteHybrid SearchReranker

Stage 3Agentic RAG

让 LLM 充当调度中枢(如 ReAct 的“思考—行动—观察”循环),不再走固定流水线:拆解子问题 → 按需调用工具(向量库、网络搜索、SQL、计算器)→ 判断证据是否充分 → 不够就换策略再检索 → 综合作答。可配合持久记忆层(如 Mem0)保留跨会话上下文。

The LLM acts as the orchestrator (for example ReAct's think–act–observe loop) instead of following a fixed pipeline: decompose the question → call tools as needed (vector store, web search, SQL, calculator) → judge whether evidence suffices → retry with a new strategy if not → synthesise. A persistent memory layer (such as Mem0) can carry context across sessions.

代价是更高的延迟与成本、更难预测和评估的行为。它是当前的主要发展方向,但简单问答场景并不一定需要它。

The price is higher latency and cost and behaviour that is harder to predict and evaluate. It is the main direction of travel, but simple Q&A rarely needs it.

ReActTool CallingMemorySelf-check
案例:Google NotebookLM,从问答工具到研究助手Case study: Google NotebookLM, from Q&A tool to research assistant

核心设计是“严格溯源(source-grounding)”:对话回答基于用户加入笔记本的资料,并以行内引用指回原文片段,方便核对。下面几项功能展示了 RAG 如何从“查资料”走向“做研究、做内容”。

Its core design is source-grounding: chat answers are based on the sources the user adds to a notebook, with inline citations pointing back to the original passages for checking. The features below show RAG moving from looking things up to doing research and producing content.

  • 深度研究(Deep Research):先拟定研究计划,再多轮检索网页、调整关键词,把找到的来源导入笔记本并生成带引用的报告——这正是上文 Agentic RAG 的产品化。
  • Deep Research: drafts a research plan, searches the web over several rounds while refining queries, imports the sources it finds into the notebook and writes a cited report—Agentic RAG turned into a product.
  • 音频概览(Audio Overviews):把资料变成两位主持人的对话播客。按 Google DeepMind 的介绍,其语音模型在含真实口语停顿与语气词(如“嗯”“啊”)的对话数据上微调,可在单颗 TPU v5e 上不到 3 秒生成约 2 分钟对话,比实时快 40 倍以上。
  • Audio Overviews: turn sources into a two-host podcast conversation. According to Google DeepMind, the speech model was fine-tuned on dialogue containing natural disfluencies ("umm", "aah") and can generate about two minutes of dialogue in under three seconds on a single TPU v5e—over 40× faster than real time.
  • 视频概览(Video Overviews):生成带旁白的幻灯片式讲解,从资料中提取图片、图表、引文与数字来辅助说明。
  • Video Overviews: narrated, slide-style explainers that pull images, diagrams, quotes and numbers from the sources.
  • 互动模式(Interactive Mode):收听音频概览时可以“插话”提问,主持人基于资料作答后再回到主线。
  • Interactive mode: listeners can "call in" with a question mid-podcast; the hosts answer from the sources and return to the main thread.

NotebookLM 功能更新频繁,以上描述以撰写时的公开介绍为准。NotebookLM changes often; the description reflects public announcements at the time of writing.

RAG 怎么评估:核心指标与基准Evaluating RAG: core metrics and benchmarks

评估 RAG 要把检索质量和生成质量分开看,否则无法判断错误出在哪一环。RAGAS 框架用 LLM 充当评审,把这两部分拆成四个常用指标:

RAG evaluation must separate retrieval quality from generation quality; otherwise you cannot tell which stage failed. The RAGAS framework uses an LLM as judge and splits the two into four common metrics:

检索侧Retrieval
上下文精确度Context precision
召回的片段里有多少真正相关、排得靠前How much of what was retrieved is relevant, and ranked high
检索侧Retrieval
上下文召回率Context recall
回答所需的信息是否都被找回(需参考答案)Whether all needed information was found (needs a reference answer)
生成侧Generation
忠实度Faithfulness
回答中的陈述能否由检索到的上下文支持Whether claims in the answer are supported by the retrieved context
生成侧Generation
回答相关性Answer relevance
回答是否切题、完整Whether the answer addresses the question

忠实度和回答相关性无需人工标注的参考答案即可计算,这是 RAGAS 降低评估成本的关键;上下文召回率仍需要参考答案。Faithfulness and answer relevance need no human-labelled reference, which is how RAGAS lowers evaluation cost; context recall still requires one.

CRAG 与 LegalBench-RAG 说明了什么What CRAG and LegalBench-RAG show

CRAG(Meta,2024):最先进的 LLM 在不检索时准确率不超过 34%;简单接入 RAG 仅提升到 44%;当时业界最好的 RAG 方案也只能在无幻觉前提下答对 63% 的问题。

CRAG (Meta, 2024): the most advanced LLMs reach at most 34% accuracy without retrieval; adding straightforward RAG lifts this only to 44%; the best industry RAG solutions of the time answered just 63% of questions without hallucination.

LegalBench-RAG(2024)是第一个专门评估 RAG 检索环节的法律领域基准:6,858 个由法律专家标注的问答对,要求从 7,900 万字符的语料中精确定位最小相关片段。作者的出发点是:检索不准,会让上下文变长变贵,并诱发模型遗忘或幻觉。

LegalBench-RAG (2024) is the first legal-domain benchmark dedicated to the retrieval step of RAG: 6,858 query–answer pairs annotated by legal experts, requiring precise location of minimal relevant snippets in a 79-million-character corpus. Its premise: imprecise retrieval makes context longer and costlier and induces the model to forget or hallucinate.

两者共同指向一个实用结论:在专业领域,检索质量往往决定 RAG 的上限。很多被归为“模型幻觉”的错误,起因是没找到或找错了证据——排查时应先查检索,再怪模型。

Together they point to a practical conclusion: in specialised domains, retrieval quality often sets the ceiling for RAG. Many errors blamed on "model hallucination" start with evidence that was missed or wrong, so debug retrieval before blaming the model.

L5–L7 · 物理 AIPhysical AI

物理 AI 与 NVIDIA CosmosPhysical AI and NVIDIA Cosmos

World foundation models for embodied intelligence

从文本走向三维空间From text to three-dimensional space

如果说 LLM 与 Agentic RAG 处理的是数字世界里的文字和逻辑,下一步就是让 AI 理解并作用于物理世界。自动驾驶、人形机器人和工业自动化需要对重力、光照、物体恒存和时空因果有可靠的“直觉”;而真实世界的数据昂贵、危险场景难以采集。世界基础模型(World Foundation Model, WFM)的思路是:先学会“世界如何演化”,再用它生成和评估训练数据。

If LLMs and agentic RAG handle text and logic in the digital world, the next step is AI that understands and acts in the physical one. Autonomous driving, humanoid robots and industrial automation need a reliable "intuition" for gravity, lighting, object permanence and cause and effect over time, yet real-world data is expensive and dangerous scenarios are hard to capture. The idea of a world foundation model (WFM) is to learn how the world evolves first, then use that to generate and evaluate training data.

GenerateCosmos Predict

世界生成与未来状态预测。根据文本、图像或一段视频,生成场景接下来如何演化的视频。第一代同时提供扩散与自回归两类模型(参数规模约 40 亿至 140 亿),后续版本以扩散模型为主。自动驾驶团队可以借此推演“如果这样避让,周围会发生什么”。

World generation and future-state prediction. From text, an image or a video clip, it generates video of how the scene evolves next. The first generation offered both diffusion and autoregressive models (roughly 4B to 14B parameters); later versions centre on diffusion. AV teams can use it to play out what happens around the car after a given manoeuvre.

TransferCosmos Transfer

可控的跨域转换,缩小仿真与现实的差距(Sim-to-Real Gap)。以分割图、深度图、边缘图等结构化信号为条件,把 3D 仿真画面转换成具有真实光照与纹理的照片级视频,并可派生出“雨夜”“黄昏”等环境变体,在保持场景几何不变的前提下大量扩充合成数据。

Controllable domain transfer that narrows the sim-to-real gap. Conditioned on structured signals such as segmentation, depth and edge maps, it turns 3D simulation renders into photorealistic video with real-world lighting and texture, and can derive variants such as "rainy night" or "dusk", multiplying synthetic data while keeping scene geometry fixed.

ReasonCosmos Reason

物理常识与推理。一个针对物理世界常识训练的视觉语言模型(VLM),通过思维链(Chain-of-Thought)分析视频中的因果关系,判断“接下来合理的动作是什么”。常用于给训练数据自动打标、筛选,或作为机器人与视觉智能体的规划模块。

Physical common sense and reasoning. A vision-language model (VLM) trained for physical-world common sense, using chain-of-thought to analyse cause and effect in video and decide what a reasonable next action would be. It is used to auto-label and curate training data, or as the planning component of robots and vision agents.

支撑这些模型的还有 Cosmos Tokenizer:把视频在时间和空间上同时压缩成 Token,使模型能以可承受的成本处理长视频。另需注意:NVIDIA 已发布将推理与生成统一在一个模型中的 Cosmos 3,产品线划分可能继续变化。Underpinning these models is the Cosmos Tokenizer, which compresses video into tokens across both time and space so long clips can be processed at manageable cost. Note also that NVIDIA has since released Cosmos 3, which unifies reasoning and generation in one model, so the product lineup may keep changing.

Cosmos 与 Sora 类视频模型:评价标准不同Cosmos vs. Sora-style video models: different yardsticks

娱乐向视频生成追求画面美感;Cosmos 优先追求物理一致性——三维空间一致、物体不会凭空消失、碰撞与运动符合常理。对训练具身智能来说,一段“好看但违反物理”的视频不仅无用,还可能教坏模型。

Entertainment-oriented video generation optimises for visual appeal; Cosmos prioritises physical consistency—coherent 3D space, objects that do not vanish, collisions and motion that make sense. For training embodied AI, a video that looks good but breaks physics is not just useless; it can teach the model the wrong lessons.

L3–L2 · 软件调度与算力Software & compute

算力基础设施与 AI 工厂Compute infrastructure and the AI factory

NVIDIA Blackwell Ultra & GB300 NVL72 as a worked example

上面所有能力,最后都要落到 GPU、显存和互联上。本节以 NVIDIA Blackwell Ultra 为例,因为它集中体现了当下的三条主线:更低的数值精度、更大的显存、以机架为单位的高速互联。NVIDIA 已推出下一代 Vera Rubin 平台,但这三条主线仍在延续。 Every capability above ultimately lands on GPUs, memory and interconnect. This section uses NVIDIA Blackwell Ultra as the example because it embodies three current trends: lower numeric precision, more memory, and rack-scale high-speed interconnect. NVIDIA has since launched the next-generation Vera Rubin platform, but the same three trends continue.

为什么需要 GPU 集群?Why GPU clusters?

训练 LLM 就像“读完全世界的书并总结规律”。CPU 像一位博士逐页精读;GPU 则像成千上万名速读助手同时开工——它擅长的正是神经网络里铺天盖地的矩阵乘法。但即便是 GPU,单卡也远远不够:

Training an LLM is like reading every book in the world and distilling the patterns. A CPU is one scholar reading page by page; a GPU is thousands of speed-readers working at once, which is exactly what the matrix multiplications in neural networks need. Even so, one GPU is nowhere near enough:

  • 装不下:GPT-3 的 1750 亿参数以 FP16 存储就需约 350 GB,而单张 H100 只有 80 GB 显存。训练时还要存梯度和优化器状态,混合精度 + Adam 下每个参数约需 16 字节,总量达 TB 级。模型必须切分到许多 GPU 上。
  • It doesn't fit: GPT-3's 175B parameters need about 350 GB in FP16, while an H100 has 80 GB. Training also stores gradients and optimiser state—roughly 16 bytes per parameter with mixed precision and Adam—pushing the total into terabytes. The model must be split across many GPUs.
  • 算不完:按常被引用的估算,用单张 V100 训练一次 GPT-3 约需 355 年。
  • It doesn't finish: by a widely cited estimate, training GPT-3 once on a single V100 would take about 355 years.
  • 通信成为瓶颈:切分之后,GPU 之间每一步都要交换梯度和激活值(All-Reduce 等集合通信)。训练是同步进行的,最慢的那条链路决定整体速度——一条链路拥塞,成千上万张 GPU 都要等它。通信占比取决于模型规模、并行策略和网络设计,这正是下一节网络要解决的问题。
  • Communication becomes the bottleneck: once split, GPUs must exchange gradients and activations at every step (All-Reduce and other collectives). Training is synchronous, so the slowest link sets the pace—one congested link can stall thousands of GPUs. How much time goes to communication depends on model size, parallelism strategy and network design, which is the problem the next section on networking addresses.
Blackwell Ultra GPU:关键规格Blackwell Ultra GPU: key specifications
2080 亿208B
晶体管Transistors
Hopper 的 2.6 倍2.6× Hopper
15 PF
NVFP4 稠密算力Dense NVFP4 compute
约为 H100 FP8 稠密算力的 7.5 倍(Hopper 不支持 FP4)~7.5× H100 dense FP8 (Hopper has no FP4)
288 GB
HBM3e 显存HBM3e memory
8 组 12-Hi 堆叠;H100 的 3.6 倍Eight 12-Hi stacks; 3.6× H100
8 TB/s
显存带宽Memory bandwidth
H100 为 3.35 TB/sH100: 3.35 TB/s
  • 双芯合一:采用 TSMC 4NP 工艺,两颗光罩极限尺寸的裸片通过 10 TB/s 的 NV-HBI 互联,对软件呈现为一个 CUDA GPU。
  • Two dies, one GPU: built on TSMC 4NP, two reticle-sized dies are joined by 10 TB/s NV-HBI and appear to software as a single CUDA GPU.
  • NVFP4 低精度:4 位浮点配合两级缩放,NVIDIA 称精度接近 FP8(差异通常小于约 1%),显存占用相比 FP8 约减少 1.8 倍、相比 FP16 约 3.5 倍。
  • NVFP4 low precision: 4-bit floating point with two-level scaling; NVIDIA reports accuracy close to FP8 (often within about 1%), with a memory footprint ~1.8× smaller than FP8 and ~3.5× smaller than FP16.
  • 缓解“内存墙”:288 GB 显存让大模型和长上下文的 KV Cache 能放进更少的 GPU;对解码阶段而言,8 TB/s 的带宽往往比峰值算力更重要。
  • Easing the memory wall: 288 GB lets large models and long-context KV caches fit on fewer GPUs; for the decode phase, 8 TB/s of bandwidth often matters more than peak compute.
GB300 NVL72:把一个机架变成一台“大 GPU”GB300 NVL72: turning a rack into one big GPU

一个全液冷机架内集成 72 颗 Blackwell Ultra GPU 与 36 颗 Grace CPU。9 个 NVLink 交换托盘把 72 颗 GPU 连成一个无阻塞的 NVLink 域,聚合带宽 130 TB/s;机架内共约 37 TB 高速内存(约 20 TB HBM3e + 17 TB LPDDR5X)。这样,规模在 72 卡以内的模型并行(如张量并行、专家并行)可以完全在机架内的 NVLink 上完成,不必走更慢的机架间网络。

One fully liquid-cooled rack integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs. Nine NVLink switch trays join the 72 GPUs into one non-blocking NVLink domain with 130 TB/s aggregate bandwidth, and the rack holds about 37 TB of fast memory (≈20 TB HBM3e + 17 TB LPDDR5X). Model parallelism within 72 GPUs (tensor or expert parallelism, for instance) can therefore run entirely over NVLink inside the rack instead of the slower inter-rack network.

50×
AI 工厂总产出AI factory output
= 10× 每用户速度 × 5× 每兆瓦吞吐,对比 Hopper 平台= 10× per-user speed × 5× throughput per MW, vs Hopper
10×
每用户 TPSTPS per user
交互体验(响应速度)Interactive responsiveness
35×
每百万 Token 成本降幅Lower cost per million tokens
“最高可达”,出现在低延迟区间;数据来自 SemiAnalysis InferenceX,经 NVIDIA 引用"Up to", in the low-latency regime; SemiAnalysis InferenceX data cited by NVIDIA
百 kW 级~100 kW+
单机架功率Rack power
必须液冷;具体数值视配置而定Liquid cooling required; exact figure depends on configuration

以上倍数均为 NVIDIA 在特定模型与配置下的“最高可达”口径,包含软件优化的贡献,不能直接外推到所有负载。All multipliers are NVIDIA "up to" figures for specific models and configurations, include software gains, and do not generalise to every workload.

AI 工厂的四个组成部分Four building blocks of an AI factory

GPU 算力GPU compute

生产线本身。常见形态是每台服务器 8 张 GPU(HGX)通过 NVLink 互联,或像 NVL72 那样以整个机架为一个 NVLink 域。

The production line itself. Typically eight GPUs per server (HGX) linked by NVLink, or, as in NVL72, a whole rack as one NVLink domain.

网络Networking

分工明确的物流系统:后端网络承载 GPU 之间的集合通信,要求高带宽、低尾延迟;前端网络连接用户与外部服务;存储网络为训练喂数据、写检查点。

A logistics system with clear roles: the back-end network carries GPU-to-GPU collectives and needs high bandwidth and low tail latency; the front-end network connects users and external services; the storage network feeds training data and writes checkpoints.

存储Storage

全闪存并行文件系统(如 WEKA、VAST、DDN 等方案)支撑海量并发读写,避免 GPU “等米下锅”。

All-flash parallel file systems (such as WEKA, VAST or DDN) sustain massive concurrent I/O so GPUs are never left waiting for data.

软件调度Software & scheduling

Kubernetes(配合 Run:ai)或 Slurm 把整个集群池化并分配给作业;CUDA 让 GPU 执行并行计算;NCCL 负责多 GPU 间的集合通信并选择最优路径。

Kubernetes (with Run:ai) or Slurm pools the cluster and allocates it to jobs; CUDA runs the parallel computation on GPUs; NCCL handles multi-GPU collectives and picks the best paths.

L1 · 网络层Network

数据中心网络与 UEC 1.0Data-centre networking and UEC 1.0

Ultra Ethernet Consortium — Ethernet re-engineered for AI scale

NVLink 解决了机架内的互联;一旦作业跨越成百上千个机架,GPU 之间的数据就得走数据中心网络。这一层的设计目标很朴素:让最慢的那条链路不要拖垮整个集群。 NVLink handles interconnect inside a rack; once a job spans hundreds or thousands of racks, GPU traffic must cross the data-centre network. The goal of this layer is simple: stop the slowest link from dragging down the whole cluster.

训练与推理对网络的要求不同Training and inference stress the network differently
维度Dimension 训练Training 推理Inference
同步性Synchrony 全局同步的“计算—交换—规约”循环,木桶效应明显Globally synchronous compute–exchange–reduce cycle; the slowest link dominates 请求之间相互独立;但单个模型跨多 GPU / 多节点部署时(张量并行、专家并行),每个 Token 内部仍需同步通信Requests are independent, but when one model spans GPUs or nodes (tensor or expert parallelism), each token still needs synchronous communication
东西向流量East–west 大块、高带宽(每 GPU 400G/800G)的 RDMA 集合通信Large, high-bandwidth RDMA collectives (400G/800G per GPU) 对延迟敏感的小消息(如 MoE 的 All-to-All);Prefill/Decode 分离部署时还要高速传输 KV CacheLatency-sensitive small messages (e.g. MoE All-to-All); disaggregated prefill/decode also moves KV caches at high speed
南北向流量North–south 高吞吐的数据加载与检查点写入High-throughput data loading and checkpointing 接收用户请求、访问向量数据库与外部工具(RAG / Agent)User requests, vector database and tool access (RAG / agents)
核心指标Key metrics 作业完成时间(JCT)、扩展效率Job completion time (JCT), scaling efficiency TTFT、ITL/TPOT、吞吐量、并发用户数TTFT, ITL/TPOT, throughput, concurrent users
RoCEv2、InfiniBand 与 UEC 1.0 对比RoCEv2 vs InfiniBand vs UEC 1.0
特性Feature RoCEv2 InfiniBand UEC 1.0 (UET)
网络模型Network model 依靠 PFC 在以太网上实现无损Lossless Ethernet via PFC 原生无损(链路级信用流控)Natively lossless (link-level credits) 面向尽力而为的以太网设计,也可运行在无损网络上Designed for best-effort Ethernet; also runs on lossless fabrics
拥塞控制Congestion control DCQCN(ECN + PFC) 信用流控 + 拥塞控制(FECN/BECN)Credits + congestion control (FECN/BECN) 发送端拥塞控制 + 可选的接收端授信;可利用交换机的报文裁剪Sender-based CC + optional receiver credits; can use switch packet trimming
负载均衡Load balancing 传统为 ECMP 按流哈希Traditionally per-flow ECMP hashing 子网管理器配置路由 + 自适应路由Subnet-manager routing + adaptive routing 逐包喷洒(Packet Spraying)Per-packet spraying
乱序与丢包恢复Reordering & loss recovery 传统实现为 Go-Back-N,乱序/丢包代价高;新一代网卡支持选择性重传Traditionally Go-Back-N, so reordering/loss is costly; newer NICs support selective repeat 传统可靠连接要求按序;配合自适应路由时由新一代网卡处理乱序Reliable connections traditionally expect order; with adaptive routing, newer NICs handle reordering 原生容忍乱序,直接数据放置(DDP)Reordering tolerated by design, with direct data placement
连接状态Connection state 基于连接(QP)Connection-based (QP) 基于连接(QP)Connection-based (QP) 临时连接:无需握手即可发送,事务结束即释放状态;API 基于 libfabricEphemeral connections: send without a handshake, discard state afterwards; libfabric-based API
生态Ecosystem 开放、多厂商,但调优复杂Open, multi-vendor, but hard to tune IBTA 标准,市场由 NVIDIA 主导IBTA standard; market dominated by NVIDIA 多厂商开放标准(AMD、Broadcom、Cisco 等成员)Multi-vendor open standard (members include AMD, Broadcom, Cisco)
网内计算In-network compute 无None SHARP INC(网内集合通信),UEC 已在规划中INC (in-network collectives), planned by UEC

表中是各技术的典型形态。厂商增强(如 Spectrum-X 以太网的自适应路由、新一代 RoCE 网卡的选择性重传)正在缩小三者之间的差距。The table shows typical forms. Vendor enhancements—adaptive routing in Spectrum-X Ethernet, selective repeat in newer RoCE NICs—are narrowing the gaps.

UEC 1.0 的核心机制Core mechanisms of UEC 1.0

UEC 的设计取舍可以概括为一句话:把复杂度放到端侧网卡,让中间的交换网络保持简单、高速。下面四个机制环环相扣:喷洒带来乱序,DDP 消化乱序,裁剪让乱序下的丢包能被快速发现。

UEC's trade-off fits in one sentence: put the complexity in the endpoint NIC and keep the switching fabric simple and fast. The four mechanisms below interlock: spraying causes reordering, DDP absorbs it, and trimming lets loss be detected quickly despite it.

Spray逐包喷洒Per-packet spraying

发送端把同一条消息的报文分散到多条等价路径上,并根据实时拥塞信息调整,路径一拥塞就立刻避开。这样就不会像 ECMP 那样把几条“大象流”哈希到同一根链路上,链路负载更均衡。

The sender spreads one message's packets across many equal-cost paths and steers by real-time congestion signals, moving off any path that congests. Unlike ECMP, elephant flows no longer collide on one link, so load is far more even.

DDP直接数据放置Direct data placement

每个报文自带足以定位目标缓冲区的信息(如内存键与偏移量)。乱序到达的报文不必先在网卡里重排,可以直接写入主机或 GPU 内存的正确位置;上层通信库(如 NCCL)看到的仍是完整、有序的消息。

Each packet carries enough information to locate its target buffer (such as a memory key and offset). Out-of-order packets need not be reassembled in the NIC; they are written straight to the right place in host or GPU memory, while libraries such as NCCL still see complete, ordered messages.

Trim报文裁剪Packet trimming

可选功能。交换机缓存溢出时不直接丢弃报文,而是裁掉载荷、只把报头转发给接收端。接收端据此立即知道“哪个包丢了”,精确重传,而不必在乱序环境中靠超时猜测。

Optional. When a switch buffer overflows, instead of dropping a packet it cuts the payload and forwards only the header. The receiver immediately knows which packet was lost and requests a precise retransmit, instead of guessing via timeouts amid reordering.

INC网内集合通信In-network collectives

让交换机在转发途中直接完成梯度聚合(如 All-Reduce 求和),理论上可将端点间需要传输的数据量减少约一半。InfiniBand 的 SHARP 已实现类似能力;UEC 将其规划为以太网上的首个开放标准,实际部署仍有待观察。

Switches aggregate gradients in flight (for example, All-Reduce sums), in theory roughly halving the data endpoints must exchange. InfiniBand's SHARP already does this; UEC plans it as the first open standard over Ethernet, and real deployments remain to be seen.

UEC 的意义与待验证之处Why UEC matters, and what is still unproven

UEC 的价值在于把 AI/HPC 所需的传输语义标准化到以太网生态中:多厂商可互通,降低对单一供应商的依赖,并沿用以太网已有的运维体系与规模经济。规范面向最多百万级端点的扩展。AMD、Broadcom 等厂商已推出宣称兼容 UEC 的网卡与交换芯片。

UEC's value lies in standardising the transport semantics AI/HPC needs within the Ethernet ecosystem: multi-vendor interoperability, less dependence on a single supplier, and reuse of Ethernet's operational tooling and economies of scale. The specification targets scaling to millions of endpoints, and vendors such as AMD and Broadcom have released NICs and switch silicon that claim UEC compatibility.

它能在多大程度上缩小与 InfiniBand 在尾延迟和大规模稳定性上的差距,还需要更多大规模生产部署的数据来回答。

How far it closes the gap with InfiniBand on tail latency and stability at scale still needs data from large production deployments.

负载均衡方案:从 ECMP 到逐包喷洒Load balancing: from ECMP to packet spraying
方案Scheme 粒度Granularity 乱序风险Reordering risk 均衡效果Balance
Static ECMP 按流(五元组哈希)Per flow (5-tuple hash) 无None 差:少量大象流易撞到同一链路Poor: a few elephant flows collide on one link
Enhanced ECMP / QP scaling 按流(增加 QP 数量与哈希因子)Per flow (more QPs, richer hash) 无None 中Medium
Flowlet / DLB 按流片段(利用流内空闲间隙换路径)Per flowlet (switch paths at idle gaps) 极低Very low 高High
Adaptive Routing 逐包,交换机按本地拥塞选路(IB、Spectrum-X 等)Per packet, switch picks by local congestion (IB, Spectrum-X, etc.) 高,需网卡处理乱序High; NIC must handle reordering 很高Very high
Packet Spraying (UET) 逐包,由发送端网卡选路Per packet, sender NIC picks paths 高,由 DDP 解决High; solved by DDP 很高,接近满链路利用(取决于拥塞反馈质量)Very high, near-full link utilisation (depends on congestion feedback)
Centralized TE 按流,由中央控制器下发路径Per flow, paths set by a central controller 无None 高,但受控制面反应速度限制High, but limited by control-plane reaction time
L7 → L1

结语:AI 是一项全栈系统工程Conclusion: AI is full-stack systems engineering

The age of full-stack AI system engineering

把七层连起来看,会发现每一层的突破都依赖其他层。NotebookLM 这类基于 Agentic RAG 的应用,正把 AI 从“信息总结器”推向能自主检索、多模态输出的研究助手;而它能否可信,首先取决于检索管线能否精确找到证据——CRAG 与 LegalBench-RAG 都说明,检索与严谨的评估是降低幻觉、构建可信企业 AI 的关键环节。

Seen end to end, every layer's progress depends on the others. Agentic-RAG applications such as NotebookLM are moving AI from summariser to research assistant that retrieves autonomously and answers in multiple media; whether it can be trusted depends first on whether the retrieval pipeline finds the right evidence. CRAG and LegalBench-RAG both show that retrieval and rigorous evaluation are key to reducing hallucination and building trustworthy enterprise AI.

当 AI 走进物理世界,NVIDIA Cosmos 这类世界模型用合成数据弥补真实数据的稀缺。而从文本到物理的所有算法创新,最终都转化为对算力和互联的需求:GB300 NVL72 通过 NVFP4 低精度、大容量 HBM 与机架级 NVLink 的软硬协同,大幅降低推理成本;UEC 1.0 则以逐包喷洒、直接数据放置和临时连接,尝试让开放以太网承担起跨机架的大规模 AI 流量。

As AI moves into the physical world, world models such as NVIDIA Cosmos offset the scarcity of real data with synthetic data. And every algorithmic advance, from text to physics, ultimately becomes demand for compute and interconnect: GB300 NVL72 co-designs NVFP4 precision, large HBM and rack-scale NVLink to cut inference cost sharply, while UEC 1.0 uses per-packet spraying, direct data placement and ephemeral connections to let open Ethernet carry large-scale AI traffic across racks.

所以,AI 的进步早已不只是“把模型做大”,而是从算法、评估、芯片、内存一直到网络协议的全链路协同。下次评估一个 AI 系统时,可以拿这张地图自上而下问三个问题:

AI progress, then, is no longer just about bigger models; it is co-design across the whole chain, from algorithms and evaluation to chips, memory and network protocols. Next time you assess an AI system, walk this map top-down with three questions:

  1. 答错了?先看检索(L6)有没有找到正确证据,再看模型(L5)有没有忠实使用它。Wrong answer? First check whether retrieval (L6) found the right evidence, then whether the model (L5) used it faithfully.
  2. 变慢了?分清是首 Token 慢(Prefill、算力)、逐字慢(Decode、显存带宽),还是多卡通信慢(NVLink 与网络,L2–L1)。Too slow? Tell apart a slow first token (prefill, compute), slow streaming (decode, memory bandwidth) and slow multi-GPU communication (NVLink and network, L2–L1).
  3. 太贵了?看数值精度、显存能否装下模型与 KV Cache、互联能否让 GPU 不空等(L2–L3)。Too expensive? Look at numeric precision, whether memory fits the model and KV cache, and whether interconnect keeps GPUs from idling (L2–L3).
参考来源Sources

主要来源Main sources

  1. Vaswani et al., Attention Is All You Need, 2017.
  2. Wei et al., Emergent Abilities of Large Language Models, 2022; Schaeffer et al., Are Emergent Abilities of Large Language Models a Mirage?, 2023.
  3. Ouyang et al., Training language models to follow instructions with human feedback, 2022.
  4. Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, 2022.
  5. Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation, 2023.
  6. Yang et al., CRAG – Comprehensive RAG Benchmark, 2024.
  7. Pipitone & Houir Alami, LegalBench-RAG, 2024.
  8. Google DeepMind, Pushing the frontiers of audio generation; Google, NotebookLM Video Overviews and an upgraded Studio.
  9. NVIDIA, Cosmos World Foundation Model Platform for Physical AI, 2025; NVIDIA Cosmos.
  10. NVIDIA Technical Blog, Inside NVIDIA Blackwell Ultra, 2025; GB300 NVL72 product page.
  11. NVIDIA Blog, SemiAnalysis InferenceX data: Blackwell Ultra up to 50x better performance and 35x lower costs.
  12. Ultra Ethernet Consortium, Ultra Ethernet Specification Update, 2024; UEC launches Specification 1.0, 2025.