|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
CPU 硬體和軟體最佳化的進步使 CPU 更適合運行傳統上由 GPU 主導的生成式 AI 聊天機器人和服務。英特爾和 Ampere 已經透過較小的語言模型展示了令人鼓舞的性能,實現了適合實際使用的令牌延遲。雖然與專用加速器相比,CPU 面臨記憶體頻寬限制,但即將推出的 MCR DIMM 和 4 位元操作等功能旨在解決這些瓶頸。這表明 CPU 可能能夠處理中等大小的人工智慧模型,將重點轉移到最佳化以實現廣泛採用。

CPUs Gain Ground in Generative AI Race as Intel and Ampere Push Limits
隨著英特爾和 Ampere 突破極限,CPU 在生成式 AI 競賽中取得進展
Introduction
介紹
The realm of generative artificial intelligence (AI) has been largely dominated by graphics processing units (GPUs) and specialized accelerators due to their unparalleled computational power. However, as smaller and more widely deployable AI models emerge within enterprises, CPU manufacturers Intel and Ampere are asserting that their products can effectively handle these tasks. Recent advancements in software optimizations and mitigation of hardware bottlenecks have paved the way for CPUs to become viable options in the AI landscape.
生成式人工智慧 (AI) 領域主要由圖形處理單元 (GPU) 和專用加速器主導,因為它們具有無與倫比的運算能力。然而,隨著企業內部出現更小且可廣泛部署的人工智慧模型,CPU 製造商英特爾和 Ampere 聲稱他們的產品可以有效地處理這些任務。軟體優化和硬體瓶頸緩解的最新進展為 CPU 成為人工智慧領域的可行選擇鋪平了道路。
Intel's Progress with Xeon Processors
英特爾至強處理器的進展
At Intel's Vision event in April, CEO Pat Gelsinger showcased the company's achievements in adapting larger language models (LLMs) for execution on its Xeon platform. A live demonstration featuring the forthcoming Granite Rapids Xeon 6 processor revealed Meta's Llama2-70B model operating at 4-bit precision with an impressive second token latency of 82 milliseconds (ms).
在 4 月的英特爾願景活動中,執行長 Pat Gelsinger 展示了該公司在採用更大的語言模型 (LLM) 以在其至強平台上執行方面取得的成就。採用即將推出的 Granite Rapids Xeon 6 處理器的現場演示揭示了 Meta 的 Llama2-70B 模型以 4 位精度運行,具有令人印象深刻的 82 毫秒 (ms) 秒令牌延遲。
Second token latency gauges the time required for an AI model to analyze a query and provide its first response. A lower latency translates to perceived performance enhancements. In terms of performance metrics, the observed 82ms latency corresponds to approximately 12 tokens per second.
第二個令牌延遲衡量人工智慧模型分析查詢並提供第一個回應所需的時間。較低的延遲意味著可感知的性能增強。就效能指標而言,觀察到的 82 毫秒延遲相當於每秒約 12 個令牌。
This result represents a significant improvement over Intel's 5th-generation Xeon processors released in December, which exhibited a second token latency of 151ms.
這一結果比英特爾去年 12 月發布的第五代至強處理器有了顯著改進,後者的第二個令牌延遲為 151 毫秒。
Oracle's Results with Ampere's CPUs
Oracle 使用 Ampere CPU 所得的結果
Oracle has also published test data related to executing the Llama2-7B model on Ampere's Altra central processing units (CPUs). Utilizing a 64-core OCI A1 instance paired with a 4-bit quantized version of the model, Oracle achieved throughput rates ranging from 33 to 119 tokens per second for batch sizes of 1 and 16, respectively.
Oracle 也發布了在 Ampere 的 Altra 中央處理單元 (CPU) 上執行 Llama2-7B 模型的相關測試資料。利用 64 核心 OCI A1 實例與模型的 4 位元量化版本配對,Oracle 在批次大小為 1 和 16 的情況下分別實現了每秒 33 到 119 個令牌的吞吐率。
In the context of conversational chatbots, a larger batch size equates to a higher capacity to concurrently handle multiple queries. Oracle's testing revealed a direct correlation between batch size and throughput; however, the larger the batch size, the slower the model generated text. For instance, at a batch size of 16, Oracle attained its highest throughput performance, but the output rate was approximately 7.5 tokens per second per query. This delay would be noticeable from the end-user's perspective.
在對話式聊天機器人的上下文中,較大的批量大小相當於同時處理多個查詢的更高容量。 Oracle 的測試揭示了批量大小和吞吐量之間的直接相關性;但是,批量大小越大,模型生成文字的速度就越慢。例如,在批次大小為 16 時,Oracle 獲得了最高的吞吐量效能,但每個查詢的輸出率約為每秒 7.5 個代幣。從最終用戶的角度來看,這種延遲是顯而易見的。
Oracle shared results across multiple batch sizes, while Intel's data is limited to batch size one. Intel has been contacted for further details on performance at higher batch sizes.
Oracle 共享多個批次大小的結果,而英特爾的資料僅限於一個批次大小。已聯繫英特爾以獲取有關更高批量大小的性能的更多詳細資訊。
Causes of Improved Performance
提升績效的原因
According to Jeff Wittich, Ampere's chief product officer, these performance gains were largely attributed to custom software libraries and optimizations to Llama.cpp, developed in collaboration with Oracle. Both Oracle and Intel have since released performance metrics for Meta's newly launched Llama3 models, demonstrating similar performance characteristics.
Ampere 首席產品長 Jeff Wittich 表示,這些效能提升很大程度上歸功於與 Oracle 合作開發的客製化軟體庫和對 Llama.cpp 的最佳化。此後,甲骨文和英特爾都發布了 Meta 新推出的 Llama3 模型的性能指標,展示了類似的性能特徵。
Pending the accuracy of these performance claims – given the test parameters and our experience running 4-bit quantized models on CPUs – CPUs appear to be a viable option for executing small-scale models. In the near future, they may also be capable of handling moderately sized models, particularly at relatively small batch sizes.
在這些效能聲明的準確性之前——考慮到測試參數和我們在 CPU 上運行 4 位元量化模型的經驗——CPU 似乎是執行小規模模型的可行選擇。在不久的將來,它們也可能能夠處理中等大小的模型,特別是相對較小的批量大小。
Limitations and Ongoing Challenges
局限性和持續的挑戰
While Intel and Ampere have successfully demonstrated LLMs running on their respective CPU platforms, it is crucial to recognize that various compute and memory limitations prevent CPUs from completely replacing GPUs or dedicated accelerators for larger-scale models.
雖然英特爾和 Ampere 已成功展示了在各自 CPU 平台上運行的法學碩士,但重要的是要認識到,各種計算和內存限制阻止 CPU 完全取代 GPU 或大型模型的專用加速器。
For models pushing the boundaries of generative AI, Ronak Shah, director of Xeon AI product management at Intel, emphasized that upcoming products like the Gaudi accelerator are specifically engineered for such tasks.
對於突破生成式人工智慧邊界的模式,英特爾至強人工智慧產品管理總監 Ronak Shah 強調,即將推出的產品(例如 Gaudi 加速器)是專門為此類任務而設計的。
Overcoming Bottlenecks
克服瓶頸
Historically, conversations surrounding the execution of LLMs on CPUs have been subdued because, despite increasing core counts, conventional processors still fall short in terms of parallelism compared to modern GPUs and accelerators designed for AI workloads.
從歷史上看,圍繞在CPU 上執行LLM 的討論一直很低調,因為儘管核心數量不斷增加,但與專為AI 工作負載設計的現代GPU 和加速器相比,傳統處理器在並行性方面仍然存在不足。
However, CPUs are undergoing significant enhancements. Modern units dedicate a substantial portion of their die space to features such as vector extensions or even specialized matrix math accelerators.
然而,CPU 正在經歷顯著的增強。現代單元將其晶片空間的很大一部分用於向量擴展甚至專用矩陣數學加速器等功能。
Intel incorporated the latter feature in its Sapphire Rapids Xeon Scalable processors released early last year. Each core is also equipped with Advanced Matrix Extensions (AMX), although not all stock-keeping units (SKUs) support AMX due to the flexibility of software-defined silicon.
英特爾在去年初發布的 Sapphire Rapids Xeon 可擴充處理器中整合了後者功能。每個核心還配備了高級矩陣擴展 (AMX),但由於軟體定義晶片的靈活性,並非所有庫存單元 (SKU) 都支援 AMX。
As the name suggests, AMX extensions are tailored to accelerate matrix math calculations prevalent in deep learning workloads. Since its initial implementation, Intel has continuously refined its AMX engines for improved performance on larger models. This advancement is likely reflected in the upcoming Intel Xeon 6 processors slated for release later this year.
顧名思義,AMX 擴展專為加速深度學習工作負載中普遍存在的矩陣數學計算而量身定制。自最初實施以來,英特爾不斷改進其 AMX 引擎,以提高大型車型的性能。這一進步可能會反映在即將於今年稍後發布的英特爾至強 6 處理器中。
While Intel heavily relies on matrix acceleration, Ampere's Wittich explained that acceptable performance can be achieved using the two 128-bit vector units embedded in each of its AmpereOne and Altra cores. These vector units support FP16, BF16, INT8, and INT16 precision levels.
雖然英特爾嚴重依賴矩陣加速,但 Ampere 的 Wittich 解釋說,使用每個 AmpereOne 和 Altra 內核中嵌入的兩個 128 位元向量單元可以獲得可接受的性能。這些向量單元支援 FP16、BF16、INT8 和 INT16 精度等級。
Memory Bottlenecks and MCR DIMMs
記憶體瓶頸和 MCR DIMM
Despite their inferior performance in executing OPS or FLOPS compared to GPUs, CPUs possess a significant advantage: their independence from expensive and capacity-constrained high-bandwidth memory (HBM) modules.
儘管與 GPU 相比,CPU 在執行 OPS 或 FLOPS 方面的效能較差,但它具有顯著的優勢:它們獨立於昂貴且容量受限的高頻寬記憶體 (HBM) 模組。
As previously discussed, operating a model at FP8/INT8 requires approximately 1 gigabyte (GB) of memory for each billion parameters. Consequently, executing a model like OpenAI's 1.7 trillion parameter GPT-4 model at FP8 would necessitate over 1.7 terabytes (TB) of memory, which would be approximately halved when quantized to 4-bits. This memory requirement exceeds the capacity of any single GPU but falls within the capabilities of modern CPUs.
如前所述,在 FP8/INT8 下運行模型時,每十億個參數需要大約 1 GB 的記憶體。因此,在 FP8 上執行 OpenAI 的 1.7 兆參數 GPT-4 模型等模型將需要超過 1.7 TB 的內存,當量化為 4 位時,內存將大約減半。此記憶體需求超出了任何單一 GPU 的容量,但在現代 CPU 的能力範圍內。
However, the drawback lies in the sluggish speed of large DRAM modules used by CPUs compared to HBM.
但缺點是CPU使用的大型DRAM模組的速度比HBM慢。
With only eight memory channels currently supported on Intel's 5th-generation Xeon and Ampere's One processors, these chips are limited to roughly 350 gigabytes per second (GB/sec) of memory bandwidth when running 5600MT/sec DIMMs. While Wittich mentioned plans for a 12-channel version of Ampere's chip with a targeted release later this year – featuring a purported 256 cores – it is not yet available.
由於英特爾第 5 代 Xeon 和 Ampere's One 處理器目前僅支援 8 個記憶體通道,因此這些晶片在運行 5600MT/秒 DIMM 時,記憶體頻寬限制為大約 350 GB/秒 (GB/秒)。雖然 Wittich 提到計劃於今年稍後發布 Ampere 晶片的 12 通道版本(據稱具有 256 個核心),但目前該版本尚未上市。
Nevertheless, all of Oracle's testing has been conducted on Ampere's Altra generation, which utilizes even slower DDR4 memory and operates at a maximum bandwidth of approximately 200GB/sec. This suggests the potential for substantial performance gains by upgrading to the newer AmpereOne cores.
儘管如此,Oracle 的所有測試都是在 Ampere 的 Altra 世代上進行的,該世代使用更慢的 DDR4 內存,並以大約 200GB/秒的最大頻寬運行。這表明昇級到較新的 AmpereOne 核心有可能大幅提升效能。
These bandwidth speeds may seem impressive – certainly faster than an SSD – but the eight HBM modules found on AMD's MI300X or Nvidia's upcoming Blackwell GPUs deliver speeds of 5.3 TB/sec and 8TB/sec, respectively. However, HBM modules are constrained by a maximum capacity of 192GB.
這些頻寬速度可能看起來令人印象深刻(當然比 SSD 更快),但 AMD MI300X 或 Nvidia 即將推出的 Blackwell GPU 上的 8 個 HBM 模組可分別提供 5.3 TB/秒和 8TB/秒的速度。然而,HBM 模組的最大容量限制為 192GB。
To illustrate this concept, consider memory capacity as a fuel tank, memory bandwidth as a fuel line, and compute as an internal combustion engine. Regardless of the size of the fuel tank or the power of the engine, if the fuel line is too narrow to supply sufficient fuel for optimal engine performance, the system will be hindered.
為了說明這個概念,請將記憶體容量視為油箱,將記憶體頻寬視為燃油管路,並將計算視為內燃機。無論油箱尺寸或引擎功率有多大,如果燃油管路太窄而無法提供足夠的燃油以實現最佳引擎性能,系統就會受到阻礙。
This limitation explains why previous attempts to execute LLMs on CPUs were largely confined to smaller models.
這項限制解釋了為什麼先前在 CPU 上執行 LLM 的嘗試主要局限於較小的模型。
Clearing the Bottlenecks: Intel's Granite Rapids Xeon 6
清除瓶頸:英特爾 Granite Rapids Xeon 6
Despite these obstacles, Intel's forthcoming Granite Rapids Xeon 6 platform provides clues on how CPUs could potentially handle larger models in the near future.
儘管存在這些障礙,英特爾即將推出的 Granite Rapids Xeon 6 平台為 CPU 如何在不久的將來處理更大的模型提供了線索。
Intel's recent demonstration showcased a single Xeon 6 processor effortlessly running Llama2-70B with a reasonable second token latency of 82ms. Crucially, many details regarding the test rig remain unknown, including the number and clock speed of the cores. These details will likely be revealed later this year – potentially in December.
英特爾最近的演示展示了單個 Xeon 6 處理器輕鬆運行 Llama2-70B,第二個令牌延遲合理為 82 毫秒。至關重要的是,有關測試裝置的許多細節仍然未知,包括內核的數量和時脈速度。這些細節可能會在今年稍後(可能是 12 月)公佈。
"The substantial advancement from 5th-generation Xeon to Xeon 6 lies in the introduction of MCR DIMMs, which effectively clears many of the bottlenecks associated with memory-bound workloads," explained Shah.
「從第 5 代 Xeon 到 Xeon 6 的重大進步在於 MCR DIMM 的引入,它有效地消除了與記憶體限制工作負載相關的許多瓶頸,」Shah 解釋道。
Multiplexer combined rank (MCR) DIMMs facilitate much faster memory access compared to standard DRAM. Intel has already demonstrated this technology running at 8,800MT/sec. With 12 memory channels equipped with MCR DIMMs, a single Granite Rapids socket would access approximately 825GB/sec of bandwidth – a significant leap from previous generations.
與標準 DRAM 相比,多工器組合列 (MCR) DIMM 可實現更快的記憶體存取速度。英特爾已經展示了該技術的運行速度為 8,800MT/秒。憑藉配備 MCR DIMM 的 12 個記憶體通道,單一 Granite Rapids 插槽可存取約 825GB/秒的頻寬,與前幾代相比實現了顯著飛躍。
Wittich noted that Ampere is also exploring the implementation of MCR DIMMs but did not provide a timeline for their inclusion in Ampere silicon.
Wittich 指出,Ampere 也在探索 MCR DIMM 的實施,但沒有提供將其納入 Ampere 晶片的時間表。
However, faster memory technology is not the sole innovation offered by Granite Rapids. Intel's AMX engine has gained support for 4-bit operations via the introduction of the MXFP4 data type, which theoretically has the potential to double effective performance.
然而,更快的記憶體技術並不是 Granite Rapids 提供的唯一創新。英特爾的 AMX 引擎透過引入 MXFP4 資料類型獲得了對 4 位元操作的支持,理論上有可能使有效性能翻倍。
Moreover, lower precision reduces the model footprint and subsequently lowers memory capacity and bandwidth requirements. Quantization techniques employed to compress models trained at higher precisions can also achieve similar reductions in footprint and bandwidth usage. Therefore, the practical benefit of supporting 4-bit mathematics in hardware primarily manifests as performance enhancements.
此外,較低的精度會減少模型佔用空間,從而降低記憶體容量和頻寬要求。用於壓縮以更高精度訓練的模型的量化技術也可以實現類似的佔地面積和頻寬使用量的減少。因此,在硬體中支援4位數學的實際好處主要表現為性能的增強。
Balancing Act: Optimizing CPU Design for AI
平衡法:針對 AI 最佳化 CPU 設計
For CPU designers, striking the right balance of AI capabilities presents a challenge. Excessive allocation of die area to features like AMX risks transforming the chip into an AI accelerator rather than a general-purpose processor.
對於 CPU 設計人員來說,在人工智慧功能之間取得適當的平衡是一項挑戰。為 AMX 等功能過度分配晶片面積可能會導致晶片轉變為 AI 加速器而不是通用處理器。
As a result, instead of aiming for CPUs capable of handling the largest and most demanding LLMs, vendors are focusing on the distribution of AI models to identify the most widely adopted models and optimizing their products to cater to these workloads.
因此,供應商不再將目標放在能夠處理最大、要求最高的 LLM 的 CPU 上,而是專注於 AI 模型的分發,以確定最廣泛採用的模型,並優化其產品以滿足這些工作負載。
"From a customer perspective, the sweet spot right now revolves around models with 7–13 billion parameters. That's where most of our attention is directed today," stated Wittich.
「從客戶的角度來看,目前的最佳點圍繞著具有 7-130 億個參數的模型。這就是我們今天大部分注意力所關注的地方,」Wittich 說。
Intel's Shah has observed a similar
英特爾的 Shah 也觀察到了類似的情況
免責聲明:info@kdj.com
所提供的資訊並非交易建議。 kDJ.com對任何基於本文提供的資訊進行的投資不承擔任何責任。加密貨幣波動性較大,建議您充分研究後謹慎投資!
如果您認為本網站使用的內容侵犯了您的版權,請立即聯絡我們(info@kdj.com),我們將及時刪除。
-
- 比特幣、eCash 分叉和空投動態:深入探討加密貨幣的最新爭議
- 2026-05-03 00:52:02
- 探索最近的 eCash 分叉、其作為高風險空投的分類,以及對比特幣和加密生態系統的更廣泛影響。
-
-
- 聯準會維持利率穩定,地緣政治緊張局勢引發比特幣價格下跌
- 2026-05-01 04:04:38
- 聯準會維持利率的決定,加上中東衝突,影響了比特幣的價格。分析近期趨勢和市場反應。
-
-
-
-
-
-

































