시가총액: $2.2316T 1.41%
거래량(24시간): $50.2264B 30.30%
  • 시가총액: $2.2316T 1.41%
  • 거래량(24시간): $50.2264B 30.30%
  • 공포와 탐욕 지수:
  • 시가총액: $2.2316T 1.41%
암호화
주제
암호화
소식
cryptostopics
비디오
최고의 뉴스
암호화
주제
암호화
소식
cryptostopics
비디오
bitcoin
bitcoin

$87959.907984 USD

1.34%

ethereum
ethereum

$2920.497338 USD

3.04%

tether
tether

$0.999775 USD

0.00%

xrp
xrp

$2.237324 USD

8.12%

bnb
bnb

$860.243768 USD

0.90%

solana
solana

$138.089498 USD

5.43%

usd-coin
usd-coin

$0.999807 USD

0.01%

tron
tron

$0.272801 USD

-1.53%

dogecoin
dogecoin

$0.150904 USD

2.96%

cardano
cardano

$0.421635 USD

1.97%

hyperliquid
hyperliquid

$32.152445 USD

2.23%

bitcoin-cash
bitcoin-cash

$533.301069 USD

-1.94%

chainlink
chainlink

$12.953417 USD

2.68%

unus-sed-leo
unus-sed-leo

$9.535951 USD

0.73%

zcash
zcash

$521.483386 USD

-2.87%

암호화폐 뉴스 기사

인텔과 암페어, AI 경쟁에 앞장서는 CPU의 타당성 입증

2024/05/01 19:24

CPU 하드웨어 및 소프트웨어 최적화의 발전으로 인해 전통적으로 GPU가 지배했던 생성 AI 챗봇 및 서비스를 실행하는 데 CPU의 실행 가능성이 더욱 높아졌습니다. Intel과 Ampere는 더 작은 언어 모델로 유망한 성능을 입증하여 실제 사용에 적합한 토큰 대기 시간을 달성했습니다. CPU는 특수 가속기에 비해 메모리 대역폭 제한에 직면해 있지만 MCR DIMM 및 4비트 작업과 같은 향후 기능은 이러한 병목 현상을 해결하는 것을 목표로 합니다. 이는 CPU가 적당한 크기의 AI 모델을 처리할 수 있게 되어 광범위한 채택을 위한 최적화에 중점을 둘 수 있음을 의미합니다.

인텔과 암페어, AI 경쟁에 앞장서는 CPU의 타당성 입증

CPUs Gain Ground in Generative AI Race as Intel and Ampere Push Limits

CPU는 Intel 및 Ampere 푸시 제한으로 생성 AI 경쟁에서 기반을 확보합니다.

Introduction

소개

The realm of generative artificial intelligence (AI) has been largely dominated by graphics processing units (GPUs) and specialized accelerators due to their unparalleled computational power. However, as smaller and more widely deployable AI models emerge within enterprises, CPU manufacturers Intel and Ampere are asserting that their products can effectively handle these tasks. Recent advancements in software optimizations and mitigation of hardware bottlenecks have paved the way for CPUs to become viable options in the AI landscape.

생성적 인공 지능(AI)의 영역은 비교할 수 없는 계산 능력으로 인해 그래픽 처리 장치(GPU)와 특수 가속기가 주로 지배해 왔습니다. 그러나 더 작고 더 광범위하게 배포 가능한 AI 모델이 기업 내에서 등장함에 따라 CPU 제조업체인 Intel과 Ampere는 자사 제품이 이러한 작업을 효과적으로 처리할 수 있다고 주장하고 있습니다. 소프트웨어 최적화 및 하드웨어 병목 현상 완화의 최근 발전으로 CPU가 AI 환경에서 실행 가능한 옵션이 될 수 있는 길이 열렸습니다.

Intel's Progress with Xeon Processors

Xeon 프로세서를 통한 Intel의 발전

At Intel's Vision event in April, CEO Pat Gelsinger showcased the company's achievements in adapting larger language models (LLMs) for execution on its Xeon platform. A live demonstration featuring the forthcoming Granite Rapids Xeon 6 processor revealed Meta's Llama2-70B model operating at 4-bit precision with an impressive second token latency of 82 milliseconds (ms).

지난 4월 Intel의 Vision 이벤트에서 Pat Gelsinger CEO는 Xeon 플랫폼 실행을 위해 LLM(대형 언어 모델)을 적용한 회사의 성과를 선보였습니다. 곧 출시될 Granite Rapids Xeon 6 프로세서를 특징으로 하는 라이브 데모에서는 82밀리초(ms)의 인상적인 두 번째 토큰 대기 시간으로 4비트 정밀도로 작동하는 Meta의 Llama2-70B 모델이 공개되었습니다.

Second token latency gauges the time required for an AI model to analyze a query and provide its first response. A lower latency translates to perceived performance enhancements. In terms of performance metrics, the observed 82ms latency corresponds to approximately 12 tokens per second.

두 번째 토큰 대기 시간은 AI 모델이 쿼리를 분석하고 첫 번째 응답을 제공하는 데 필요한 시간을 측정합니다. 대기 시간이 짧을수록 인지된 성능 향상으로 해석됩니다. 성능 지표 측면에서 관찰된 82ms 대기 시간은 초당 약 12개의 토큰에 해당합니다.

This result represents a significant improvement over Intel's 5th-generation Xeon processors released in December, which exhibited a second token latency of 151ms.

이 결과는 12월에 출시된 Intel의 5세대 Xeon 프로세서(151ms의 두 번째 토큰 대기 시간)에 비해 크게 개선되었음을 나타냅니다.

Oracle's Results with Ampere's CPUs

Ampere의 CPU를 사용한 Oracle의 결과

Oracle has also published test data related to executing the Llama2-7B model on Ampere's Altra central processing units (CPUs). Utilizing a 64-core OCI A1 instance paired with a 4-bit quantized version of the model, Oracle achieved throughput rates ranging from 33 to 119 tokens per second for batch sizes of 1 and 16, respectively.

Oracle은 또한 Ampere의 Altra 중앙 처리 장치(CPU)에서 Llama2-7B 모델을 실행하는 것과 관련된 테스트 데이터를 게시했습니다. Oracle은 모델의 4비트 양자화 버전과 결합된 64코어 OCI A1 인스턴스를 활용하여 배치 크기 1과 16에 대해 초당 33~119개 토큰 범위의 처리 속도를 달성했습니다.

In the context of conversational chatbots, a larger batch size equates to a higher capacity to concurrently handle multiple queries. Oracle's testing revealed a direct correlation between batch size and throughput; however, the larger the batch size, the slower the model generated text. For instance, at a batch size of 16, Oracle attained its highest throughput performance, but the output rate was approximately 7.5 tokens per second per query. This delay would be noticeable from the end-user's perspective.

대화형 챗봇의 맥락에서 배치 크기가 클수록 여러 쿼리를 동시에 처리할 수 있는 용량이 더 커집니다. Oracle의 테스트에서는 배치 크기와 처리량 간의 직접적인 상관 관계가 밝혀졌습니다. 그러나 배치 크기가 클수록 모델이 텍스트를 생성하는 속도가 느려집니다. 예를 들어, 배치 크기 16에서 Oracle은 가장 높은 처리량 성능을 달성했지만 출력 속도는 쿼리당 초당 약 7.5토큰이었습니다. 이러한 지연은 최종 사용자의 관점에서 눈에 띄게 나타납니다.

Oracle shared results across multiple batch sizes, while Intel's data is limited to batch size one. Intel has been contacted for further details on performance at higher batch sizes.

Oracle은 여러 배치 크기에 걸쳐 결과를 공유했지만 Intel의 데이터는 배치 크기 1로 제한되었습니다. 더 높은 배치 크기의 성능에 대한 자세한 내용을 알아보기 위해 인텔에 연락했습니다.

Causes of Improved Performance

성능 향상의 원인

According to Jeff Wittich, Ampere's chief product officer, these performance gains were largely attributed to custom software libraries and optimizations to Llama.cpp, developed in collaboration with Oracle. Both Oracle and Intel have since released performance metrics for Meta's newly launched Llama3 models, demonstrating similar performance characteristics.

Ampere의 최고 제품 책임자인 Jeff Wittich에 따르면 이러한 성능 향상은 주로 Oracle과 공동으로 개발된 Llama.cpp에 대한 맞춤형 소프트웨어 라이브러리 및 최적화 덕분이라고 합니다. Oracle과 Intel은 이후 Meta가 새로 출시한 Llama3 모델에 대한 성능 지표를 발표하여 유사한 성능 특성을 보여주었습니다.

Pending the accuracy of these performance claims – given the test parameters and our experience running 4-bit quantized models on CPUs – CPUs appear to be a viable option for executing small-scale models. In the near future, they may also be capable of handling moderately sized models, particularly at relatively small batch sizes.

테스트 매개변수와 CPU에서 4비트 양자화 모델을 실행한 경험을 고려할 때 이러한 성능 주장의 정확성이 유지될 때까지 CPU는 소규모 모델을 실행하는 데 실행 가능한 옵션인 것으로 보입니다. 가까운 미래에는 중간 크기의 모델, 특히 상대적으로 작은 배치 크기의 모델을 처리할 수도 있습니다.

Limitations and Ongoing Challenges

한계와 지속적인 과제

While Intel and Ampere have successfully demonstrated LLMs running on their respective CPU platforms, it is crucial to recognize that various compute and memory limitations prevent CPUs from completely replacing GPUs or dedicated accelerators for larger-scale models.

Intel과 Ampere는 각자의 CPU 플랫폼에서 실행되는 LLM을 성공적으로 시연했지만, 다양한 컴퓨팅 및 메모리 제한으로 인해 CPU가 대규모 모델의 GPU 또는 전용 가속기를 완전히 대체할 수 없다는 점을 인식하는 것이 중요합니다.

For models pushing the boundaries of generative AI, Ronak Shah, director of Xeon AI product management at Intel, emphasized that upcoming products like the Gaudi accelerator are specifically engineered for such tasks.

생성 AI의 경계를 넓히는 모델의 경우 Intel의 Xeon AI 제품 관리 이사인 Ronak Shah는 Gaudi 가속기와 같은 곧 출시될 제품이 이러한 작업을 위해 특별히 설계되었다고 강조했습니다.

Overcoming Bottlenecks

병목 현상 극복

Historically, conversations surrounding the execution of LLMs on CPUs have been subdued because, despite increasing core counts, conventional processors still fall short in terms of parallelism compared to modern GPUs and accelerators designed for AI workloads.

역사적으로 CPU에서의 LLM 실행을 둘러싼 대화는 코어 수가 증가함에도 불구하고 AI 워크로드용으로 설계된 최신 GPU 및 가속기에 비해 병렬성 측면에서 여전히 부족하기 때문에 차분해졌습니다.

However, CPUs are undergoing significant enhancements. Modern units dedicate a substantial portion of their die space to features such as vector extensions or even specialized matrix math accelerators.

그러나 CPU는 상당한 개선을 겪고 있습니다. 현대 장치는 다이 공간의 상당 부분을 벡터 확장이나 특수 행렬 수학 가속기와 같은 기능에 할당합니다.

Intel incorporated the latter feature in its Sapphire Rapids Xeon Scalable processors released early last year. Each core is also equipped with Advanced Matrix Extensions (AMX), although not all stock-keeping units (SKUs) support AMX due to the flexibility of software-defined silicon.

Intel은 작년 초에 출시된 Sapphire Rapids Xeon Scalable 프로세서에 후자 기능을 통합했습니다. 소프트웨어 정의 실리콘의 유연성으로 인해 모든 SKU(재고 관리 장치)가 AMX를 지원하는 것은 아니지만 각 코어에는 AMX(Advanced Matrix Extensions)도 탑재되어 있습니다.

As the name suggests, AMX extensions are tailored to accelerate matrix math calculations prevalent in deep learning workloads. Since its initial implementation, Intel has continuously refined its AMX engines for improved performance on larger models. This advancement is likely reflected in the upcoming Intel Xeon 6 processors slated for release later this year.

이름에서 알 수 있듯이 AMX 확장은 딥 러닝 워크로드에서 널리 사용되는 행렬 수학 계산을 가속화하도록 맞춤화되었습니다. Intel은 초기 구현 이후 더 큰 모델의 성능을 향상시키기 위해 AMX 엔진을 지속적으로 개선해 왔습니다. 이러한 발전은 올해 후반에 출시될 Intel Xeon 6 프로세서에 반영될 가능성이 높습니다.

While Intel heavily relies on matrix acceleration, Ampere's Wittich explained that acceptable performance can be achieved using the two 128-bit vector units embedded in each of its AmpereOne and Altra cores. These vector units support FP16, BF16, INT8, and INT16 precision levels.

Intel은 매트릭스 가속에 크게 의존하고 있지만 Ampere의 Wittich는 AmpereOne 및 Altra 코어 각각에 내장된 2개의 128비트 벡터 장치를 사용하여 수용 가능한 성능을 달성할 수 있다고 설명했습니다. 이러한 벡터 단위는 FP16, BF16, INT8 및 INT16 정밀도 수준을 지원합니다.

Memory Bottlenecks and MCR DIMMs

메모리 병목 현상 및 MCR DIMM

Despite their inferior performance in executing OPS or FLOPS compared to GPUs, CPUs possess a significant advantage: their independence from expensive and capacity-constrained high-bandwidth memory (HBM) modules.

GPU에 비해 ​​OPS 또는 FLOPS 실행 성능이 열등함에도 불구하고 CPU는 값비싸고 용량이 제한된 고대역폭 메모리(HBM) 모듈로부터 독립된다는 상당한 이점을 가지고 있습니다.

As previously discussed, operating a model at FP8/INT8 requires approximately 1 gigabyte (GB) of memory for each billion parameters. Consequently, executing a model like OpenAI's 1.7 trillion parameter GPT-4 model at FP8 would necessitate over 1.7 terabytes (TB) of memory, which would be approximately halved when quantized to 4-bits. This memory requirement exceeds the capacity of any single GPU but falls within the capabilities of modern CPUs.

이전에 설명한 대로 FP8/INT8에서 모델을 작동하려면 10억 개의 매개변수마다 약 1GB의 메모리가 필요합니다. 결과적으로 OpenAI의 1조 7천억 매개변수 GPT-4 모델과 같은 모델을 FP8에서 실행하려면 1.7테라바이트(TB) 이상의 메모리가 필요하며, 4비트로 양자화하면 대략 절반으로 줄어듭니다. 이 메모리 요구 사항은 단일 GPU의 용량을 초과하지만 최신 CPU의 성능에 속합니다.

However, the drawback lies in the sluggish speed of large DRAM modules used by CPUs compared to HBM.

하지만 CPU가 사용하는 대형 DRAM 모듈은 HBM에 비해 속도가 느린 것이 단점이다.

With only eight memory channels currently supported on Intel's 5th-generation Xeon and Ampere's One processors, these chips are limited to roughly 350 gigabytes per second (GB/sec) of memory bandwidth when running 5600MT/sec DIMMs. While Wittich mentioned plans for a 12-channel version of Ampere's chip with a targeted release later this year – featuring a purported 256 cores – it is not yet available.

현재 Intel의 5세대 Xeon 및 Ampere의 One 프로세서에서 8개의 메모리 채널만 지원되는 이 칩은 5600MT/초 DIMM을 실행할 때 약 350GB/초의 메모리 대역폭으로 제한됩니다. Wittich는 올해 후반에 256개 코어를 탑재할 예정인 Ampere 칩의 12채널 버전에 대한 계획을 언급했지만 아직 출시되지 않았습니다.

Nevertheless, all of Oracle's testing has been conducted on Ampere's Altra generation, which utilizes even slower DDR4 memory and operates at a maximum bandwidth of approximately 200GB/sec. This suggests the potential for substantial performance gains by upgrading to the newer AmpereOne cores.

그럼에도 불구하고 Oracle의 모든 테스트는 훨씬 더 느린 DDR4 메모리를 활용하고 약 200GB/초의 최대 대역폭에서 작동하는 Ampere의 Altra 세대에서 수행되었습니다. 이는 최신 AmpereOne 코어로 업그레이드하면 상당한 성능 향상이 가능함을 의미합니다.

These bandwidth speeds may seem impressive – certainly faster than an SSD – but the eight HBM modules found on AMD's MI300X or Nvidia's upcoming Blackwell GPUs deliver speeds of 5.3 TB/sec and 8TB/sec, respectively. However, HBM modules are constrained by a maximum capacity of 192GB.

이러한 대역폭 속도는 확실히 SSD보다 빠르며 인상적으로 보일 수 있지만 AMD의 MI300X 또는 곧 출시될 Nvidia의 Blackwell GPU에 있는 8개의 HBM 모듈은 각각 5.3TB/초 및 8TB/초의 속도를 제공합니다. 그러나 HBM 모듈은 최대 용량이 192GB로 제한됩니다.

To illustrate this concept, consider memory capacity as a fuel tank, memory bandwidth as a fuel line, and compute as an internal combustion engine. Regardless of the size of the fuel tank or the power of the engine, if the fuel line is too narrow to supply sufficient fuel for optimal engine performance, the system will be hindered.

이 개념을 설명하기 위해 메모리 용량을 연료 탱크로, 메모리 대역폭을 연료 라인으로, 컴퓨팅을 내연 기관으로 간주합니다. 연료 탱크의 크기나 엔진 출력에 관계없이 연료 라인이 너무 좁아 최적의 엔진 성능을 위한 충분한 연료를 공급할 수 없으면 시스템이 방해를 받게 됩니다.

This limitation explains why previous attempts to execute LLMs on CPUs were largely confined to smaller models.

이러한 제한은 이전에 CPU에서 LLM을 실행하려는 시도가 주로 더 작은 모델에 국한되었던 이유를 설명합니다.

Clearing the Bottlenecks: Intel's Granite Rapids Xeon 6

병목 현상 제거: Intel의 Granite Rapids Xeon 6

Despite these obstacles, Intel's forthcoming Granite Rapids Xeon 6 platform provides clues on how CPUs could potentially handle larger models in the near future.

이러한 장애물에도 불구하고 Intel이 곧 출시할 Granite Rapids Xeon 6 플랫폼은 가까운 미래에 CPU가 어떻게 더 큰 모델을 처리할 수 있는지에 대한 단서를 제공합니다.

Intel's recent demonstration showcased a single Xeon 6 processor effortlessly running Llama2-70B with a reasonable second token latency of 82ms. Crucially, many details regarding the test rig remain unknown, including the number and clock speed of the cores. These details will likely be revealed later this year – potentially in December.

Intel의 최근 시연에서는 82ms의 합리적인 두 번째 토큰 대기 시간으로 Llama2-70B를 쉽게 실행하는 단일 Xeon 6 프로세서를 선보였습니다. 결정적으로 코어 수와 클럭 속도를 포함하여 테스트 장비에 관한 많은 세부 사항이 알려지지 않았습니다. 이러한 세부 정보는 올해 말, 12월에 공개될 가능성이 높습니다.

"The substantial advancement from 5th-generation Xeon to Xeon 6 lies in the introduction of MCR DIMMs, which effectively clears many of the bottlenecks associated with memory-bound workloads," explained Shah.

Shah는 "5세대 Xeon에서 Xeon 6으로의 실질적인 발전은 MCR DIMM의 도입에 있습니다. 이는 메모리 바인딩된 작업 부하와 관련된 많은 병목 현상을 효과적으로 제거합니다."라고 설명했습니다.

Multiplexer combined rank (MCR) DIMMs facilitate much faster memory access compared to standard DRAM. Intel has already demonstrated this technology running at 8,800MT/sec. With 12 memory channels equipped with MCR DIMMs, a single Granite Rapids socket would access approximately 825GB/sec of bandwidth – a significant leap from previous generations.

MCR(멀티플렉서 결합 랭크) DIMM은 표준 DRAM에 비해 훨씬 빠른 메모리 액세스를 지원합니다. 인텔은 이미 8,800MT/초로 실행되는 이 기술을 시연했습니다. MCR DIMM이 장착된 12개의 메모리 채널을 사용하면 단일 Granite Rapids 소켓은 약 825GB/초의 대역폭에 액세스할 수 있습니다. 이는 이전 세대에 비해 크게 향상된 것입니다.

Wittich noted that Ampere is also exploring the implementation of MCR DIMMs but did not provide a timeline for their inclusion in Ampere silicon.

Wittich는 Ampere가 MCR DIMM 구현을 모색하고 있지만 Ampere 실리콘에 포함하기 위한 일정을 제공하지 않았다고 지적했습니다.

However, faster memory technology is not the sole innovation offered by Granite Rapids. Intel's AMX engine has gained support for 4-bit operations via the introduction of the MXFP4 data type, which theoretically has the potential to double effective performance.

그러나 더 빠른 메모리 기술은 Granite Rapids가 제공하는 유일한 혁신이 아닙니다. Intel의 AMX 엔진은 이론적으로 효과적인 성능을 두 배로 늘릴 수 있는 MXFP4 데이터 유형의 도입을 통해 4비트 작업에 대한 지원을 얻었습니다.

Moreover, lower precision reduces the model footprint and subsequently lowers memory capacity and bandwidth requirements. Quantization techniques employed to compress models trained at higher precisions can also achieve similar reductions in footprint and bandwidth usage. Therefore, the practical benefit of supporting 4-bit mathematics in hardware primarily manifests as performance enhancements.

또한 정밀도가 낮을수록 모델 공간이 줄어들고 결과적으로 메모리 용량 및 대역폭 요구 사항도 낮아집니다. 더 높은 정밀도로 훈련된 모델을 압축하는 데 사용되는 양자화 기술도 공간 및 대역폭 사용량을 비슷한 수준으로 줄일 수 있습니다. 따라서 하드웨어에서 4비트 연산을 지원하는 실질적인 이점은 주로 성능 향상으로 나타납니다.

Balancing Act: Optimizing CPU Design for AI

균형 조정법: AI를 위한 CPU 설계 최적화

For CPU designers, striking the right balance of AI capabilities presents a challenge. Excessive allocation of die area to features like AMX risks transforming the chip into an AI accelerator rather than a general-purpose processor.

CPU 설계자에게 AI 기능의 적절한 균형을 맞추는 것은 어려운 일입니다. AMX와 같은 기능에 다이 영역을 과도하게 할당하면 칩이 범용 프로세서가 아닌 AI 가속기로 변환될 위험이 있습니다.

As a result, instead of aiming for CPUs capable of handling the largest and most demanding LLMs, vendors are focusing on the distribution of AI models to identify the most widely adopted models and optimizing their products to cater to these workloads.

결과적으로 공급업체는 가장 크고 가장 까다로운 LLM을 처리할 수 있는 CPU를 목표로 하는 대신 AI 모델 배포에 집중하여 가장 널리 채택되는 모델을 식별하고 이러한 워크로드를 충족하도록 제품을 최적화하고 있습니다.

"From a customer perspective, the sweet spot right now revolves around models with 7–13 billion parameters. That's where most of our attention is directed today," stated Wittich.

Wittich는 "고객 관점에서 볼 때 현재 최적의 지점은 70억~130억 개의 매개변수를 가진 모델을 중심으로 돌아가고 있습니다. 오늘날 우리의 관심이 가장 집중되는 부분이 바로 여기에 있습니다"라고 Wittich는 말했습니다.

Intel's Shah has observed a similar

Intel의 Shah도 비슷한 현상을 관찰했습니다.

부인 성명:info@kdj.com

제공된 정보는 거래 조언이 아닙니다. kdj.com은 이 기사에 제공된 정보를 기반으로 이루어진 투자에 대해 어떠한 책임도 지지 않습니다. 암호화폐는 변동성이 매우 높으므로 철저한 조사 후 신중하게 투자하는 것이 좋습니다!

본 웹사이트에 사용된 내용이 귀하의 저작권을 침해한다고 판단되는 경우, 즉시 당사(info@kdj.com)로 연락주시면 즉시 삭제하도록 하겠습니다.

2026年07月28日 에 게재된 다른 기사