시가총액: $2.2316T 1.41%
거래량(24시간): $50.2264B 30.30%
  • 시가총액: $2.2316T 1.41%
  • 거래량(24시간): $50.2264B 30.30%
  • 공포와 탐욕 지수:
  • 시가총액: $2.2316T 1.41%
암호화
주제
암호화
소식
cryptostopics
비디오
최고의 뉴스
암호화
주제
암호화
소식
cryptostopics
비디오
bitcoin
bitcoin

$87959.907984 USD

1.34%

ethereum
ethereum

$2920.497338 USD

3.04%

tether
tether

$0.999775 USD

0.00%

xrp
xrp

$2.237324 USD

8.12%

bnb
bnb

$860.243768 USD

0.90%

solana
solana

$138.089498 USD

5.43%

usd-coin
usd-coin

$0.999807 USD

0.01%

tron
tron

$0.272801 USD

-1.53%

dogecoin
dogecoin

$0.150904 USD

2.96%

cardano
cardano

$0.421635 USD

1.97%

hyperliquid
hyperliquid

$32.152445 USD

2.23%

bitcoin-cash
bitcoin-cash

$533.301069 USD

-1.94%

chainlink
chainlink

$12.953417 USD

2.68%

unus-sed-leo
unus-sed-leo

$9.535951 USD

0.73%

zcash
zcash

$521.483386 USD

-2.87%

암호화폐 뉴스 기사

NVIDIA의 HGX H200 AI 가속기는 NVIDIA 독점 디코딩 알고리즘 "Medusa"를 통해 Llama 3.1 추론을 크게 향상시킵니다.

2024/09/08 17:00

성능은 초고속 GPU 간 통신과 고급 소프트웨어를 갖춘 "하나의 강력한 GPU"로 요청을 처리하는 결합된 GPU의 능력에 따라 달라집니다.

NVIDIA의 HGX H200 AI 가속기는 NVIDIA 독점 디코딩 알고리즘 "Medusa"를 통해 Llama 3.1 추론을 크게 향상시킵니다.

NVIDIA's HGX H200 AI accelerators are getting a big boost in Llama 3.1 inferencing, thanks to a new NVIDIA-exclusive decoding algorithm called "Medusa."

NVIDIA의 HGX H200 AI 가속기는 "Medusa"라는 새로운 NVIDIA 독점 디코딩 알고리즘 덕분에 Llama 3.1 추론에서 큰 향상을 얻고 있습니다.

As large language models (LLMs) continue to grow in size and complexity, multi-GPU compute is a must-have to deliver the low latency and high throughput that real-time generative AI applications demand.

대규모 언어 모델(LLM)의 크기와 복잡성이 지속적으로 증가함에 따라 실시간 생성 AI 애플리케이션이 요구하는 짧은 대기 시간과 높은 처리량을 제공하려면 다중 GPU 컴퓨팅이 필수입니다.

Performance depends on the combined GPUs' ability to work together as “one mighty GPU” with ultra-fast GPU-to-GPU communication and advanced software able to take full advantage of the multiple GPUs. By splitting the calculations of each model layer across the available GPUs using a technique called tensor parallelism in tandem with advanced algorithms like speculative decoding, token generation latency can be reduced, delivering an interactive user experience.

성능은 초고속 GPU 간 통신과 여러 GPU를 최대한 활용할 수 있는 고급 소프트웨어를 갖춘 "하나의 강력한 GPU"로 함께 작동하는 결합된 GPU의 능력에 따라 달라집니다. 추측적 디코딩과 같은 고급 알고리즘과 함께 텐서 병렬성이라는 기술을 사용하여 사용 가능한 GPU에서 각 모델 계층의 계산을 분할함으로써 토큰 생성 대기 시간을 줄이고 대화형 사용자 경험을 제공할 수 있습니다.

For very low latency Llama 3.1 serving, cloud services can use a full NVIDIA HGX H200 server, each incorporating eight H200 Tensor Core GPUs and four all-to-all NVLink Switch chips. Each GPU within the server can communicate at the full 900 GB/s bandwidth to any other GPU via NVLink Switch. High GPU-to-GPU fabric bandwidth is required to keep multi-GPU communication from becoming the bottleneck in interactive use cases.

대기 시간이 매우 짧은 Llama 3.1 서비스를 위해 클라우드 서비스는 각각 8개의 H200 Tensor Core GPU와 4개의 전체 NVLink 스위치 칩을 통합하는 전체 NVIDIA HGX H200 서버를 사용할 수 있습니다. 서버 내의 각 GPU는 NVLink 스위치를 통해 전체 900GB/s 대역폭으로 다른 GPU와 통신할 수 있습니다. 대화형 사용 사례에서 다중 GPU 통신이 병목 현상을 일으키지 않도록 하려면 높은 GPU-GPU 패브릭 대역폭이 필요합니다.

To efficiently implement optimization algorithms on NVIDIA H200 HGX systems, NVIDIA TensorRT-LLM is used. TensorRT-LLM is an open-source TensorRT library that delivers state-of-the-art inference performance on the latest LLMs using a variety of techniques, including tensor parallelism and speculative decoding.

NVIDIA H200 HGX 시스템에서 최적화 알고리즘을 효율적으로 구현하기 위해 NVIDIA TensorRT-LLM이 사용됩니다. TensorRT-LLM은 텐서 병렬 처리 및 추측 디코딩을 포함한 다양한 기술을 사용하여 최신 LLM에 대한 최첨단 추론 성능을 제공하는 오픈 소스 TensorRT 라이브러리입니다.

Upcoming TensorRT-LLM optimizations, including the improvement of a speculative decoding algorithm called Medusa, provide outstanding low latency performance on Llama 3.1 70B and Llama 3.1 405B of 268 tokens/second/user and 108 tokens/second/user, respectively on HGX H200.

Medusa라는 예측적 디코딩 알고리즘의 개선을 포함하여 향후 TensorRT-LLM 최적화는 Llama 3.1 70B 및 Llama 3.1 405B에서 HGX H200에서 각각 268개 토큰/초/사용자 및 108개 토큰/초/사용자의 탁월한 낮은 대기 시간 성능을 제공합니다.

Medusa boosts token generation by up to 1.9x on NVIDIA HGX H200

Medusa는 NVIDIA HGX H200에서 토큰 생성을 최대 1.9배 향상시킵니다.

Transformer-based LLMs are auto-regressive, meaning that tokens need to be generated sequentially, limiting throughput per generation step to just one token. Typically, during LLM inference, the rate at which a single token is generated depends on how quickly model weights are loaded into memory. This means that the workload can leave the substantial Tensor Core capabilities of H200 GPUs underutilized.

Transformer 기반 LLM은 자동 회귀적입니다. 즉, 토큰을 순차적으로 생성해야 하며 생성 단계당 처리량을 하나의 토큰으로 제한합니다. 일반적으로 LLM 추론 중에 단일 토큰이 생성되는 속도는 모델 가중치가 메모리에 로드되는 속도에 따라 달라집니다. 이는 워크로드로 인해 H200 GPU의 상당한 Tensor Core 기능이 제대로 활용되지 않을 수 있음을 의미합니다.

Speculative decoding is a technique that increases token generation throughput per token generation step by using a “draft model” to try to predict multiple subsequent tokens beyond the next token. The target LLM then “batches” the prediction candidates and validates them in parallel with the next token, making more effective use of available parallel GPU compute resources. If the original LLM accepts any candidate sequence, multiple tokens are generated in the generation step and therefore accelerate token generation.

추측적 디코딩은 "초안 모델"을 사용하여 다음 토큰 이후에 여러 후속 토큰을 예측하려고 시도하여 토큰 생성 단계당 토큰 생성 처리량을 늘리는 기술입니다. 그런 다음 대상 LLM은 예측 후보를 "일괄 처리"하고 다음 토큰과 병렬로 유효성을 검사하여 사용 가능한 병렬 GPU 컴퓨팅 리소스를 보다 효과적으로 활용합니다. 원래 LLM이 후보 시퀀스를 수락하면 생성 단계에서 여러 토큰이 생성되므로 토큰 생성이 가속화됩니다.

Medusa, described in this paper, is a speculative decoding algorithm that uses the original model as the draft model, avoiding the system complexity and distribution discrepancy of using a separate draft model. This technique employs additional decoding “heads”, called Medusa heads, to predict candidate tokens beyond the next token. Each Medusa head generates a distribution of tokens beyond the previous.

본 논문에서 설명하는 Medusa는 원본 모델을 초안 모델로 사용하여 별도의 초안 모델을 사용하는 데 따른 시스템 복잡성과 분포 불일치를 피하는 추측적 디코딩 알고리즘입니다. 이 기술은 Medusa 헤드라고 하는 추가 디코딩 "헤드"를 사용하여 다음 토큰 이후의 후보 토큰을 예측합니다. 각 메두사 헤드는 이전 것보다 더 많은 토큰 배포를 생성합니다.

With Medusa, an HGX H200 can produce 268 tokens per second per user for Llama 3.1 70B and 108 for Llama 3.1 405B. This is over 1.5x faster on Llama 3.1 70B and over 1.9x faster on Llama 3.1 405B than without Medusa. Although there is variability in the Medusa acceptance rate between tasks depending on how the heads are fine-tuned, its overall performance is generalized across a wide range of tasks.

Medusa를 사용하면 HGX H200은 Llama 3.1 70B의 경우 사용자당 초당 268개의 토큰을, Llama 3.1 405B의 경우 108개의 토큰을 생성할 수 있습니다. 이는 Medusa가 없을 때보다 Llama 3.1 70B에서는 1.5배 이상 빠르고 Llama 3.1 405B에서는 1.9배 이상 빠릅니다. 헤드를 어떻게 미세 조정하느냐에 따라 작업 간 메두사 수용률에 차이가 있지만 전반적인 성능은 광범위한 작업에 걸쳐 일반화됩니다.

Medusa heads for both Llama 3.1 70B and Llama 3.1 405B were trained using the NVIDIA TensorRT Model Optimizer integration with the NVIDIA NeMo framework. The Medusa head training used a frozen backbone, ensuring that the use of Medusa yields identical accuracy to the base model.

Llama 3.1 70B 및 Llama 3.1 405B의 Medusa 헤드는 NVIDIA NeMo 프레임워크와 NVIDIA TensorRT Model Optimizer 통합을 사용하여 교육되었습니다. Medusa 머리 훈련에서는 고정 백본을 사용하여 Medusa 사용이 기본 모델과 동일한 정확도를 제공하도록 보장했습니다.

NVIDIA full-stack innovation never stops

NVIDIA 풀스택 혁신은 결코 멈추지 않습니다

NVIDIA HGX H200 with NVLink Switch and TensorRT-LLM already deliver excellent real-time inference performance on popular and demanding community models. To continue improving user experiences and reduce inference costs, we relentlessly innovate across every layer of the technology stack – chips, systems, software libraries, algorithms, and more.

NVLink 스위치 및 TensorRT-LLM을 탑재한 NVIDIA HGX H200은 이미 인기 있고 까다로운 커뮤니티 모델에서 탁월한 실시간 추론 성능을 제공하고 있습니다. 사용자 경험을 지속적으로 개선하고 추론 비용을 줄이기 위해 우리는 칩, 시스템, 소프트웨어 라이브러리, 알고리즘 등 기술 스택의 모든 계층에서 끊임없이 혁신을 이루고 있습니다.

We look forward to sharing future updates on our low latency inference performance as both our platform and the LLM ecosystem advances.

플랫폼과 LLM 생태계가 모두 발전함에 따라 낮은 대기 시간 추론 성능에 대한 향후 업데이트를 공유할 수 있기를 기대합니다.

원본 소스:wccftech

부인 성명:info@kdj.com

제공된 정보는 거래 조언이 아닙니다. kdj.com은 이 기사에 제공된 정보를 기반으로 이루어진 투자에 대해 어떠한 책임도 지지 않습니다. 암호화폐는 변동성이 매우 높으므로 철저한 조사 후 신중하게 투자하는 것이 좋습니다!

본 웹사이트에 사용된 내용이 귀하의 저작권을 침해한다고 판단되는 경우, 즉시 당사(info@kdj.com)로 연락주시면 즉시 삭제하도록 하겠습니다.

2026年07月28日 에 게재된 다른 기사