市值: $2.2032T 1.06%
成交额(24h): $37.9282B -32.34%
  • 市值: $2.2032T 1.06%
  • 成交额(24h): $37.9282B -32.34%
  • 恐惧与贪婪指数:
  • 市值: $2.2032T 1.06%
加密货币
话题
百科
资讯
加密话题
视频
热门新闻
加密货币
话题
百科
资讯
加密话题
视频
bitcoin
bitcoin

$87959.907984 USD

1.34%

ethereum
ethereum

$2920.497338 USD

3.04%

tether
tether

$0.999775 USD

0.00%

xrp
xrp

$2.237324 USD

8.12%

bnb
bnb

$860.243768 USD

0.90%

solana
solana

$138.089498 USD

5.43%

usd-coin
usd-coin

$0.999807 USD

0.01%

tron
tron

$0.272801 USD

-1.53%

dogecoin
dogecoin

$0.150904 USD

2.96%

cardano
cardano

$0.421635 USD

1.97%

hyperliquid
hyperliquid

$32.152445 USD

2.23%

bitcoin-cash
bitcoin-cash

$533.301069 USD

-1.94%

chainlink
chainlink

$12.953417 USD

2.68%

unus-sed-leo
unus-sed-leo

$9.535951 USD

0.73%

zcash
zcash

$521.483386 USD

-2.87%

加密货币新闻

NVIDIA 正在帮助 Apple 构建更快、更好的 AI 体验

2024/12/20 19:52

如果您通过 BGR 链接购买,我们可能会赚取附属佣金,帮助支持我们的专家产品实验室。 Apple 和 NVIDIA 分享了合作细节

NVIDIA 正在帮助 Apple 构建更快、更好的 AI 体验

Tech giants Apple and NVIDIA have joined forces to enhance the performance of Large Language Models (LLMs) by introducing a new text generation technique for AI.

科技巨头 Apple 和 NVIDIA 联手通过引入新的 AI 文本生成技术来增强大型语言模型 (LLM) 的性能。

According to Apple, accelerating LLM inference is a crucial ML research problem. This is because auto-regressive token generation is computationally expensive and relatively slow. As a result, improving inference efficiency can reduce latency for users.

Apple 表示,加速 LLM 推理是一个至关重要的 ML 研究问题。这是因为自回归令牌生成的计算成本较高且相对较慢。因此,提高推理效率可以减少用户的延迟。

In addition to ongoing efforts to accelerate inference on Apple silicon, the company has recently made significant progress in accelerating LLM inference for the NVIDIA GPUs widely used for production applications across the industry, the company writes in a research paper.

该公司在一份研究论文中写道,除了持续努力加速 Apple 芯片上的推理之外,该公司最近还在加速广泛用于整个行业生产应用的 NVIDIA GPU 的 LLM 推理方面取得了重大进展。

Earlier this year, Apple published and open-sourced Recurrent Drafter (ReDrafter), which is a novel approach to speculative decoding that “achieves state of the art performance.”

今年早些时候,Apple 发布并开源了 Recurrent Drafter (ReDrafter),这是一种新颖的推测解码方法,“实现了最先进的性能”。

According to the company, ReDrafter uses an RNN draft model, and combines beam search with dynamic tree attention to speed up LLM token generation by up to 3.5 tokens per generation step for open source models, surpassing the performance of prior speculative decoding techniques.

据该公司称,ReDrafter 使用 RNN 草案模型,并将波束搜索与动态树注意力相结合,将开源模型的 LLM 令牌生成速度提高到每生成步 3.5 个令牌,超越了之前的推测解码技术的性能。

“In benchmarking a tens-of-billions parameter production model on NVIDIA GPUs, using the NVIDIA TensorRT-LLM inference acceleration framework with ReDrafter, we have seen 2.7x speed-up in generated tokens per second for greedy decoding,” Apple papers show.

苹果论文显示:“在 NVIDIA GPU 上对数百亿个参数生产模型进行基准测试时,使用 NVIDIA TensorRT-LLM 推理加速框架和 ReDrafter,我们发现每秒生成的贪婪解码令牌速度提高了 2.7 倍。”

With that, this technology could signifanctly reduce latency users may experience, while also using fewer GPUs and consuming less power.

这样,该技术可以显着减少用户可能遇到的延迟,同时使用更少的 GPU 并消耗更少的电量。

原文来源:bgr

免责声明:info@kdj.com

所提供的信息并非交易建议。根据本文提供的信息进行的任何投资,kdj.com不承担任何责任。加密货币具有高波动性,强烈建议您深入研究后,谨慎投资!

如您认为本网站上使用的内容侵犯了您的版权,请立即联系我们(info@kdj.com),我们将及时删除。

2026年07月26日 发表的其他文章