市值: $2.1622T -0.84%
體積(24小時): $50.0999B -16.65%
  • 市值: $2.1622T -0.84%
  • 體積(24小時): $50.0999B -16.65%
  • 恐懼與貪婪指數:
  • 市值: $2.1622T -0.84%
加密
主題
加密植物
資訊
加密術
影片
頭號新聞
加密
主題
加密植物
資訊
加密術
影片
bitcoin
bitcoin

$87959.907984 USD

1.34%

ethereum
ethereum

$2920.497338 USD

3.04%

tether
tether

$0.999775 USD

0.00%

xrp
xrp

$2.237324 USD

8.12%

bnb
bnb

$860.243768 USD

0.90%

solana
solana

$138.089498 USD

5.43%

usd-coin
usd-coin

$0.999807 USD

0.01%

tron
tron

$0.272801 USD

-1.53%

dogecoin
dogecoin

$0.150904 USD

2.96%

cardano
cardano

$0.421635 USD

1.97%

hyperliquid
hyperliquid

$32.152445 USD

2.23%

bitcoin-cash
bitcoin-cash

$533.301069 USD

-1.94%

chainlink
chainlink

$12.953417 USD

2.68%

unus-sed-leo
unus-sed-leo

$9.535951 USD

0.73%

zcash
zcash

$521.483386 USD

-2.87%

加密貨幣新聞文章

Zyda:用於語言建模的突破性 1.3 兆代幣開放資料集

2024/06/08 10:39

Zyphra 宣布發布 Zyda,這是一個用於語言建模的突破性的包含 1.3 兆代幣的開放資料集。這創新資料集將重新定義語言模型訓練和研究的標準

Zyda:用於語言建模的突破性 1.3 兆代幣開放資料集

Zyphra, a pioneer in the field of natural language processing (NLP), has announced the release of Zyda, a groundbreaking 1.3 trillion-token open dataset for language modeling. This massive and meticulously curated dataset is poised to redefine the standards of language model training and research, offering an unparalleled combination of size, quality, and accessibility.

自然語言處理 (NLP) 領域的先驅 Zyphra 宣布發布 Zyda,這是一個突破性的、包含 1.3 兆代幣的語言建模開放資料集。這個龐大且精心策劃的資料集將重新定義語言模型訓練和研究的標準,提供無與倫比的規模、品質和可訪問性組合。

To create Zyda, several high-quality open datasets were combined and refined through a rigorous filtering and deduplication process. The resulting dataset boasts an impressive token count while maintaining the highest data quality standards. Zyda is primarily designed to facilitate advanced language modeling experiments and training at a scale that was previously unattainable with open datasets.

為了創建 Zyda,我們透過嚴格的過濾和重複資料刪除流程組合併完善了多個高品質的開放資料集。產生的資料集擁有令人印象深刻的令牌數量,同時保持最高的資料品質標準。 Zyda 的主要目的是促進高階語言建模實驗和訓練,其規模是以前使用開放資料集無法實現的。

In comprehensive ablation studies, Zyda has consistently outperformed existing datasets, including Dolma, Fineweb, Pile, RefinedWeb, and SlimPajama. This makes Zyda a crucial resource for researchers and developers looking to contribute to the field of language modeling.

在綜合消融研究中,Zyda 始終優於現有資料集,包括 Dolma、Fineweb、Pile、RefinedWeb 和 SlimPajama。這使得 Zyda 成為希望為語言建模領域做出貢獻的研究人員和開發人員的重要資源。

Key Features of Zyda

Zyda 的主要特點

Zyda was meticulously crafted by merging seven well-respected open language modeling datasets: RefinedWeb, Starcoder, C4, Pile, Slimpajama, pe2so, and arXiv. Each dataset underwent a uniform post-processing pipeline designed to enhance quality and coherence.

Zyda 是透過合併七個備受推崇的開放語言建模資料集精心打造的:RefinedWeb、Starcoder、C4、Pile、Slimpajama、pe2so 和 arXiv。每個資料集都經過統一的後處理流程,旨在提高品質和一致性。

The creation process involved thorough syntactic filtering to eliminate low-quality documents, followed by an aggressive cross-deduplication pass. Many datasets contained significant overlaps due to common data sources like Common Crawl, making cross-deduplication particularly important. This extensive cleaning process reduced the initial 2 trillion tokens to a more refined and manageable 1.3 trillion.

建立過程涉及徹底的語法過濾以消除低品質文檔,然後進行積極的交叉重複資料刪除。由於 Common Crawl 等常見資料來源,許多資料集包含大量重疊,因此交叉重複資料刪除尤其重要。這種廣泛的清理過程將最初的 2 兆代幣減少到更精細和易於管理的 1.3 兆。

The effectiveness of Zyda is evident in the performance of Zamba, a language model trained on Zyda. When compared to models trained on competing datasets, Zamba demonstrates significant strength on a per-token basis. This serves as a testament to Zyda’s superior quality and potential to drive language modeling advancements.

Zyda 的有效性在 Zamba(一種在 Zyda 上訓練的語言模型)的表現中得到了體現。與在競爭資料集上訓練的模型相比,Zamba 在每個代幣的基礎上表現出了顯著的優勢。這證明了 Zyda 的卓越品質和推動語言建模進步的潛力。

In conclusion, Zyda represents a monumental leap forward in the field of language modeling. By providing a massive, high-quality, open dataset, Zyphra is paving the way for the next generation of NLP research and applications. The release of Zyda not only underscores Zyphra’s leadership in the field but also sets a new benchmark for what is possible with open datasets.

總之,Zyda 代表了語言建模領域的巨大飛躍。透過提供大量、高品質、開放的資料集,Zyphra 正在為下一代 NLP 研究和應用鋪平道路。 Zyda 的發布不僅凸顯了 Zyphra 在該領域的領導地位,也為開放資料集的可能性樹立了新的基準。

免責聲明:info@kdj.com

所提供的資訊並非交易建議。 kDJ.com對任何基於本文提供的資訊進行的投資不承擔任何責任。加密貨幣波動性較大,建議您充分研究後謹慎投資!

如果您認為本網站使用的內容侵犯了您的版權,請立即聯絡我們(info@kdj.com),我們將及時刪除。

2026年08月02日 其他文章發表於