|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Zyphra は、言語モデリング用の画期的な 1 兆 3,000 億トークンのオープン データセットである Zyda のリリースを発表しました。この革新的なデータセットは、言語モデルのトレーニングと研究の標準を再定義するために設定されています

Zyphra, a pioneer in the field of natural language processing (NLP), has announced the release of Zyda, a groundbreaking 1.3 trillion-token open dataset for language modeling. This massive and meticulously curated dataset is poised to redefine the standards of language model training and research, offering an unparalleled combination of size, quality, and accessibility.
自然言語処理 (NLP) 分野のパイオニアである Zyphra は、言語モデリング用の画期的な 1 兆 3,000 億トークンのオープン データセットである Zyda のリリースを発表しました。この大規模で細心の注意を払って厳選されたデータセットは、言語モデルのトレーニングと研究の標準を再定義する準備ができており、サイズ、品質、アクセシビリティの比類のない組み合わせを提供します。
To create Zyda, several high-quality open datasets were combined and refined through a rigorous filtering and deduplication process. The resulting dataset boasts an impressive token count while maintaining the highest data quality standards. Zyda is primarily designed to facilitate advanced language modeling experiments and training at a scale that was previously unattainable with open datasets.
Zyda を作成するために、いくつかの高品質のオープン データセットが結合され、厳密なフィルタリングと重複排除のプロセスを通じて洗練されました。結果として得られるデータセットは、最高のデータ品質基準を維持しながら、驚異的なトークン数を誇ります。 Zyda は主に、オープン データセットでは以前は達成できなかった規模での高度な言語モデリングの実験とトレーニングを容易にするように設計されています。
In comprehensive ablation studies, Zyda has consistently outperformed existing datasets, including Dolma, Fineweb, Pile, RefinedWeb, and SlimPajama. This makes Zyda a crucial resource for researchers and developers looking to contribute to the field of language modeling.
包括的なアブレーション研究において、Zyda は、Dolma、Fineweb、Pile、RefinedWeb、SlimPajama などの既存のデータセットを常に上回っています。このため、言語モデリングの分野に貢献したいと考えている研究者や開発者にとって、Zyda は重要なリソースになります。
Key Features of Zyda
Zydaの主な特徴
Zyda was meticulously crafted by merging seven well-respected open language modeling datasets: RefinedWeb, Starcoder, C4, Pile, Slimpajama, pe2so, and arXiv. Each dataset underwent a uniform post-processing pipeline designed to enhance quality and coherence.
Zyda は、7 つの定評あるオープン言語モデリング データセット (RefinedWeb、Starcoder、C4、Pile、Slimpajama、pe2so、arXiv) を統合して細心の注意を払って作成されました。各データセットは、品質と一貫性を高めるために設計された均一な後処理パイプラインを受けました。
The creation process involved thorough syntactic filtering to eliminate low-quality documents, followed by an aggressive cross-deduplication pass. Many datasets contained significant overlaps due to common data sources like Common Crawl, making cross-deduplication particularly important. This extensive cleaning process reduced the initial 2 trillion tokens to a more refined and manageable 1.3 trillion.
作成プロセスには、低品質のドキュメントを排除するための徹底的な構文フィルタリングと、それに続く積極的な相互重複排除パスが含まれていました。多くのデータセットには、Common Crawl などの一般的なデータ ソースによる重大な重複が含まれており、相互重複排除が特に重要になっています。この大規模なクリーニング プロセスにより、最初の 2 兆トークンが、より洗練され管理しやすい 1 兆 3,000 億トークンに減りました。
The effectiveness of Zyda is evident in the performance of Zamba, a language model trained on Zyda. When compared to models trained on competing datasets, Zamba demonstrates significant strength on a per-token basis. This serves as a testament to Zyda’s superior quality and potential to drive language modeling advancements.
Zyda の有効性は、Zyda でトレーニングされた言語モデルである Zamba のパフォーマンスで明らかです。競合するデータセットでトレーニングされたモデルと比較すると、Zamba はトークンごとに大きな強みを示します。これは、Zyda の優れた品質と言語モデリングの進歩を推進する可能性の証拠として機能します。
In conclusion, Zyda represents a monumental leap forward in the field of language modeling. By providing a massive, high-quality, open dataset, Zyphra is paving the way for the next generation of NLP research and applications. The release of Zyda not only underscores Zyphra’s leadership in the field but also sets a new benchmark for what is possible with open datasets.
結論として、Zyda は言語モデリングの分野における画期的な進歩を表しています。 Zyphra は、大規模で高品質のオープン データセットを提供することで、次世代の NLP 研究とアプリケーションへの道を切り開いています。 Zyda のリリースは、この分野における Zyphra のリーダーシップを強調するだけでなく、オープン データセットで何が可能になるかについての新しいベンチマークを設定します。
免責事項:info@kdj.com
提供される情報は取引に関するアドバイスではありません。 kdj.com は、この記事で提供される情報に基づいて行われた投資に対して一切の責任を負いません。暗号通貨は変動性が高いため、十分な調査を行った上で慎重に投資することを強くお勧めします。
このウェブサイトで使用されているコンテンツが著作権を侵害していると思われる場合は、直ちに当社 (info@kdj.com) までご連絡ください。速やかに削除させていただきます。

































