시가총액: $2.2001T 1.31%
거래량(24시간): $63.9638B -6.10%
  • 시가총액: $2.2001T 1.31%
  • 거래량(24시간): $63.9638B -6.10%
  • 공포와 탐욕 지수:
  • 시가총액: $2.2001T 1.31%
암호화
주제
암호화
소식
cryptostopics
비디오
최고의 뉴스
암호화
주제
암호화
소식
cryptostopics
비디오
bitcoin
bitcoin

$87959.907984 USD

1.34%

ethereum
ethereum

$2920.497338 USD

3.04%

tether
tether

$0.999775 USD

0.00%

xrp
xrp

$2.237324 USD

8.12%

bnb
bnb

$860.243768 USD

0.90%

solana
solana

$138.089498 USD

5.43%

usd-coin
usd-coin

$0.999807 USD

0.01%

tron
tron

$0.272801 USD

-1.53%

dogecoin
dogecoin

$0.150904 USD

2.96%

cardano
cardano

$0.421635 USD

1.97%

hyperliquid
hyperliquid

$32.152445 USD

2.23%

bitcoin-cash
bitcoin-cash

$533.301069 USD

-1.94%

chainlink
chainlink

$12.953417 USD

2.68%

unus-sed-leo
unus-sed-leo

$9.535951 USD

0.73%

zcash
zcash

$521.483386 USD

-2.87%

암호화폐 뉴스 기사

Zyda: 언어 모델링을 위한 획기적인 1조 3천억 토큰 공개 데이터 세트

2024/06/08 10:39

Zyphra는 언어 모델링을 위한 획기적인 1조 3천억 토큰 공개 데이터 세트인 Zyda의 출시를 발표했습니다. 이 혁신적인 데이터 세트는 언어 모델 훈련 및 연구의 표준을 재정의하도록 설정되었습니다.

Zyda: 언어 모델링을 위한 획기적인 1조 3천억 토큰 공개 데이터 세트

Zyphra, a pioneer in the field of natural language processing (NLP), has announced the release of Zyda, a groundbreaking 1.3 trillion-token open dataset for language modeling. This massive and meticulously curated dataset is poised to redefine the standards of language model training and research, offering an unparalleled combination of size, quality, and accessibility.

자연어 처리(NLP) 분야의 선구자인 Zyphra는 언어 모델링을 위한 획기적인 1조 3천억 토큰 공개 데이터 세트인 Zyda의 출시를 발표했습니다. 이 방대하고 꼼꼼하게 선별된 데이터 세트는 크기, 품질 및 접근성의 비교할 수 없는 조합을 제공하여 언어 모델 훈련 및 연구의 표준을 재정의할 준비가 되어 있습니다.

To create Zyda, several high-quality open datasets were combined and refined through a rigorous filtering and deduplication process. The resulting dataset boasts an impressive token count while maintaining the highest data quality standards. Zyda is primarily designed to facilitate advanced language modeling experiments and training at a scale that was previously unattainable with open datasets.

Zyda를 만들기 위해 여러 가지 고품질 공개 데이터 세트를 결합하고 엄격한 필터링 및 중복 제거 프로세스를 통해 개선했습니다. 결과 데이터 세트는 최고의 데이터 품질 표준을 유지하면서 인상적인 토큰 수를 자랑합니다. Zyda는 주로 이전에 공개 데이터세트로는 달성할 수 없었던 규모로 고급 언어 모델링 실험과 교육을 용이하게 하도록 설계되었습니다.

In comprehensive ablation studies, Zyda has consistently outperformed existing datasets, including Dolma, Fineweb, Pile, RefinedWeb, and SlimPajama. This makes Zyda a crucial resource for researchers and developers looking to contribute to the field of language modeling.

포괄적인 절제 연구에서 Zyda는 Dolma, Fineweb, Pile, RefinedWeb 및 SlimPajama를 포함한 기존 데이터세트보다 지속적으로 뛰어난 성능을 보여왔습니다. 이로 인해 Zyda는 언어 모델링 분야에 기여하려는 연구자와 개발자에게 중요한 리소스가 되었습니다.

Key Features of Zyda

Zyda의 주요 기능

Zyda was meticulously crafted by merging seven well-respected open language modeling datasets: RefinedWeb, Starcoder, C4, Pile, Slimpajama, pe2so, and arXiv. Each dataset underwent a uniform post-processing pipeline designed to enhance quality and coherence.

Zyda는 RefinedWeb, Starcoder, C4, Pile, Slimpajama, pe2so 및 arXiv 등 7개의 존경받는 개방형 언어 모델링 데이터세트를 병합하여 꼼꼼하게 제작되었습니다. 각 데이터세트는 품질과 일관성을 향상시키도록 설계된 균일한 후처리 파이프라인을 거쳤습니다.

The creation process involved thorough syntactic filtering to eliminate low-quality documents, followed by an aggressive cross-deduplication pass. Many datasets contained significant overlaps due to common data sources like Common Crawl, making cross-deduplication particularly important. This extensive cleaning process reduced the initial 2 trillion tokens to a more refined and manageable 1.3 trillion.

생성 프로세스에는 품질이 낮은 문서를 제거하기 위한 철저한 구문 필터링과 공격적인 교차 중복 제거 통과가 포함되었습니다. 많은 데이터 세트에는 Common Crawl과 같은 공통 데이터 소스로 인해 상당한 중복이 포함되어 있어 교차 중복 제거가 특히 중요했습니다. 이 광범위한 정리 과정을 통해 초기 2조 토큰을 보다 세련되고 관리하기 쉬운 1조 3천억 토큰으로 줄였습니다.

The effectiveness of Zyda is evident in the performance of Zamba, a language model trained on Zyda. When compared to models trained on competing datasets, Zamba demonstrates significant strength on a per-token basis. This serves as a testament to Zyda’s superior quality and potential to drive language modeling advancements.

Zyda의 효율성은 Zyda에서 훈련된 언어 모델인 Zamba의 성능에서 분명하게 드러납니다. 경쟁 데이터 세트로 훈련된 모델과 비교할 때 Zamba는 토큰 단위로 상당한 강점을 보여줍니다. 이는 Zyda의 우수한 품질과 언어 모델링 발전을 주도할 잠재력을 입증하는 역할을 합니다.

In conclusion, Zyda represents a monumental leap forward in the field of language modeling. By providing a massive, high-quality, open dataset, Zyphra is paving the way for the next generation of NLP research and applications. The release of Zyda not only underscores Zyphra’s leadership in the field but also sets a new benchmark for what is possible with open datasets.

결론적으로 Zyda는 언어 모델링 분야에서 기념비적인 도약을 나타냅니다. 대규모의 고품질 개방형 데이터 세트를 제공함으로써 Zyphra는 차세대 NLP 연구 및 애플리케이션을 위한 길을 닦고 있습니다. Zyda의 출시는 해당 분야에서 Zyphra의 리더십을 강조할 뿐만 아니라 공개 데이터 세트로 가능한 것에 대한 새로운 벤치마크를 설정합니다.

부인 성명:info@kdj.com

제공된 정보는 거래 조언이 아닙니다. kdj.com은 이 기사에 제공된 정보를 기반으로 이루어진 투자에 대해 어떠한 책임도 지지 않습니다. 암호화폐는 변동성이 매우 높으므로 철저한 조사 후 신중하게 투자하는 것이 좋습니다!

본 웹사이트에 사용된 내용이 귀하의 저작권을 침해한다고 판단되는 경우, 즉시 당사(info@kdj.com)로 연락주시면 즉시 삭제하도록 하겠습니다.

2026年07月30日 에 게재된 다른 기사