시가총액: $2.1641T -1.03%
거래량(24시간): $55.4357B -12.30%
  • 시가총액: $2.1641T -1.03%
  • 거래량(24시간): $55.4357B -12.30%
  • 공포와 탐욕 지수:
  • 시가총액: $2.1641T -1.03%
암호화
주제
암호화
소식
cryptostopics
비디오
최고의 뉴스
암호화
주제
암호화
소식
cryptostopics
비디오
bitcoin
bitcoin

$87959.907984 USD

1.34%

ethereum
ethereum

$2920.497338 USD

3.04%

tether
tether

$0.999775 USD

0.00%

xrp
xrp

$2.237324 USD

8.12%

bnb
bnb

$860.243768 USD

0.90%

solana
solana

$138.089498 USD

5.43%

usd-coin
usd-coin

$0.999807 USD

0.01%

tron
tron

$0.272801 USD

-1.53%

dogecoin
dogecoin

$0.150904 USD

2.96%

cardano
cardano

$0.421635 USD

1.97%

hyperliquid
hyperliquid

$32.152445 USD

2.23%

bitcoin-cash
bitcoin-cash

$533.301069 USD

-1.94%

chainlink
chainlink

$12.953417 USD

2.68%

unus-sed-leo
unus-sed-leo

$9.535951 USD

0.73%

zcash
zcash

$521.483386 USD

-2.87%

암호화폐 뉴스 기사

카멜레온: 원활한 다중 모드 문서 생성을 위한 혁신적인 모델

2024/05/18 15:04

메타 연구자들은 이미지와 텍스트 시퀀스를 완벽하게 통합하는 다중 모드 기반 모델인 Chameleon을 개발했습니다. 양식을 분리하는 기존 모델과 달리 Chameleon의 통합 아키텍처는 텍스트와 같은 이미지를 토큰화하여 양식 전반에 걸쳐 포괄적인 문서 모델링과 인터리브 추론을 가능하게 합니다. 이러한 초기 융합 접근 방식의 최적화 문제를 해결하기 위해 연구원들은 아키텍처 개선 및 훈련 기술을 도입하여 Meta의 대규모 데이터 세트에 대한 성공적인 훈련을 달성했습니다. Chameleon의 평가는 텍스트 전용 작업에서 경쟁력 있는 성능을 보여주며 상식 추론 및 수학에서 LLaMa-2를 능가합니다. 이미지 관련 작업에서는 이미지 캡션 작성 및 시각적 질문 답변에 탁월하여 더 적은 수의 샷으로 더 큰 모델을 능가하거나 일치시킵니다.

카멜레온: 원활한 다중 모드 문서 생성을 위한 혁신적인 모델

Chameleon: A Unified Foundation Model for Seamless Multimodal Document Modeling

카멜레온: 원활한 다중 모드 문서 모델링을 위한 통합 기반 모델

Recent multimodal foundation models have revolutionized the field of artificial intelligence, enabling impressive advancements in tasks such as natural language processing, image recognition, and question answering. However, these models often treat different modalities separately, employing specific encoders or decoders for each. This approach inherently limits their ability to effectively fuse information across modalities and produce comprehensive multimodal documents that seamlessly integrate diverse sequences of images and text.

최근의 다중 모드 기반 모델은 인공 지능 분야에 혁명을 일으켜 자연어 처리, 이미지 인식, 질문 응답과 같은 작업에서 인상적인 발전을 가능하게 했습니다. 그러나 이러한 모델은 각각에 대해 특정 인코더 또는 디코더를 사용하여 다양한 양식을 별도로 처리하는 경우가 많습니다. 이러한 접근 방식은 여러 양식에 걸쳐 정보를 효과적으로 융합하고 다양한 이미지와 텍스트 시퀀스를 원활하게 통합하는 포괄적인 다중 모드 문서를 생성하는 능력을 본질적으로 제한합니다.

Meta researchers have addressed this challenge by introducing Chameleon, a groundbreaking mixed-modal foundation model designed to facilitate seamless reasoning and generation with interleaved textual and image sequences. Unlike traditional models, Chameleon employs a unified architecture that treats both modalities equally by tokenizing images akin to text. This approach, termed early fusion, allows for seamless flow of information across modalities but also presents unique optimization challenges.

메타 연구자들은 인터리브된 텍스트 및 이미지 시퀀스를 통해 원활한 추론 및 생성을 용이하게 하도록 설계된 획기적인 혼합 모드 기반 모델인 Chameleon을 도입하여 이러한 문제를 해결했습니다. 기존 모델과 달리 Chameleon은 텍스트와 유사한 이미지를 토큰화하여 두 가지 양식을 동일하게 처리하는 통합 아키텍처를 사용합니다. 초기 융합이라고 하는 이 접근 방식은 양식 전반에 걸쳐 정보의 원활한 흐름을 허용하지만 고유한 최적화 문제도 제시합니다.

To overcome these challenges, the researchers have incorporated architectural enhancements and innovative training techniques into Chameleon's design. By adapting transformer architecture and fine-tuning strategies, they have laid the foundation for a model capable of addressing the complexities of multimodal document modeling.

이러한 과제를 극복하기 위해 연구원들은 Chameleon의 디자인에 아키텍처 개선과 혁신적인 훈련 기술을 통합했습니다. 변환기 아키텍처와 미세 조정 전략을 적용하여 다중 모드 문서 모델링의 복잡성을 해결할 수 있는 모델의 기반을 마련했습니다.

To capture the visual information effectively, they developed a novel image tokenizer that encodes high-resolution 512 × 512 images into 1024 tokens from an 8192-codebook. This tokenization process ensures that images are represented in a manner compatible with the text tokens, facilitating seamless reasoning across modalities. They further refined their tokenizer by introducing a BPE tokenizer with a 65,536-vocabulary, including image tokens, trained using the sentencepiece library. This enhanced tokenizer improved the model's performance and stability.

시각적 정보를 효과적으로 캡처하기 위해 고해상도 512 × 512 이미지를 8192 코드북의 1024개 토큰으로 인코딩하는 새로운 이미지 토크나이저를 개발했습니다. 이 토큰화 프로세스는 이미지가 텍스트 토큰과 호환되는 방식으로 표현되도록 보장하여 양식 전반에 걸쳐 원활한 추론을 촉진합니다. 그들은 문장 라이브러리를 사용하여 훈련된 이미지 토큰을 포함하여 65,536개의 어휘가 있는 BPE 토크나이저를 도입하여 토크나이저를 더욱 개선했습니다. 이 향상된 토크나이저는 모델의 성능과 안정성을 향상시켰습니다.

During training, they employed QK-Norm, dropout, and z-loss regularization techniques to address stability issues. Additionally, they leveraged Meta's RSC platform for efficient and scalable training. Inference in Chameleon is streamlined using PyTorch and xformers, supporting both streaming and non-streaming modes with token masking for conditional logic.

훈련 중에 안정성 문제를 해결하기 위해 QK-Norm, 드롭아웃 및 z-손실 정규화 기술을 사용했습니다. 또한 효율적이고 확장 가능한 교육을 위해 Meta의 RSC 플랫폼을 활용했습니다. Chameleon의 추론은 조건부 논리를 위한 토큰 마스킹을 통해 스트리밍 및 비스트리밍 모드를 모두 지원하는 PyTorch 및 xformers를 사용하여 간소화됩니다.

To enhance model capabilities and safety, Chameleon underwent a rigorous alignment stage, undergoing fine-tuning on diverse datasets encompassing Text, Code, Visual Chat, and Safety. This comprehensive training regimen aimed to improve the model's generalizability and mitigate potential risks. They curated high-quality images for Image Generation using an aesthetic classifier, ensuring that the model generates visually appealing and relevant images.

모델 기능과 안전성을 강화하기 위해 Chameleon은 엄격한 정렬 단계를 거쳐 텍스트, 코드, 시각적 채팅 및 안전을 포괄하는 다양한 데이터 세트를 미세 조정했습니다. 이 포괄적인 훈련 방식은 모델의 일반화 가능성을 향상하고 잠재적인 위험을 완화하는 것을 목표로 했습니다. 그들은 미적 분류자를 사용하여 이미지 생성을 위한 고품질 이미지를 선별하여 모델이 시각적으로 매력적이고 관련성 있는 이미지를 생성하도록 했습니다.

Supervised Fine-Tuning (SFT) played a crucial role in Chameleon's training process. By balancing data across modalities and employing a cosine learning rate schedule with a weight decay of 0.1, the researchers optimized the model's performance exclusively based on answer prompts. Dropout of 0.05 and z-loss regularization further enhanced the model's stability. Images in prompts were resized with border padding, while those in answers underwent center-cropping for optimal image generation quality.

감독된 미세 조정(SFT)은 카멜레온의 훈련 과정에서 중요한 역할을 했습니다. 연구자들은 여러 양식에 걸쳐 데이터의 균형을 맞추고 가중치 감소가 0.1인 코사인 학습률 일정을 사용하여 답변 프롬프트를 기반으로 모델의 성능을 최적화했습니다. 0.05의 드롭아웃과 z-손실 정규화로 모델의 안정성이 더욱 향상되었습니다. 프롬프트의 이미지는 테두리 패딩을 사용하여 크기가 조정되었으며, 답변의 이미지는 최적의 이미지 생성 품질을 위해 중앙 자르기를 거쳤습니다.

Chameleon's versatility extends beyond its core multimodal capabilities, as evidenced by its strong performance on various text-only tasks. The model achieved competitive results on commonsense reasoning and math tasks, surpassing LLaMa-2 on many benchmarks. This improvement can be attributed to Chameleon's robust pre-training and the inclusion of code data in its training regimen.

다양한 텍스트 전용 작업에 대한 강력한 성능에서 알 수 있듯이 Chameleon의 다용성은 핵심 다중 모드 기능 이상으로 확장됩니다. 이 모델은 상식 추론 및 수학 작업에서 경쟁력 있는 결과를 달성했으며 많은 벤치마크에서 LLaMa-2를 능가했습니다. 이러한 개선은 Chameleon의 강력한 사전 훈련과 훈련 계획에 코드 데이터를 포함했기 때문일 수 있습니다.

In image-to-text tasks, Chameleon excelled in image captioning, matching or surpassing larger models like Flamingo-80B and IDEFICS-80B with fewer shots. Its performance in visual question answering (VQA) approached that of top models, even though LLaMa-1.5 slightly outperformed VQA-v2.

이미지-텍스트 작업에서 Chameleon은 이미지 캡션 작성에 탁월하여 Flamingo-80B 및 IDEFICS-80B와 같은 대형 모델을 더 적은 수의 샷으로 일치하거나 능가했습니다. LLaMa-1.5가 VQA-v2보다 약간 뛰어난 성능을 보였지만 시각적 질문 응답(VQA) 성능은 상위 모델의 성능에 근접했습니다.

Overall, Chameleon demonstrates remarkable versatility and efficiency across different tasks, requiring fewer training examples and smaller model sizes compared to its counterparts.

전반적으로 Chameleon은 다양한 작업 전반에 걸쳐 놀라운 다양성과 효율성을 보여 주며, 대응하는 모델에 비해 훈련 예제가 적고 모델 크기가 더 작습니다.

In summary, Chameleon is a token-based model that seamlessly integrates image and text tokens, enabling superior performance in vision-language tasks. Its unified architecture facilitates joint reasoning over modalities, surpassing late-fusion models like Flamingo and IDEFICS in image captioning and visual question answering. Chameleon's early-fusion approach introduces novel techniques for stable training, addressing scalability challenges faced by previous multimodal models. It opens up new possibilities for multimodal interaction, as demonstrated by its strong performance on mixed-modal open-ended QA benchmarks.

요약하면, 카멜레온은 이미지와 텍스트 토큰을 완벽하게 통합하여 비전 언어 작업에서 뛰어난 성능을 제공하는 토큰 기반 모델입니다. 통합 아키텍처는 양식에 대한 공동 추론을 용이하게 하며 이미지 캡션 및 시각적 질문 답변 분야에서 Flamingo 및 IDEFICS와 같은 후기 융합 모델을 능가합니다. Chameleon의 초기 융합 접근 방식은 안정적인 훈련을 위한 새로운 기술을 도입하여 이전 다중 모드 모델이 직면한 확장성 문제를 해결합니다. 혼합 모드 개방형 QA 벤치마크에서 강력한 성능을 보여주듯이 다중 모드 상호 작용에 대한 새로운 가능성을 열어줍니다.

This groundbreaking research paves the way for the development of even more advanced multimodal models capable of handling complex tasks and producing highly coherent and informative content. Researchers, practitioners, and industry leaders alike eagerly anticipate the transformative potential of Chameleon and its potential impact on the future of AI-driven multimodal applications.

이 획기적인 연구는 복잡한 작업을 처리하고 일관되고 유익한 콘텐츠를 생성할 수 있는 더욱 발전된 다중 모드 모델을 개발할 수 있는 길을 열어줍니다. 연구원, 실무자 및 업계 리더 모두 카멜레온의 혁신적인 잠재력과 AI 기반 다중 모드 애플리케이션의 미래에 대한 잠재적 영향을 간절히 기대하고 있습니다.

부인 성명:info@kdj.com

제공된 정보는 거래 조언이 아닙니다. kdj.com은 이 기사에 제공된 정보를 기반으로 이루어진 투자에 대해 어떠한 책임도 지지 않습니다. 암호화폐는 변동성이 매우 높으므로 철저한 조사 후 신중하게 투자하는 것이 좋습니다!

본 웹사이트에 사용된 내용이 귀하의 저작권을 침해한다고 판단되는 경우, 즉시 당사(info@kdj.com)로 연락주시면 즉시 삭제하도록 하겠습니다.

2026年08月02日 에 게재된 다른 기사