|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Meta AI와 암스테르담 대학교의 최근 연구에 따르면 인기 있는 신경망 아키텍처인 변환기는 대부분의 최신 컴퓨터 비전 모델에 존재하는 국소성 유도 편향에 의존하지 않고 이미지의 개별 픽셀에서 직접 작동할 수 있는 것으로 나타났습니다.
![]()
Meta AI and researchers from the University of Amsterdam have demonstrated that transformers, a popular neural network architecture, can operate directly on individual pixels of an image, without relying on the locality inductive bias present in most modern computer vision models.
Meta AI와 암스테르담 대학교 연구원들은 널리 사용되는 신경망 아키텍처인 변환기가 대부분의 현대 컴퓨터 비전 모델에 존재하는 국소성 유도 편향에 의존하지 않고 이미지의 개별 픽셀에서 직접 작동할 수 있음을 입증했습니다.
Their study, titled "Transformers on Individual Pixels," challenges the long-held belief that locality – the notion that neighboring pixels are more related than distant ones – is a fundamental requirement for vision tasks.
"개별 픽셀의 변환기(Transformers on Individual Pixels)"라는 제목의 그들의 연구는 인접 픽셀이 먼 픽셀보다 더 관련되어 있다는 개념인 지역성이 비전 작업의 기본 요구 사항이라는 오랜 믿음에 도전합니다.
Traditionally, computer vision architectures like Convolutional Neural Networks (ConvNets) and Vision Transformers (ViTs) have incorporated locality bias through techniques such as convolutional kernels, pooling operations, and patchification, assuming neighboring pixels are more related.
전통적으로 ConvNets(Convolutional Neural Networks) 및 ViTs(Vision Transformers)와 같은 컴퓨터 비전 아키텍처는 인접 픽셀이 더 관련되어 있다고 가정하여 Convolutional 커널, 풀링 작업 및 패치화와 같은 기술을 통해 지역성 편향을 통합했습니다.
In contrast, the researchers introduced Pixel Transformers (PiTs), which treat each pixel as an individual token, removing any assumptions about the 2D grid structure of images. Surprisingly, PiTs achieved highly performant results across various tasks.
이와 대조적으로 연구원들은 각 픽셀을 개별 토큰으로 처리하여 이미지의 2D 그리드 구조에 대한 모든 가정을 제거하는 픽셀 변환기(PiT)를 도입했습니다. 놀랍게도 PiT는 다양한 작업에서 매우 뛰어난 결과를 달성했습니다.
For instance, when PiTs were applied to image generation tasks using latent token spaces from VQGAN, they outperformed their locality-biased counterparts on quality metrics like Fréchet Inception Distance (FID) and Inception Score (IS).
예를 들어 VQGAN의 잠재 토큰 공간을 사용하여 이미지 생성 작업에 PiT를 적용했을 때 FID(Fréchet Inception Distance) 및 IS(Inception Score)와 같은 품질 지표에서 지역성 편향된 대응 항목보다 성능이 뛰어났습니다.
While PiTs, operating on the lines of Perceiver IO Transformers, can be computationally expensive due to longer sequences, they challenge the need for locality bias in vision models. As advances in handling large sequence lengths are made, PiTs may become more practical.
Perceiver IO Transformers 라인에서 작동하는 PiT는 더 긴 시퀀스로 인해 계산 비용이 많이 들 수 있지만 비전 모델의 지역성 편향에 대한 필요성에 도전합니다. 긴 시퀀스 길이를 처리하는 기술이 발전함에 따라 PiT가 더욱 실용적이 될 수 있습니다.
The study ultimately highlights the potential benefits of reducing inductive biases in neural architectures, which could lead to more versatile and capable systems for diverse vision tasks and data modalities.
이 연구는 궁극적으로 신경 구조의 유도 편향을 줄이는 것의 잠재적인 이점을 강조하며, 이를 통해 다양한 비전 작업 및 데이터 양식을 위한 보다 다재다능하고 유능한 시스템을 만들 수 있습니다.
부인 성명:info@kdj.com
제공된 정보는 거래 조언이 아닙니다. kdj.com은 이 기사에 제공된 정보를 기반으로 이루어진 투자에 대해 어떠한 책임도 지지 않습니다. 암호화폐는 변동성이 매우 높으므로 철저한 조사 후 신중하게 투자하는 것이 좋습니다!
본 웹사이트에 사용된 내용이 귀하의 저작권을 침해한다고 판단되는 경우, 즉시 당사(info@kdj.com)로 연락주시면 즉시 삭제하도록 하겠습니다.

































