|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Meta AI 和阿姆斯特丹大學的最新研究表明,變壓器(一種流行的神經網路架構)可以直接對影像的各個像素進行操作,而不依賴大多數現代電腦視覺模型中存在的局部歸納偏差。
![]()
Meta AI and researchers from the University of Amsterdam have demonstrated that transformers, a popular neural network architecture, can operate directly on individual pixels of an image, without relying on the locality inductive bias present in most modern computer vision models.
Meta AI 和阿姆斯特丹大學的研究人員證明,變壓器(一種流行的神經網路架構)可以直接對影像的各個像素進行操作,而不依賴大多數現代電腦視覺模型中存在的局部歸納偏差。
Their study, titled "Transformers on Individual Pixels," challenges the long-held belief that locality – the notion that neighboring pixels are more related than distant ones – is a fundamental requirement for vision tasks.
他們的研究題為“單一像素上的變形金剛”,挑戰了人們長期以來的信念,即局部性(即相鄰像素比遠處像素更相關)是視覺任務的基本要求。
Traditionally, computer vision architectures like Convolutional Neural Networks (ConvNets) and Vision Transformers (ViTs) have incorporated locality bias through techniques such as convolutional kernels, pooling operations, and patchification, assuming neighboring pixels are more related.
傳統上,卷積神經網路 (ConvNets) 和視覺變換器 (ViTs) 等電腦視覺架構透過卷積核、池化操作和修補程式化等技術納入局部偏差,假設相鄰像素更相關。
In contrast, the researchers introduced Pixel Transformers (PiTs), which treat each pixel as an individual token, removing any assumptions about the 2D grid structure of images. Surprisingly, PiTs achieved highly performant results across various tasks.
相較之下,研究人員引入了像素變換器 (PiT),它將每個像素視為單獨的標記,消除了有關影像 2D 網格結構的任何假設。令人驚訝的是,PiT 在各種任務中都取得了高性能的結果。
For instance, when PiTs were applied to image generation tasks using latent token spaces from VQGAN, they outperformed their locality-biased counterparts on quality metrics like Fréchet Inception Distance (FID) and Inception Score (IS).
例如,當 PiT 使用 VQGAN 的潛在標記空間應用於影像生成任務時,它們在 Fréchet 起始距離 (FID) 和起始分數 (IS) 等品質指標上優於局部偏向的同行。
While PiTs, operating on the lines of Perceiver IO Transformers, can be computationally expensive due to longer sequences, they challenge the need for locality bias in vision models. As advances in handling large sequence lengths are made, PiTs may become more practical.
雖然 PiT 在 Perceiver IO Transformer 上運行,由於序列較長,計算成本可能很高,但它們挑戰了視覺模型中對局部性偏差的需求。隨著處理大序列長度的進步,PiT 可能變得更加實用。
The study ultimately highlights the potential benefits of reducing inductive biases in neural architectures, which could lead to more versatile and capable systems for diverse vision tasks and data modalities.
該研究最終強調了減少神經架構中歸納偏差的潛在好處,這可能會導致針對不同視覺任務和資料模式的更通用和強大的系統。
免責聲明:info@kdj.com
所提供的資訊並非交易建議。 kDJ.com對任何基於本文提供的資訊進行的投資不承擔任何責任。加密貨幣波動性較大,建議您充分研究後謹慎投資!
如果您認為本網站使用的內容侵犯了您的版權,請立即聯絡我們(info@kdj.com),我們將及時刪除。
-
- 比特幣、eCash 分叉和空投動態:深入探討加密貨幣的最新爭議
- 2026-05-03 00:52:02
- 探索最近的 eCash 分叉、其作為高風險空投的分類,以及對比特幣和加密生態系統的更廣泛影響。
-
-
- 聯準會維持利率穩定,地緣政治緊張局勢引發比特幣價格下跌
- 2026-05-01 04:04:38
- 聯準會維持利率的決定,加上中東衝突,影響了比特幣的價格。分析近期趨勢和市場反應。
-
-
-
-
-
-

































