|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Meta AI 和阿姆斯特丹大学的最新研究表明,变压器(一种流行的神经网络架构)可以直接对图像的各个像素进行操作,而不依赖于大多数现代计算机视觉模型中存在的局部归纳偏差。
![]()
Meta AI and researchers from the University of Amsterdam have demonstrated that transformers, a popular neural network architecture, can operate directly on individual pixels of an image, without relying on the locality inductive bias present in most modern computer vision models.
Meta AI 和阿姆斯特丹大学的研究人员证明,变压器(一种流行的神经网络架构)可以直接对图像的各个像素进行操作,而不依赖于大多数现代计算机视觉模型中存在的局部归纳偏差。
Their study, titled "Transformers on Individual Pixels," challenges the long-held belief that locality – the notion that neighboring pixels are more related than distant ones – is a fundamental requirement for vision tasks.
他们的研究题为“单个像素上的变形金刚”,挑战了人们长期以来的信念,即局部性(即相邻像素比远处像素更相关)是视觉任务的基本要求。
Traditionally, computer vision architectures like Convolutional Neural Networks (ConvNets) and Vision Transformers (ViTs) have incorporated locality bias through techniques such as convolutional kernels, pooling operations, and patchification, assuming neighboring pixels are more related.
传统上,卷积神经网络 (ConvNets) 和视觉变换器 (ViTs) 等计算机视觉架构通过卷积核、池化操作和补丁化等技术纳入局部偏差,假设相邻像素更相关。
In contrast, the researchers introduced Pixel Transformers (PiTs), which treat each pixel as an individual token, removing any assumptions about the 2D grid structure of images. Surprisingly, PiTs achieved highly performant results across various tasks.
相比之下,研究人员引入了像素变换器 (PiT),它将每个像素视为一个单独的标记,消除了有关图像 2D 网格结构的任何假设。令人惊讶的是,PiT 在各种任务中都取得了高性能的结果。
For instance, when PiTs were applied to image generation tasks using latent token spaces from VQGAN, they outperformed their locality-biased counterparts on quality metrics like Fréchet Inception Distance (FID) and Inception Score (IS).
例如,当 PiT 使用 VQGAN 的潜在标记空间应用于图像生成任务时,它们在 Fréchet 起始距离 (FID) 和起始分数 (IS) 等质量指标上优于局部偏向的同行。
While PiTs, operating on the lines of Perceiver IO Transformers, can be computationally expensive due to longer sequences, they challenge the need for locality bias in vision models. As advances in handling large sequence lengths are made, PiTs may become more practical.
虽然 PiT 在 Perceiver IO Transformer 上运行,由于序列较长,计算成本可能会很高,但它们挑战了视觉模型中对局部性偏差的需求。随着处理大序列长度的进步,PiT 可能变得更加实用。
The study ultimately highlights the potential benefits of reducing inductive biases in neural architectures, which could lead to more versatile and capable systems for diverse vision tasks and data modalities.
该研究最终强调了减少神经架构中归纳偏差的潜在好处,这可能会导致针对不同视觉任务和数据模式的更加通用和强大的系统。
免责声明:info@kdj.com
所提供的信息并非交易建议。根据本文提供的信息进行的任何投资,kdj.com不承担任何责任。加密货币具有高波动性,强烈建议您深入研究后,谨慎投资!
如您认为本网站上使用的内容侵犯了您的版权,请立即联系我们(info@kdj.com),我们将及时删除。
-
- 比特币、eCash 分叉和空投动态:深入探讨加密货币的最新争议
- 2026-05-03 00:52:02
- 探索最近的 eCash 分叉、其作为高风险空投的分类,以及对比特币和加密生态系统的更广泛影响。
-
-
- 美联储维持利率稳定,地缘政治紧张局势引发比特币价格下跌
- 2026-05-01 04:04:38
- 美联储维持利率的决定,加上中东冲突,影响了比特币的价格。分析近期趋势和市场反应。
-
-
-
-
-
-

































