版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
FLattenTransformer:VisionTransformerusingFocusedLinearAttention
DongchenHan*XuranPan*YizengHan
ShijiSongGaoHuang
V
NXd
V
NXd
SoftmaxAttention
SoftmaxAttentiono=softmax(QKT)V
GiventheinputNtokensXeRN根c,withineachhead,self-attentioncanbewrittenas:
Q=xWQ,K=xWK,V=xWV,
oi=
j=1
Sim(Qi,kj)
Σ1Sim(Qi,kj)vj,
whereWQ,WK,WVeRc根careprojectionmatricesandSim(Q,K)=exp(QKT/d).
NXd
KT
N
d
Q
X
QKT
NXN
Highexpressivecapability
收Quadraticcomplexity0(N2d)
收Usuallyusingsparseattentionorwindowattentionpatternstoreducecomplexity
2
SoftmaxAttention
SparseGlobalAttentionPVT/PVTv2
1.WenhaiWang,EnzeXie,XiangLi,Deng-PingFan,KaitaoSong,DingLiang,TongLu,PingLuo,andLingShao.Pyramidvisiontransformer:Aversatilebackbonefordensepredictionwithoutconvolutions.InProceedingsoftheIEEE/CVFInternationalConferenceonComputerVision,pages568–578,2021.
3
2.WenhaiWang,EnzeXie,XiangLi,Deng-PingFan,KaitaoSong,DingLiang,TongLu,PingLuo,andLingShao.Pvtv2:Improvedbaselineswithpyramidvisiontransformer.ComputationalVisualMedia,8(3):415–424,2022.
SoftmaxAttention
ShiftedWindowAttentionSwinTransformer
1.ZeLiu,YutongLin,YueCao,HanHu,YixuanWei,ZhengZhang,StephenLin,andBainingGuo.Swintransformer:Hierarchicalvisiontransformerusingshiftedwindows.InProceedingsoftheIEEE/CVFInternationalConferenceonComputerVision,pages10012–10022,2021.
4
SoftmaxAttention
Cross-ShapedWindowAttentionCSwinTransformer
1.XiaoyiDong,JianminBao,DongdongChen,WeimingZhang,NenghaiYu,LuYuan,DongChen,andBainingGuo.Cswintransformer:Ageneralvisiontransformerbackbonewithcross-shapedwindows.InProceedingsoftheIEEE/CVFConferenceonComputerVisionandPatternRecognition,pages12124–12134,2022.
5
V
Q
d
N
T
K
×
N×d
N×d
Q
N×d
LinearAttention
LinearAttentiono=Φ(Q)(Φ(K)TV)
Carefullydesignedkernelsareintroducedastheapproximationoftheoriginalsimilarityfunction:
Sim(Q,K)=Φ(Q)Φ(K)T,
Nφ(Qi)φ(kj)T
oi=Σj=1Σ1φ(Qi)φ(kj)Tvj
Σ1(φ(Qi)φ(kj)T)vj
=
φ(1φ(kj)T
φ(
=
φ(1φ(kj)T).
KTV
d×d
收Inferiorperformance
LinearcomplexityΘ(Nd2)
Canenjoyalargereceptivefieldwhile
maintainingalowamountofcomputation
6
1.KatharopoulosA,VyasA,PappasN,etal.Transformersarernns:Fastautoregressivetransformerswithlinearattention[C]//InternationalConferenceonMachineLearning.PMLR,2020:5156-5165.
LinearAttention
HydraAttention0(Nd)
1.DanielBolya,Cheng-YangFu,XiaoliangDai,PeizhaoZhang,andJudyHofman.Hydraattention:Eficientattentionwithmanyheads.InComputerVision–ECCV2022Workshops:TelAviv,Israel,October23–27,2022,Proceedings,PartVII,pages35–49.Springer,2023.
7
LinearAttention
EfficientAttention0(Nd2)
1.ZhuoranShen,MingyuanZhang,HaiyuZhao,ShuaiYi,andHongshengLi.Eficientattention:Attentionwithlinearcomplexities.InProceedingsoftheIEEE/CVFWinterConferenceonApplicationsofComputerVision,pages3531–3539,2021.
8
IMotivation
Todesignlinearattentionmodulethatperforms
onparwithorevenbetterthanSoftmaxattention.
9
Linear
Attention
FocusAbility
Original
Image
Softmax
Attention
Sharp
Focusoncertainregions
收Smooth
收Closertotheaverageofalltokens
FocusedLinear
Attention
(Ours)
10
FocusAbility
FocusedFunctionfp
Sim(Qi,kj)=φp(Qi)φp(kj)T,
whereφp(x)=fp(ReLU(Q)),afp(x)=⋅x∗∗p.
FeatureDirectionAdjustmentwithfp
Letx=(x1,x2,x3,⋯,xn)∈ℝn,ay=(y1,y2,y3,⋯,yn)∈ℝn
Assume:(1)∃!m,s.t.xm=1xi),∃!n,s.t.yn=1yj),
(2)0<x,y<IxIIyI.
Forapairoffeatures{x,y}withm=n:∃p>1,s.t.φp(x),φp(y)>x,y.
Forapairoffeatures{x,y}withm≠n:∃p>1,s.t.φp(x),φp(y)<x,y.
11
FocusAbility
Proof:
φp(x)=fp(ReLU(x))=fp(x),aφp(y)=fp(ReLU(y))=fp(y),
fp(x)=x∗∗p=x,afp(y)=y.
Therefore,wehave:
〈φp(x),φp(y)〉=〈fp(x),fp(y)〉=fp(x)fp(y)〈u,V〉
=xy〈u,V〉,
where
〈u,V〉=〈,〉
12
1
xp
FocusAbility
Proof:
u,V=II,II=
=
Σ1xy
Σ1ab
1ap1bp),
ai=,bi=,ai,bi∈[0,1].
Basedonourassumption,wehave:
∃!m,s.t.am=1,∃!n,s.t.bn=1.
Therefore,
a={,b={.
13
FocusAbility
Considerthefollowingtwocases:
(1)m=n:
〈u,v〉=
==1.
1根1
〈φp(x),φp(y)〉=xy〈u,v〉=xy>〈x,y〉.
Thuswehave,
∃p>1,s.t.〈Φp(x),Φp(y)〉>〈x,y〉.
1根1
(2)m≠n:
〈u,v〉=
==0.
1根1
〈φp(x),φp(y)〉=xy〈u,v〉=0<〈x,y〉.
Thuswehave,
∃p>1,s.t.〈Φp(x),Φp(y)〉<〈x,y〉.
1根0+0根1
14
FocusAbility
Exampletoshowtheefectsoffp:
k2
k4
f3(k1)f3(k2)
f3(k3)f3(k4)
k1
k3
(b)AttentionMap
(a)QueryandKeys
15
FeatureDiversity
(a)SoftmaxAttentionMapMatrixRank:196
Softmaxattentioncanlearn
afull-rankattentionmap
(b)LinearAttentionMapMatrixRank:54
收Linearattentioncannotlearnanattentionmapwitharankgreaterthanheaddim64
LinearAttention:
O=φ(Q)φ(K)TV
Rank(φ(Q)φ(K)T)
≤min{Rank(φ(Q)),Rank(φ(K)T)}
≤iͺ{N,d}
16
RankRestorationModel(DWC)
O=φ(Q)φ(K)TV+DWC(V),
O=(φ(Q)φ(K)T+MDWC)V=MeqV.
FocusedLinearAttention(FLatten)
O=Sim(Q,K)V
=Φ(Q)Φ(K)TV+퐃W͗(V),
whereφp(x)=fp(ReLU(Q)),fp(x)=⋅x∗∗p.
FeatureDiversity
(c)Ours
MatrixRank:196
17
IFocusedLinearAttention
O=Sim(Q,K)V=Φ(Q)Φ(K)TV+퐃W͗(V)
FLatten
Transformer
{
Lowcomputationcomplexityaslinearattention.
HighexpressivecapabilityasSoftmaxattention.
Largereceptivefield.
Flexibility.
18
ExperimentResults
PerformancesonImageNetClassification:
19
ExperimentResults
PerformancesonCOCOobjectdetection:
20
ExperimentResults
Comparisonwithotherlinearattentiondesigns:
21
IExperimentResults
Accuracy-RuntimecurveonImageNet:
22
ExperimentResults
Ablationoneachmodule
basedonDeiT-T.
Ablationonwindowsize
based
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2026年全国交管12123驾驶证学法减分(学法免分)考试题附含参考答案
- 7.6 CNN 过拟合欠拟合
- 脊髓损伤的并发症护理
- 2026年电力通信网络规划员考试真题解析卷
- 2025年广东省吴川市高二生物下册期末考试模拟试卷及参考答案1套
- 平板显示膜涂布工安全知识宣贯考核试卷含答案
- 衡器装配调试工道德考核试卷含答案
- 酒精发酵工安全管理强化考核试卷含答案
- 压缩机操作工安全实践竞赛考核试卷含答案
- 籽晶片制造工岗前操作考核试卷含答案
- 海水集中取水项目施工方案
- 2026年秋季开学情绪管理心理危机干预培训课件
- 小学主题班会课件:小小少年扣好人生第一粒扣子
- 2026年江苏省苏州市中考语文试题(原卷版)
- 2026-2030中国铬矿行业市场发展趋势与前景展望战略分析研究报告
- 《锂离子电池化成分容系统》-全文及说明
- 美国律师职业责任制度
- 2025华能澜沧江水电股份有限公司大学毕业生招聘17人笔试参考题库附带答案详解(3卷)
- 感染性疾病科医生进修汇报
- 泥浆清运合同(2025版)
- 睿达杯二试试题和答案
评论
0/150
提交评论