版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
arXiv:2304.06712v1[cs.CV]13Apr2023
WhatdoesCLIPknowaboutaredcircle?
VisualprmptengineeringforVLMs
AleksandarShtedritskiChristianRupprechtAndreaVedaldi
VisualGeometryGroup,UniversityofOxford
(suny,chrisr,vedaldi}@robots.ox.ac.uk
Abstract
Large-scaleVision-LanguageModels,suchasCLIP,learnpowerfulimage-textrepresentationsthathavefoundnumerousapplications,fromzero-shotclassificationtotext-to-imagegeneration.Despitethat,theircapabilitiesforsolvingnoveldiscriminativetasksviapromptingfallbehindthoseoflargelanguagemodels,suchasGPT-3.Hereweex-ploretheideaofvisualpromptengineeringforsolvingcom-putervisiontasksbeyondclassificationbyeditinginimagespaceinsteadoftext.Inparticular,wediscoveranemer-gentabilityofCLIP,where,bysimplydrawingaredcir-clearoundanobject,wecandirectthemodel’sattentiontothatregion,whilealsomaintainingglobalinformation.Weshowthepowerofthissimpleapproachbyachievingstate-of-the-artinzero-shotreferringexpressionscomprehensionandstrongperformanceinkeypointlocalizationtasks.Fi-nally,wedrawattentiontosomepotentialethicalconcernsoflargelanguage-visionmodels.
1.Introduction
LargeLanguageModels(LLMs)suchasGPT-2/3[
7
,
40
]andChatGPT[
1
]havedemonstratedsurprisingemergingbehaviours.Forexample,thesemodelscanperformlan-guagetranslationwithoutbeingexplicitlytrainedforit,inazero-shotmanner.Thiscanbepartiallyexplainedbythefactthatoccurrencesofthedesiredbehaviours,suchastranslatingbetweentwolanguages,naturallyoccurintheirenormoustrainingcorpus,whichis,essentially,theInternet.
InterestingemergentbehaviourshavebeenobservedinlargeVision-LanguageModels(VLMs)likeCLIP[
39
]too.Forexample,CLIPcanbeusedforzero-shotclassifica-tionbycheckingthecompatibilityofagivenimagewithpromptssuchas“animageofaX”,whereXisoneofasetofclasshypothesestobetested.
EmergentbehavioursareelicitedbysupplyingsuitablycraftedinputstotheVLMs,oftencalledprompts.Asintheexampleabove,researchershavemostlyfocusedonengi-
ThecowthatisthesmallestThewhitelittlelamb
forehead
righteye
mouth
nose
Figure1:VisualPromptEngineering.WedrawmultipleannotationsoveranimageandhaveCLIPchoosethecorrectonegivenacaption.Hereweshowpredictionsforthegivenexpressions.Top:ExamplesfromRefCOCOgonreferring
expressionsdetection.Bottom:ExamplefromSPair71konkeypointlocalization.
neeringtextualprompts,manipulatingthetextualinputofthemodel.ThisapproachisinspiredbyLLMs,wherema-nipulatingthetextualmodalityistheonlyavailableoption.However,VLMsareinherentlymultimodalandofferthepossibilityofmanipulatingbothmodalities,textualandvi-sual.Whilethetextualmodalityisthenaturalchoiceforexpressingsemantics,thevisualmodalitycanbebetterforexpressinggeometricpropertiessuchaslocation.
Inthispaper,wethusexplorevisualpromptengineer-ing
1
.Wedosowithtwogoals.Thefirstgoalistocon-tributeonemorepracticaltoolforextractingusefulinfor-mationfromVLMsinazero-shotmanner.Wedemon-stratethisbyobtainingstate-of-the-artzero-shotresultsinreferringexpressionscomprehensionbyengineeringvisualprompts.Thesecondgoalistocharacteriseinterestingand
1Notethedifferencebetweenvisualprompttuning,asettingpreviouslyexplored,wherethepromptsaretask-specificlearnabletokens,andvisualpromptengineering,whereweapplyafixedaugmentationinpixelspace.
unexpectedpropertiesoftheVLMsandtheirtrainingdata,includingidentifyingsomebehavioursthatcanraiseethicalconcerns.
Perhapsthemostsurprisingofourfindingsistheeffec-tivenessofaparticulartypeofvisualprompting:drawingaplainredcircleontopoftheimage(Fig.
1
).WeshowthatthissimpleinterventionsteerstheVLMtoanalyse/talkabouttheimageregioncontainedinthecircle.Thisbe-haviourcanthenbeusedfortaskssuchasnamingaspe-cificobjectorobjectpartordetectingparticularimagere-gionsbasedonadescription.Thelatter,forinstance,isachievedbymarkingeachobjectproposalwitharedcircleandusingtheVLMtofindthebestmatchwithrespecttotheprovidedreferringexpression,achievingstrongresultsonmultiplebenchmarksintheunsupervisedregime.Fur-thermore,weshowthatpromptingwithacirclealsoworksforfiner-grainedlocalization,markingspecificobjectpartsorkeypointsinsteadofjustwholeobjects.
Wefurthercontrastmarkinganimagewiththealterna-tiveofcroppingit,which,fromslidingwindowclassifierstoregionneuralnetworks,isthecanonicalapproachtosteerthefocusofanimage-levelpredictortoaparticularimageregion.Weshowthat,forVLMsatleast,markingissig-nificantlymoreeffectivethancropping,possiblybecauseitdoesnotlosecontextualinformationlikethelatter.
Apartfromthepracticalapplications,ourfindingsrevealunexpectedandintriguingpropertiesofVLMs.Weshowempirically,thatmarkingwitharedcircleisoptimalamongaselectionofpossiblemarkers(variantsofthecircle,boxes,arrows,etc.).Presumably,theVLMsunderstandredcirclesoutoftheboxbecausetheseappearsufficientlyfrequentlyinthetrainingcorpus,i.e.,theInternet.WhilewedonothaveaccesstothefulltrainingdataofCLIP,wecorrobo-ratethisintuitionbyseekingexamplesofsuchimagesinYFCC15M,adatasetofCC-BYimages.
Ouranalysisshowsthatredcirclesareindeedpresentevenina(comparativelysmall)datasetofimageslikeYFCC15M,buttheyarerare.Itisatestamenttotheex-traordinarycapacityofVLMsthatsuchabehaviourcanbelearnedfromsuchrareevents,withoutanexplicitfocusondoingso.Wetestmodelsofdifferentsizes/capacitiesandshowthatonlythelargermodelsexhibitthisbehaviourre-liably,whichwellcorrespondstoourintuition.
Finally,wenotethattheabilityofVLMstolearnevenfromrareeventssuch“redcircles”canacquirebothdesir-ableandundesirablebehaviours.Redcircles,inparticular,canhaveanegativeconnotationinthetrainingdataastheyareoftenusedbynewsoutletstomarkmissingpeopleorcriminalsand,evidently,themodellearnsfromsuchexam-ples.Asaresult,weshowthatdrawingaredcircleinanimageincreasestheprobabilitythatthemodelwouldchar-acteriseapersonasacriminalorasamissingperson.
Tosummarise,wemakethefollowingmaincontribu-
tions:(1)Weproposemarkingasanewformofvisualpromptengineeringwhichiseffectiveinextractinguse-fulemergentbehavioursinVLMslikeCLIP;(2)Weusethelattertoachievestate-of-the-artzero-shotreferringex-pressionscomprehensionusingaVLM;(3)Weprovideananalysisofwhymarkingiseffectiveforthesemodels,andlinkthattothetrainingdataandlargemodelcapacity;(4)Weshowthatvisualpromptengineeringcanalsoelicitun-wantedbehaviours,suchastriggeringproblematicbiasesintheVLMs,revealingpotentialethicalissues.
2.Relatedwork
EmergentBehaviourfromLargeScalePretraining
hasmainlybeenobservedinLargeLanguageModels(LLMs).Mostnotably,GPT-2[
40
],GPT-3[
7
],andChat-GPT[
1
]havebeenshowntobecapableoftaskssuchaszero-shottranslation,questionanswering,arithmetic,aswellasplanningactionsforembodiedagents[
18
].Fine-tuningLLMscanalsoleadtomodelsthatcangeneratecodefromdocstrings[
8
]orsolvemathproblems[
13
,
24
].Onlyafewemergentzero-shotbehaviourshavebeenre-portedforVLMslikeCLIP,mainlyforclassification[
39
]andOCR[
34
].GenerativeVLMslikeFLAMINGO[
3
]andBLIP[
25
]excelincaptioningandvisualquestion-answeringtasks,butalsohavenowayofsolvingpixel-levelcomputervisiontasks.
PromptingVLMsismostcommonlyperformedbyprependingasetoflearnabletokenstothetextinput[
16
,
21
,
63
,
64
],visioninput[
19
,
47
,
62
],orbothtextandvi-sioninputs[
42
,
60
],inordertoeasilysteerafrozenCLIPmodeltosolveadesiredtask.[
4
]learnaugmentationsinpixelspace,suchaspaddingaroundtheimage,orchang-ingapatchoftheimage,whichareoptimizedwithgradientdescentonadownstreamtask.[
5
]castimageinpaintingasavisualpromptingtask,usingagenerativemodeltrainedonfiguresfromacademicpapers.ColorfulPromptTuning(CPT)[
57
]colorregionsofanimageanduseacaptioningmodeltopredictwhichobjectinanimageanexpressionreferstobypredictingitscolor.SimilarlytoCPT,weaug-menttheinputimageinpixelspaceandperformzero-shotinference.However,weannotatetheimageinahuman-likemannerandshowthatourmethodismorepowerful
andmoreflexiblethanCPT.
ReferringExpressionComprehension(REC)aimstolocalizeatargetobjectinanimagethatcorrespondstoatex-tualdescription.MostapproachestoRECstartwithobjectproposals,forexample,generatedwithFaster-RCNN[
41
],andlearntoscorethem[
17
,
31
,
32
,
49
,
54
].RECissome-timesconsideredtogetherwithreferringexpressiongener-ation—thetaskofgeneratingadescriptionofagivenre-gion.[
32
]useacomprehensionmodeltoguideagenerator,whereas[
9
]jointlytrainadetectorwithacaptiongenera-tor.Someworksmodelthesceneasagraph[
31
,
49
,
54
]or
uselanguageparsersandgrammar-basedmethods[
12
,
30
],leadingtoamoreinterpretableresult.Morerecently,trans-formerarchitectureshavebeenused[
14
,
22
,
26
,
51
,
55
].[
14
,
22
,
26
,
55
]performtext-modulatedobjectdetection,whereatransformerdecodertakesthereferringexpressionasaninputandpredictsaboundingbox.[
51
]trainwithatext-to-pixelcontrastiveloss,whichallowsforatext-drivensegmentationordetectionattesttime.
UnsupervisedReferringExpressionComprehension
isalessexploredarea,onlymadepossiblewiththein-troductionoflargepretrainedmodelssuchasCLIP[
39
].ReCLIP[
45
]cropsobjectproposalsandranksthemus-ingCLIPbeforeanad-hocpostprocessingsteptotakeintoaccountrelationssuchasleft/right,smaller/bigger,etc.CPT[
57
]colorsobjectproposalboxesanduseapre-trainedcaptioningmodel[
61
]toauto-regressivelypredictwhichcoloredproposalcorrespondstothequerydescrip-tion.Pseudo-Q[
20
]generatesdescriptionsformultipleob-jectsinanimage,whichisusedtotrainaRECnetwork.
However,thismodelisnotfullyunsupervisedasthepseudodescriptionsitusesaregeneratedusingacaptioningmodel
trainedonCOCO.
VisualReasoningUsingLargePretrainedModelshasbeenanareaofsignificantinterestinthelastfewyears.Inadditiontoreferringexpressiondetection[
45
],CLIP[
39
]hasbeenusedforsemanticsegmentation[
28
,
38
].[
38
]useCLIPtoassigntextlabelstoobjectpartsafterdoingpartco-segmentationinthelatentspaceofaGAN.[
28
]utilizeCLIPforopen-vocabularysegmentationbyusingageneral-purposemaskproposalnetworkandCLIPasaclassifier.CLIPhasalsobeenusedforunsupervisedobjectproposalgeneration[
43
]andopen-setdetection[
15
].Semanticseg-mentationalsoemergesfromimageonly[
35
,
50
]orimage-text[
53
]self-supervision.
BiasofVLMsisanincreasinglypopularareaofre-search,asdownstreamapplicationscomewiththeriskofperpetuatingbiasesandstereotypesexistinginthetrainingdata.However,methodsforassessingthebiasofaVLMarestillnotwellestablished.[
2
]measurethemisclassificationrateofCLIPoffacesofpeopleofdifferentraceswithnon-humanandcriminalcategories,whereas[
6
,
11
,
48
]measurefairnessinretrievalresults.Here,weshowadifferentkindofbias,wheretheadditionofaredcircleoverapersoncantriggeranegativeconnotation.
3.Method
OurgoalistodevelopvisualpromptinginVision-
LanguageModels(VLMs).VLMssolvepredictiontasksthatinvolvejointlyprocessingtextandimages.Forex-ample,modelssuchasCLIParetrainedtomatchtextandimagesamples.TheinputtosuchaVLMisanimageieR3×H×WandtextteΣ*,whereΣisanalphabet.Theoutputisascores(i,t)thatexpressesthedegreeof
compatibilitybetweenthesuppliedimageandtext.
3.1.Promptengineering
OneofthemoststrikingcapabilitiesofVLMsistheirabilitytosolveavarietyofclassificationtaskswithlittletonofurthertrainingatall,inazero-shotmanner.ThisisdonebyreducingthetaskofinteresttothatofevaluatingtheVLMonsuitably-engineeredimageandtextpairs.
Forexample,givenanimage-captionpair(i,t),considertheproblemoflocalizinganamedobjectkeypointintheim-age.Wecancastthisasaquestion-answerproblem,wherethequestionqeQisthenameoftheobjectkeypoint(e.g.,“rightear”,“frontleftleg”,...)andtheansweraeAisoneofadiscretesetofimagelocations.
BecausetheVLMcomputesacompatibilityscores(i,t)betweenanimageiandthetextt,itcannotbeusedtomapthequestionqtotheansweradirectly.However,viapromptengineering,wecanusetheVLMtoconstructacompati-bilityscores(q,a|i,t)betweenquestionandanswer,con-ditionedontheinputimage-textpair(i,t).Thisscoreisingeneralgivenbytheexpression
s(q,a|i,t)=s(iqa,tqa)(1)
whereiqaandtqaareversionsoftheinputimageandtext,obtainedbytransformingthelattertoreflectthequestion-answerpair(q,a).
ThespecificwayEq.(
1
)shouldbeappliedtoaprob-lemdependsonthespecificnatureofthelatter.Forex-ample,intheproblemoflocalizingthenamedkeypoints,itisnaturaltoencodethenameofthekeypointviathetextualmodalityandits2Dlocationviathevisualmodal-ity.Forinstance,inordertoanswerthequestionq=“rightear”foragiveninputimageiwithcaptiont=“dog”,wecanengineerthetextualprompttqa=tq=“animageoftherightearofadog”toencodeadescriptionofthenamedentity.Likewise,wecanengineerthevisualpromptiqa=iainsuchawayasto‘select’thelocationaintheimage,usingoneofthemethodsdiscussedinSec-tion
3.2
.Withthis,wecananswerthequestionbyfind-ing(q|i,t)=argmaxaeAs(q,a|i,t),thatmaximizesthescores(q,a|i,t)=s(ia,tq),whichspecializesEq.(
1
).
Inthefollowingsections,weprovidefurtherdetailsandapplytheseideastoafewconcretetasks.
3.2.Visualpromptingviamarking
Theusualwayofencodinglocationinformationinavi-sualpromptistocroptheimagearoundthedesiredlocation,meaningthatiaistheimagecroppedarounda.ThisideahasbeenusedextensivelywithVLMs,includingtointer-pretreferringexpressions,wheremaximizingascoreoftheforms(ia,tq)seeksfortheimagecropthatbestmatchesthereferringexpressiontq.
...
...
cat
VLM
Thisisa
...
bird
bear
...
VLM
eareyenose
Classification
Q
A
...
bear
Referringexpressionscomprehension
A
Q
VLM
Thecub
ontheright
Namingkeypoints
A/Q
Q/A
The
ear
eye
nose
ofabear
Figure2:PromptengineeringforVLMs.Wecastzero-shotinferencewithVLMsasaQ/Aproblem,eachrequiringspecificpromptengineering.Inthefigure,QisQuestionandAisAnswer(asetofpossibleanswers).Left:textpromptengineeringforclassification.Thiswidelyusedmethodcanbeinterpretedasfollowsinourframework:Theimageisthequestion,andclassesaretheavailableanswers,whichareengineeredintoprompts.Middle:visualpromptengineeringforreferringexpressionscomprehension.Thequestionisthereferringexpression,andtheavailableanswersaretheboxproposals,whichweengineerintovisualprompts.Right:visualandtextpromptengineeringforkeypointmatching.Forkeypointlocalization,weuseasimilarsetuptoreferringexpressions,wherethequestionisakeypointinplaintextandthepossibleanswersareall2Dlocationsintheimage.
Inthispaper,weexploreanalternativeapproachforvi-sualpromptingthatusestheconceptofmarkingthedesiredregionintheimage.Markingquiteliterallymeansover-layingtotheimageiaacircle,abox,oranarrow,whichvisuallyindicatesthedesiredlocationa.
Whiletheideaofmarkingmaysoundstrange,itisinter-estingfortworeasons.First,differentlyfromcropping,amarkedimageiapreservesalmostalltheinformationcon-tainedintheinputimagei,includingcontextualinforma-tionthatcropslack.Second,weshowthatmarkingworkswellwithVLMs,outperformingcropping-basedprompten-gineeringinsomepredictiontasks.
Whilethesimplestmarkingconsistingofaredcircleisparticularlyeffective,inSection
4
weexploreseveraldif-ferentwaysofgeneratingmarkings.Wereferthereadertothatsectionforfurtherdetailsandexamples.
3.3.Tasks
Westudytheideaofmark-basedpromptengineeringbyconsideringseveralzero-shotpredictiontasks,fromsimpletaskssuchasmatchingkeypointstotheirnamestomorecomplexonessuchasreferringexpressioncomprehension.
NamingKeypoints.Thefirstandsimplesttaskthatweconsiderismatchingthenameofthekeypointsofanobjecttotheir2Dlocationsinanimage.Theinputisanimagei,asetofkeypointnamesQ,andasetofcorrespondingkeypointlocationsAc{0,...,H_1}x{0,...,W_1}.Thenumberofnamesandlocationsisthesame(m=|Q|=|A|)andthegoalistomatchthetwo.WeexpressthelatteraspredictingthesquarepermutationmatrixΠeSmthat
associateseachnameqtoitscorrespondinglocationa(i.e.,Πqa=1).
InordertopredictΠ,weuseEq.(
1
)todefinethecostofassociatingnameqtolocationaasCqa=s(ia,tq)whereiaisobtainedeitherviacroppingormarkingandtqisjustthenameofthekeypointsprefixedbythestring“animageof”.Forthisproblem,theroleofquestionsandanswersissymmetricandwedecodethecostmatrixCintoapermu-tationmatrixΠviaoptimaltransport:
(i,Q,A)=argmaxΠqaexp(_τCqa)
ΠeS‰qeQ,aeA
whereτ>0isatemperatureparameter.Thisoptimiza-tionproblemissolvedefficientlyviatheSinkhorn-Knoppalgorithm[
44
],whichrenormalizesmatrixC.
KeypointLocalization.Thesecondtaskisamoreusefulanddifficultvariantofthefirst.Thegoalisstilltolocalizeanamedkeypointqinanimage,butthistimethelocationsAareasubsetofamxmregulargrid.Thesearefurtherrestrictedtoasalientimageregionextractedbyusingtheunsupervisedsaliencymethodof[
50
]toavoidtestingirrel-evantlocationsinthebackground.Thedifferencecomparedtonamingkeypointsisthatthisversionoftheproblemdoesnotassumepriorknowledgeofthepossiblelocationsofthekeypoints.Giventhenameqofakeypoint,itslocationaisthenobtainedas(i,q)=argmaxaeAs(ia,tq)whereiaandtqareasdefinedpreviously.
ReferringExpressionComprehension.Comprehendingareferringexpressionmeansdetectinganobjectinanim-agethatcorrespondstoatextualspecificationthatexplicitly
Keypoint-to-name
Name-to-keypoint
Method
CUB
Spair71k
CUB
SPair71k
Random
8.2
16.8
15.0
10.5
9.4
15.1
11.9
8.2
16.8
15.0
10.5
9.4
15.1
11.9
Cropw/oSK
15.8
28.5
28.5
28.5
20.1
26.1
29.7
18.7
22.4
19.1
24.0
14.9
27.3
25.1
Cropw/SK
25.5
35.1
37.5
34.6
23.9
32.9
36.3
25.8
36.1
32.5
32.7
19.8
35.3
32.5
RedCirclew/oSK
46.5
54.8
53.1
51.6
40.1
47.4
45.2
29.5
26.8
24.9
36.9
18.8
31.8
28.9
RedCirclew/SK
58.2
67.6
60.1
59.3
53.1
56.7
52.8
56.5
67.2
54.4
59.7
49.8
56.6
53.0
Table1:NamingkeypointsresultsonCUBandSPair71k.OnSPair71k,weshowresultsonallanimalclasses—bird,cat,dog,horse,sheep,cow.Weshowthepercentageofcorrectlymatchedkeypointsandnames,givenalistofthem.Wecomparetorandomlyguessingthecorrectcorrespondenceandcroppingaroundtheregionofinterest,ratherthandrawinganannotation.WecompareallmethodswithandwithoutnormalizationwiththeSinkhorn-Knopp(SK)algorithm.
referstoit(e.g.,“fourthdogfromtheright”).Similarlytopriorwork[
20
,
45
,
57
],givenanimagei,weapproachthisproblembyextractingfirstasetofobjectproposalsusingthemethodfrom[
58
]andinterpretthoseasthesetofpossi-bleanswersA.ThesetofquestionsQisinsteadacollectionofreferringexpressionsextractedfromagivenbenchmarkdataset.Foreachreferringexpression,thebestmatchingproposalisthengivenby
(i,q)=argmaxaeA
s(ia,tq)_eQs(ia,t).
TheengineeredpromptsiaandtqaredefinedasinSec-tion
3.3
.Inthiscase,wefounditusefultosubtractfromthescoretheaveragewithrespecttoallpossiblereferringexpressionsQ.Thisweighsdownhypothesesasuchasfacesthatarevisuallyverysalientandtendtorespondverystronglytoallquestionsq.
4.Experiments
WestudythepropertiesofvisualmarkinginVLMsbyconsideringfirstthethreetasksofSection
3.3
:namingkey-points,localizingkeypoints,andreferringexpressioncom-prehension.
4.1.NamingKeypoints
Namingkeypointsisacomparativelysimpleproblemthathasnodirectapplication;however,itissimplerandfastertoevaluatethantheothertasks,soweuseittoablatevariousaspectsofourmethod.
Dataandimplementationdetails.Forthistask,wecon-sidertheCUB-200-2011(CUB)[
52
]andSPair71k[
36
]datasets.Thefirstcontainsnamedkeypointannotationsforeachimage,whereasthesecondonlyannotatesmatchingkeypointsinpairsofimages,butdoesnotnamethem.Wethusaugmentthelatter,manuallynamingeachkeypointin-stanceineachanimalimage.WefurthercroptheimagesfromSPair71kwiththeprovidedboundingboxes.Forthe
tailbeakcrownforehead
leftleg
rightleg
righteye
nape
Figure3:QualitativeResultsonLocalizingKeypointsonanimagefromSPair71k.Greenandred(dashed)bor-dersareforcorrectandwrongpredictionsaccordingtoPCKwithα=0.1.Theredcircleshownisthickerthantheoneused;seethesup.matt.forexamplesoftheactualthickness.
VLM,weusetheViT-L/14@336pxbackbone.Pleaseseethesup.matt.fordetails.
Results.Recallthat,inthistask,theoutputofthepre-dictorisapermutationmatrixΠassociatingeachkeypointlocationtoacorrespondingname.Wereport(i)theratioofkeypointnamesthataremappedtothecorrectlocationand(ii)theratioofkeypointlocationsthataremappedtothecorrectnames.Tothebestofourknowledge,therearenopriorworksthatassociatekeypointswiththeirnames.Wethuscomparetheresultofthisnewtaskto(a)randomchoiceand(b)abaselinewhereiaisobtainedbycropping. AsseeninTable
1
,promptingviavisualmarking(redcircles)significantlyoutperformsthebaselines,achievingalmosttwicetheaccuracy.UsingtheSinkhorn-Knopp(SK)
algorithmtonormalizethematchingscorefurtherboostsre-sults,mainlyimprovingresultsforpointsthatareambigu-ousandclosetoeachother,e.g.,mouthandnose.
Whatisthebestvisualmarker?Wecomparetheuseof(i)differentshapesforhighlightingalocation:circle,rectangle,cross,arrow,(ii)differentsizes,and(iii)dif-ferentcolorsoftheannotations,andshowsomeexamplesinFig.
4
.WecomparedifferentshapesandcolorsinTable
2
Markershape
Mean
Best
Circle
33.5士4.5
46.5
Arrow
283士3.1
36.3
Square
24.1士3.6
36.3
Cross
21.5士6.3
34.5
CLIP[39]ViTB/32400M87M191267191256
Circlecolor
Mean
Best
Red
36.4士5.1
46.5
Green
343士4.2
43.3
Purple
34.0士3.7
41.9
Blue
32.7士3.9
41.1
Yellow
32.4士4.0
40.8
Table2:Ablationofannotationtypesfornamingkey-point.We
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 安徽六安皋城中学2026-2027学年九年级上学期阶段性目标检测语文试题(一)(含答案)
- 海水淡化工常识能力考核试卷含答案
- 电器附件零部件制造工岗中安全强化考核试卷含答案
- 印花电脑分色工安全应急能力考核试卷含答案
- 电切削工安全知识竞赛评优考核试卷含答案
- 电商平台客户关系管理实务手册
- 物联网安装调试员岗位技能掌握考核试卷含答案
- 白蚁防治工岗中应急响应预案考考核试卷含答案
- 巧克力塑形师安全宣教竞赛考核试卷含答案
- 矿用高空作业车司机岗中水平强化考核试卷含答案
- T-CECS 486-2017《数据中心供配电设计规程》
- 儿科护理工作压力管理
- 员工调动管理制度
- 链家员工合同
- CAM制造软件厂商竞争格局研究市场调研报告
- 四不伤害及反三违安全培训课件
- 建筑工程技术课程
- 量力而行议论文
- 《心灯录》完整版
- 2026届新高考英语热点冲刺复习:定语从句
- 《盗窃案件的侦查》课件
评论
0/150
提交评论