大型语言-视觉模型的视觉提示工程 What does CLIP know about a red circle Visual prompt engineering for VLMs_第1页
大型语言-视觉模型的视觉提示工程 What does CLIP know about a red circle Visual prompt engineering for VLMs_第2页
大型语言-视觉模型的视觉提示工程 What does CLIP know about a red circle Visual prompt engineering for VLMs_第3页
大型语言-视觉模型的视觉提示工程 What does CLIP know about a red circle Visual prompt engineering for VLMs_第4页
大型语言-视觉模型的视觉提示工程 What does CLIP know about a red circle Visual prompt engineering for VLMs_第5页
已阅读5页,还剩30页未读, 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

arXiv:2304.06712v1[cs.CV]13Apr2023

WhatdoesCLIPknowaboutaredcircle?

VisualprmptengineeringforVLMs

AleksandarShtedritskiChristianRupprechtAndreaVedaldi

VisualGeometryGroup,UniversityofOxford

(suny,chrisr,vedaldi}@robots.ox.ac.uk

Abstract

Large-scaleVision-LanguageModels,suchasCLIP,learnpowerfulimage-textrepresentationsthathavefoundnumerousapplications,fromzero-shotclassificationtotext-to-imagegeneration.Despitethat,theircapabilitiesforsolvingnoveldiscriminativetasksviapromptingfallbehindthoseoflargelanguagemodels,suchasGPT-3.Hereweex-ploretheideaofvisualpromptengineeringforsolvingcom-putervisiontasksbeyondclassificationbyeditinginimagespaceinsteadoftext.Inparticular,wediscoveranemer-gentabilityofCLIP,where,bysimplydrawingaredcir-clearoundanobject,wecandirectthemodel’sattentiontothatregion,whilealsomaintainingglobalinformation.Weshowthepowerofthissimpleapproachbyachievingstate-of-the-artinzero-shotreferringexpressionscomprehensionandstrongperformanceinkeypointlocalizationtasks.Fi-nally,wedrawattentiontosomepotentialethicalconcernsoflargelanguage-visionmodels.

1.Introduction

LargeLanguageModels(LLMs)suchasGPT-2/3[

7

,

40

]andChatGPT[

1

]havedemonstratedsurprisingemergingbehaviours.Forexample,thesemodelscanperformlan-guagetranslationwithoutbeingexplicitlytrainedforit,inazero-shotmanner.Thiscanbepartiallyexplainedbythefactthatoccurrencesofthedesiredbehaviours,suchastranslatingbetweentwolanguages,naturallyoccurintheirenormoustrainingcorpus,whichis,essentially,theInternet.

InterestingemergentbehaviourshavebeenobservedinlargeVision-LanguageModels(VLMs)likeCLIP[

39

]too.Forexample,CLIPcanbeusedforzero-shotclassifica-tionbycheckingthecompatibilityofagivenimagewithpromptssuchas“animageofaX”,whereXisoneofasetofclasshypothesestobetested.

EmergentbehavioursareelicitedbysupplyingsuitablycraftedinputstotheVLMs,oftencalledprompts.Asintheexampleabove,researchershavemostlyfocusedonengi-

ThecowthatisthesmallestThewhitelittlelamb

forehead

righteye

mouth

nose

Figure1:VisualPromptEngineering.WedrawmultipleannotationsoveranimageandhaveCLIPchoosethecorrectonegivenacaption.Hereweshowpredictionsforthegivenexpressions.Top:ExamplesfromRefCOCOgonreferring

expressionsdetection.Bottom:ExamplefromSPair71konkeypointlocalization.

neeringtextualprompts,manipulatingthetextualinputofthemodel.ThisapproachisinspiredbyLLMs,wherema-nipulatingthetextualmodalityistheonlyavailableoption.However,VLMsareinherentlymultimodalandofferthepossibilityofmanipulatingbothmodalities,textualandvi-sual.Whilethetextualmodalityisthenaturalchoiceforexpressingsemantics,thevisualmodalitycanbebetterforexpressinggeometricpropertiessuchaslocation.

Inthispaper,wethusexplorevisualpromptengineer-ing

1

.Wedosowithtwogoals.Thefirstgoalistocon-tributeonemorepracticaltoolforextractingusefulinfor-mationfromVLMsinazero-shotmanner.Wedemon-stratethisbyobtainingstate-of-the-artzero-shotresultsinreferringexpressionscomprehensionbyengineeringvisualprompts.Thesecondgoalistocharacteriseinterestingand

1Notethedifferencebetweenvisualprompttuning,asettingpreviouslyexplored,wherethepromptsaretask-specificlearnabletokens,andvisualpromptengineering,whereweapplyafixedaugmentationinpixelspace.

unexpectedpropertiesoftheVLMsandtheirtrainingdata,includingidentifyingsomebehavioursthatcanraiseethicalconcerns.

Perhapsthemostsurprisingofourfindingsistheeffec-tivenessofaparticulartypeofvisualprompting:drawingaplainredcircleontopoftheimage(Fig.

1

).WeshowthatthissimpleinterventionsteerstheVLMtoanalyse/talkabouttheimageregioncontainedinthecircle.Thisbe-haviourcanthenbeusedfortaskssuchasnamingaspe-cificobjectorobjectpartordetectingparticularimagere-gionsbasedonadescription.Thelatter,forinstance,isachievedbymarkingeachobjectproposalwitharedcircleandusingtheVLMtofindthebestmatchwithrespecttotheprovidedreferringexpression,achievingstrongresultsonmultiplebenchmarksintheunsupervisedregime.Fur-thermore,weshowthatpromptingwithacirclealsoworksforfiner-grainedlocalization,markingspecificobjectpartsorkeypointsinsteadofjustwholeobjects.

Wefurthercontrastmarkinganimagewiththealterna-tiveofcroppingit,which,fromslidingwindowclassifierstoregionneuralnetworks,isthecanonicalapproachtosteerthefocusofanimage-levelpredictortoaparticularimageregion.Weshowthat,forVLMsatleast,markingissig-nificantlymoreeffectivethancropping,possiblybecauseitdoesnotlosecontextualinformationlikethelatter.

Apartfromthepracticalapplications,ourfindingsrevealunexpectedandintriguingpropertiesofVLMs.Weshowempirically,thatmarkingwitharedcircleisoptimalamongaselectionofpossiblemarkers(variantsofthecircle,boxes,arrows,etc.).Presumably,theVLMsunderstandredcirclesoutoftheboxbecausetheseappearsufficientlyfrequentlyinthetrainingcorpus,i.e.,theInternet.WhilewedonothaveaccesstothefulltrainingdataofCLIP,wecorrobo-ratethisintuitionbyseekingexamplesofsuchimagesinYFCC15M,adatasetofCC-BYimages.

Ouranalysisshowsthatredcirclesareindeedpresentevenina(comparativelysmall)datasetofimageslikeYFCC15M,buttheyarerare.Itisatestamenttotheex-traordinarycapacityofVLMsthatsuchabehaviourcanbelearnedfromsuchrareevents,withoutanexplicitfocusondoingso.Wetestmodelsofdifferentsizes/capacitiesandshowthatonlythelargermodelsexhibitthisbehaviourre-liably,whichwellcorrespondstoourintuition.

Finally,wenotethattheabilityofVLMstolearnevenfromrareeventssuch“redcircles”canacquirebothdesir-ableandundesirablebehaviours.Redcircles,inparticular,canhaveanegativeconnotationinthetrainingdataastheyareoftenusedbynewsoutletstomarkmissingpeopleorcriminalsand,evidently,themodellearnsfromsuchexam-ples.Asaresult,weshowthatdrawingaredcircleinanimageincreasestheprobabilitythatthemodelwouldchar-acteriseapersonasacriminalorasamissingperson.

Tosummarise,wemakethefollowingmaincontribu-

tions:(1)Weproposemarkingasanewformofvisualpromptengineeringwhichiseffectiveinextractinguse-fulemergentbehavioursinVLMslikeCLIP;(2)Weusethelattertoachievestate-of-the-artzero-shotreferringex-pressionscomprehensionusingaVLM;(3)Weprovideananalysisofwhymarkingiseffectiveforthesemodels,andlinkthattothetrainingdataandlargemodelcapacity;(4)Weshowthatvisualpromptengineeringcanalsoelicitun-wantedbehaviours,suchastriggeringproblematicbiasesintheVLMs,revealingpotentialethicalissues.

2.Relatedwork

EmergentBehaviourfromLargeScalePretraining

hasmainlybeenobservedinLargeLanguageModels(LLMs).Mostnotably,GPT-2[

40

],GPT-3[

7

],andChat-GPT[

1

]havebeenshowntobecapableoftaskssuchaszero-shottranslation,questionanswering,arithmetic,aswellasplanningactionsforembodiedagents[

18

].Fine-tuningLLMscanalsoleadtomodelsthatcangeneratecodefromdocstrings[

8

]orsolvemathproblems[

13

,

24

].Onlyafewemergentzero-shotbehaviourshavebeenre-portedforVLMslikeCLIP,mainlyforclassification[

39

]andOCR[

34

].GenerativeVLMslikeFLAMINGO[

3

]andBLIP[

25

]excelincaptioningandvisualquestion-answeringtasks,butalsohavenowayofsolvingpixel-levelcomputervisiontasks.

PromptingVLMsismostcommonlyperformedbyprependingasetoflearnabletokenstothetextinput[

16

,

21

,

63

,

64

],visioninput[

19

,

47

,

62

],orbothtextandvi-sioninputs[

42

,

60

],inordertoeasilysteerafrozenCLIPmodeltosolveadesiredtask.[

4

]learnaugmentationsinpixelspace,suchaspaddingaroundtheimage,orchang-ingapatchoftheimage,whichareoptimizedwithgradientdescentonadownstreamtask.[

5

]castimageinpaintingasavisualpromptingtask,usingagenerativemodeltrainedonfiguresfromacademicpapers.ColorfulPromptTuning(CPT)[

57

]colorregionsofanimageanduseacaptioningmodeltopredictwhichobjectinanimageanexpressionreferstobypredictingitscolor.SimilarlytoCPT,weaug-menttheinputimageinpixelspaceandperformzero-shotinference.However,weannotatetheimageinahuman-likemannerandshowthatourmethodismorepowerful

andmoreflexiblethanCPT.

ReferringExpressionComprehension(REC)aimstolocalizeatargetobjectinanimagethatcorrespondstoatex-tualdescription.MostapproachestoRECstartwithobjectproposals,forexample,generatedwithFaster-RCNN[

41

],andlearntoscorethem[

17

,

31

,

32

,

49

,

54

].RECissome-timesconsideredtogetherwithreferringexpressiongener-ation—thetaskofgeneratingadescriptionofagivenre-gion.[

32

]useacomprehensionmodeltoguideagenerator,whereas[

9

]jointlytrainadetectorwithacaptiongenera-tor.Someworksmodelthesceneasagraph[

31

,

49

,

54

]or

uselanguageparsersandgrammar-basedmethods[

12

,

30

],leadingtoamoreinterpretableresult.Morerecently,trans-formerarchitectureshavebeenused[

14

,

22

,

26

,

51

,

55

].[

14

,

22

,

26

,

55

]performtext-modulatedobjectdetection,whereatransformerdecodertakesthereferringexpressionasaninputandpredictsaboundingbox.[

51

]trainwithatext-to-pixelcontrastiveloss,whichallowsforatext-drivensegmentationordetectionattesttime.

UnsupervisedReferringExpressionComprehension

isalessexploredarea,onlymadepossiblewiththein-troductionoflargepretrainedmodelssuchasCLIP[

39

].ReCLIP[

45

]cropsobjectproposalsandranksthemus-ingCLIPbeforeanad-hocpostprocessingsteptotakeintoaccountrelationssuchasleft/right,smaller/bigger,etc.CPT[

57

]colorsobjectproposalboxesanduseapre-trainedcaptioningmodel[

61

]toauto-regressivelypredictwhichcoloredproposalcorrespondstothequerydescrip-tion.Pseudo-Q[

20

]generatesdescriptionsformultipleob-jectsinanimage,whichisusedtotrainaRECnetwork.

However,thismodelisnotfullyunsupervisedasthepseudodescriptionsitusesaregeneratedusingacaptioningmodel

trainedonCOCO.

VisualReasoningUsingLargePretrainedModelshasbeenanareaofsignificantinterestinthelastfewyears.Inadditiontoreferringexpressiondetection[

45

],CLIP[

39

]hasbeenusedforsemanticsegmentation[

28

,

38

].[

38

]useCLIPtoassigntextlabelstoobjectpartsafterdoingpartco-segmentationinthelatentspaceofaGAN.[

28

]utilizeCLIPforopen-vocabularysegmentationbyusingageneral-purposemaskproposalnetworkandCLIPasaclassifier.CLIPhasalsobeenusedforunsupervisedobjectproposalgeneration[

43

]andopen-setdetection[

15

].Semanticseg-mentationalsoemergesfromimageonly[

35

,

50

]orimage-text[

53

]self-supervision.

BiasofVLMsisanincreasinglypopularareaofre-search,asdownstreamapplicationscomewiththeriskofperpetuatingbiasesandstereotypesexistinginthetrainingdata.However,methodsforassessingthebiasofaVLMarestillnotwellestablished.[

2

]measurethemisclassificationrateofCLIPoffacesofpeopleofdifferentraceswithnon-humanandcriminalcategories,whereas[

6

,

11

,

48

]measurefairnessinretrievalresults.Here,weshowadifferentkindofbias,wheretheadditionofaredcircleoverapersoncantriggeranegativeconnotation.

3.Method

OurgoalistodevelopvisualpromptinginVision-

LanguageModels(VLMs).VLMssolvepredictiontasksthatinvolvejointlyprocessingtextandimages.Forex-ample,modelssuchasCLIParetrainedtomatchtextandimagesamples.TheinputtosuchaVLMisanimageieR3×H×WandtextteΣ*,whereΣisanalphabet.Theoutputisascores(i,t)thatexpressesthedegreeof

compatibilitybetweenthesuppliedimageandtext.

3.1.Promptengineering

OneofthemoststrikingcapabilitiesofVLMsistheirabilitytosolveavarietyofclassificationtaskswithlittletonofurthertrainingatall,inazero-shotmanner.ThisisdonebyreducingthetaskofinteresttothatofevaluatingtheVLMonsuitably-engineeredimageandtextpairs.

Forexample,givenanimage-captionpair(i,t),considertheproblemoflocalizinganamedobjectkeypointintheim-age.Wecancastthisasaquestion-answerproblem,wherethequestionqeQisthenameoftheobjectkeypoint(e.g.,“rightear”,“frontleftleg”,...)andtheansweraeAisoneofadiscretesetofimagelocations.

BecausetheVLMcomputesacompatibilityscores(i,t)betweenanimageiandthetextt,itcannotbeusedtomapthequestionqtotheansweradirectly.However,viapromptengineering,wecanusetheVLMtoconstructacompati-bilityscores(q,a|i,t)betweenquestionandanswer,con-ditionedontheinputimage-textpair(i,t).Thisscoreisingeneralgivenbytheexpression

s(q,a|i,t)=s(iqa,tqa)(1)

whereiqaandtqaareversionsoftheinputimageandtext,obtainedbytransformingthelattertoreflectthequestion-answerpair(q,a).

ThespecificwayEq.(

1

)shouldbeappliedtoaprob-lemdependsonthespecificnatureofthelatter.Forex-ample,intheproblemoflocalizingthenamedkeypoints,itisnaturaltoencodethenameofthekeypointviathetextualmodalityandits2Dlocationviathevisualmodal-ity.Forinstance,inordertoanswerthequestionq=“rightear”foragiveninputimageiwithcaptiont=“dog”,wecanengineerthetextualprompttqa=tq=“animageoftherightearofadog”toencodeadescriptionofthenamedentity.Likewise,wecanengineerthevisualpromptiqa=iainsuchawayasto‘select’thelocationaintheimage,usingoneofthemethodsdiscussedinSec-tion

3.2

.Withthis,wecananswerthequestionbyfind-ing(q|i,t)=argmaxaeAs(q,a|i,t),thatmaximizesthescores(q,a|i,t)=s(ia,tq),whichspecializesEq.(

1

).

Inthefollowingsections,weprovidefurtherdetailsandapplytheseideastoafewconcretetasks.

3.2.Visualpromptingviamarking

Theusualwayofencodinglocationinformationinavi-sualpromptistocroptheimagearoundthedesiredlocation,meaningthatiaistheimagecroppedarounda.ThisideahasbeenusedextensivelywithVLMs,includingtointer-pretreferringexpressions,wheremaximizingascoreoftheforms(ia,tq)seeksfortheimagecropthatbestmatchesthereferringexpressiontq.

...

...

cat

VLM

Thisisa

...

bird

bear

...

VLM

eareyenose

Classification

Q

A

...

bear

Referringexpressionscomprehension

A

Q

VLM

Thecub

ontheright

Namingkeypoints

A/Q

Q/A

The

ear

eye

nose

ofabear

Figure2:PromptengineeringforVLMs.Wecastzero-shotinferencewithVLMsasaQ/Aproblem,eachrequiringspecificpromptengineering.Inthefigure,QisQuestionandAisAnswer(asetofpossibleanswers).Left:textpromptengineeringforclassification.Thiswidelyusedmethodcanbeinterpretedasfollowsinourframework:Theimageisthequestion,andclassesaretheavailableanswers,whichareengineeredintoprompts.Middle:visualpromptengineeringforreferringexpressionscomprehension.Thequestionisthereferringexpression,andtheavailableanswersaretheboxproposals,whichweengineerintovisualprompts.Right:visualandtextpromptengineeringforkeypointmatching.Forkeypointlocalization,weuseasimilarsetuptoreferringexpressions,wherethequestionisakeypointinplaintextandthepossibleanswersareall2Dlocationsintheimage.

Inthispaper,weexploreanalternativeapproachforvi-sualpromptingthatusestheconceptofmarkingthedesiredregionintheimage.Markingquiteliterallymeansover-layingtotheimageiaacircle,abox,oranarrow,whichvisuallyindicatesthedesiredlocationa.

Whiletheideaofmarkingmaysoundstrange,itisinter-estingfortworeasons.First,differentlyfromcropping,amarkedimageiapreservesalmostalltheinformationcon-tainedintheinputimagei,includingcontextualinforma-tionthatcropslack.Second,weshowthatmarkingworkswellwithVLMs,outperformingcropping-basedprompten-gineeringinsomepredictiontasks.

Whilethesimplestmarkingconsistingofaredcircleisparticularlyeffective,inSection

4

weexploreseveraldif-ferentwaysofgeneratingmarkings.Wereferthereadertothatsectionforfurtherdetailsandexamples.

3.3.Tasks

Westudytheideaofmark-basedpromptengineeringbyconsideringseveralzero-shotpredictiontasks,fromsimpletaskssuchasmatchingkeypointstotheirnamestomorecomplexonessuchasreferringexpressioncomprehension.

NamingKeypoints.Thefirstandsimplesttaskthatweconsiderismatchingthenameofthekeypointsofanobjecttotheir2Dlocationsinanimage.Theinputisanimagei,asetofkeypointnamesQ,andasetofcorrespondingkeypointlocationsAc{0,...,H_1}x{0,...,W_1}.Thenumberofnamesandlocationsisthesame(m=|Q|=|A|)andthegoalistomatchthetwo.WeexpressthelatteraspredictingthesquarepermutationmatrixΠeSmthat

associateseachnameqtoitscorrespondinglocationa(i.e.,Πqa=1).

InordertopredictΠ,weuseEq.(

1

)todefinethecostofassociatingnameqtolocationaasCqa=s(ia,tq)whereiaisobtainedeitherviacroppingormarkingandtqisjustthenameofthekeypointsprefixedbythestring“animageof”.Forthisproblem,theroleofquestionsandanswersissymmetricandwedecodethecostmatrixCintoapermu-tationmatrixΠviaoptimaltransport:

(i,Q,A)=argmaxΠqaexp(_τCqa)

ΠeS‰qeQ,aeA

whereτ>0isatemperatureparameter.Thisoptimiza-tionproblemissolvedefficientlyviatheSinkhorn-Knoppalgorithm[

44

],whichrenormalizesmatrixC.

KeypointLocalization.Thesecondtaskisamoreusefulanddifficultvariantofthefirst.Thegoalisstilltolocalizeanamedkeypointqinanimage,butthistimethelocationsAareasubsetofamxmregulargrid.Thesearefurtherrestrictedtoasalientimageregionextractedbyusingtheunsupervisedsaliencymethodof[

50

]toavoidtestingirrel-evantlocationsinthebackground.Thedifferencecomparedtonamingkeypointsisthatthisversionoftheproblemdoesnotassumepriorknowledgeofthepossiblelocationsofthekeypoints.Giventhenameqofakeypoint,itslocationaisthenobtainedas(i,q)=argmaxaeAs(ia,tq)whereiaandtqareasdefinedpreviously.

ReferringExpressionComprehension.Comprehendingareferringexpressionmeansdetectinganobjectinanim-agethatcorrespondstoatextualspecificationthatexplicitly

Keypoint-to-name

Name-to-keypoint

Method

CUB

Spair71k

CUB

SPair71k

Random

8.2

16.8

15.0

10.5

9.4

15.1

11.9

8.2

16.8

15.0

10.5

9.4

15.1

11.9

Cropw/oSK

15.8

28.5

28.5

28.5

20.1

26.1

29.7

18.7

22.4

19.1

24.0

14.9

27.3

25.1

Cropw/SK

25.5

35.1

37.5

34.6

23.9

32.9

36.3

25.8

36.1

32.5

32.7

19.8

35.3

32.5

RedCirclew/oSK

46.5

54.8

53.1

51.6

40.1

47.4

45.2

29.5

26.8

24.9

36.9

18.8

31.8

28.9

RedCirclew/SK

58.2

67.6

60.1

59.3

53.1

56.7

52.8

56.5

67.2

54.4

59.7

49.8

56.6

53.0

Table1:NamingkeypointsresultsonCUBandSPair71k.OnSPair71k,weshowresultsonallanimalclasses—bird,cat,dog,horse,sheep,cow.Weshowthepercentageofcorrectlymatchedkeypointsandnames,givenalistofthem.Wecomparetorandomlyguessingthecorrectcorrespondenceandcroppingaroundtheregionofinterest,ratherthandrawinganannotation.WecompareallmethodswithandwithoutnormalizationwiththeSinkhorn-Knopp(SK)algorithm.

referstoit(e.g.,“fourthdogfromtheright”).Similarlytopriorwork[

20

,

45

,

57

],givenanimagei,weapproachthisproblembyextractingfirstasetofobjectproposalsusingthemethodfrom[

58

]andinterpretthoseasthesetofpossi-bleanswersA.ThesetofquestionsQisinsteadacollectionofreferringexpressionsextractedfromagivenbenchmarkdataset.Foreachreferringexpression,thebestmatchingproposalisthengivenby

(i,q)=argmaxaeA

s(ia,tq)_eQs(ia,t).

TheengineeredpromptsiaandtqaredefinedasinSec-tion

3.3

.Inthiscase,wefounditusefultosubtractfromthescoretheaveragewithrespecttoallpossiblereferringexpressionsQ.Thisweighsdownhypothesesasuchasfacesthatarevisuallyverysalientandtendtorespondverystronglytoallquestionsq.

4.Experiments

WestudythepropertiesofvisualmarkinginVLMsbyconsideringfirstthethreetasksofSection

3.3

:namingkey-points,localizingkeypoints,andreferringexpressioncom-prehension.

4.1.NamingKeypoints

Namingkeypointsisacomparativelysimpleproblemthathasnodirectapplication;however,itissimplerandfastertoevaluatethantheothertasks,soweuseittoablatevariousaspectsofourmethod.

Dataandimplementationdetails.Forthistask,wecon-sidertheCUB-200-2011(CUB)[

52

]andSPair71k[

36

]datasets.Thefirstcontainsnamedkeypointannotationsforeachimage,whereasthesecondonlyannotatesmatchingkeypointsinpairsofimages,butdoesnotnamethem.Wethusaugmentthelatter,manuallynamingeachkeypointin-stanceineachanimalimage.WefurthercroptheimagesfromSPair71kwiththeprovidedboundingboxes.Forthe

tailbeakcrownforehead

leftleg

rightleg

righteye

nape

Figure3:QualitativeResultsonLocalizingKeypointsonanimagefromSPair71k.Greenandred(dashed)bor-dersareforcorrectandwrongpredictionsaccordingtoPCKwithα=0.1.Theredcircleshownisthickerthantheoneused;seethesup.matt.forexamplesoftheactualthickness.

VLM,weusetheViT-L/14@336pxbackbone.Pleaseseethesup.matt.fordetails.

Results.Recallthat,inthistask,theoutputofthepre-dictorisapermutationmatrixΠassociatingeachkeypointlocationtoacorrespondingname.Wereport(i)theratioofkeypointnamesthataremappedtothecorrectlocationand(ii)theratioofkeypointlocationsthataremappedtothecorrectnames.Tothebestofourknowledge,therearenopriorworksthatassociatekeypointswiththeirnames.Wethuscomparetheresultofthisnewtaskto(a)randomchoiceand(b)abaselinewhereiaisobtainedbycropping. AsseeninTable

1

,promptingviavisualmarking(redcircles)significantlyoutperformsthebaselines,achievingalmosttwicetheaccuracy.UsingtheSinkhorn-Knopp(SK)

algorithmtonormalizethematchingscorefurtherboostsre-sults,mainlyimprovingresultsforpointsthatareambigu-ousandclosetoeachother,e.g.,mouthandnose.

Whatisthebestvisualmarker?Wecomparetheuseof(i)differentshapesforhighlightingalocation:circle,rectangle,cross,arrow,(ii)differentsizes,and(iii)dif-ferentcolorsoftheannotations,andshowsomeexamplesinFig.

4

.WecomparedifferentshapesandcolorsinTable

2

Markershape

Mean

Best

Circle

33.5士4.5

46.5

Arrow

283士3.1

36.3

Square

24.1士3.6

36.3

Cross

21.5士6.3

34.5

CLIP[39]ViTB/32400M87M191267191256

Circlecolor

Mean

Best

Red

36.4士5.1

46.5

Green

343士4.2

43.3

Purple

34.0士3.7

41.9

Blue

32.7士3.9

41.1

Yellow

32.4士4.0

40.8

Table2:Ablationofannotationtypesfornamingkey-point.We

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论