评估基于代码训练的大型语言模型_第1页
评估基于代码训练的大型语言模型_第2页
评估基于代码训练的大型语言模型_第3页
评估基于代码训练的大型语言模型_第4页
评估基于代码训练的大型语言模型_第5页
已阅读5页,还剩58页未读 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

EvaluatingLargeLanguageModelsTrainedonCode

MarkChen*¹JerryTworek*¹HeewooJun*¹QimingYuan*¹HenriquePondedeOliveiraPinto*1

JaredKaplan*²HarriEdwards¹YuriBurda¹NicholasJoseph²GregBrockman¹AlexRay¹RaulPuri¹

GretchenKrueger¹MichaelPetrov¹HeidyKhlaaf³GirishSastry¹PamelaMishkin¹BrookeChan¹

ScottGray¹NickRyder¹MikhailPavlov¹AletheaPower¹LukaszKaiser¹MohammadBavarian¹

ClemensWinter¹PhilippeTillet¹FelipePetroskiSuch¹DaveCummings¹MatthiasPlappert¹

FotiosChantzis¹ElizabethBarnes¹ArielHerbert-Voss¹WilliamHebgenGuss¹AlexNichol¹AlexPaino¹

2021

NikolasTezak¹JieTang¹IgorBabuschkin¹SuchirBalaji¹ShantanuJain¹WilliamSaunders¹

ChristopherHesse¹AndrewN.Carr¹JanLeike¹JoshAchiam¹VedantMisra¹EvanMorikawa¹

AlecRadford¹MatthewKnight¹MilesBrundage¹MiraMurati¹KatieMayer¹PeterWelinder¹

BobMcGrew¹DarioAmodei²SamMcCandlish²IlyaSutskever¹WojciechZaremba¹

arXiv:2107.03374v2[cs.LG]14Jul

Abstract

WeintroduceCodex,aGPTlanguagemodelfine-tunedonpubliclyavailablecodefromGitHub,andstudyitsPythoncode-writingcapabilities.AdistinctproductionversionofCodexpowersGitHubCopilot.OnHumanEval,anewevalua-tionsetwereleasetomeasurefunctionalcorrect-nessforsynthesizingprogramsfromdocstrings,ourmodelsolves28.8%oftheproblems,whileGPT-3solves0%andGPT-Jsolves11.4%.Fur-thermore,wefindthatrepeatedsamplingfromthemodelisasurprisinglyeffectivestrategyforpro-ducingworkingsolutionstodifficultprompts.Us-ingthismethod,wesolve70.2%ofourproblemswith100samplesperproblem.Carefulinvestiga-tionofourmodelrevealsitslimitations,includingdifficultywithdocstringsdescribinglongchainsofoperationsandwithbindingoperationstovari-ables.Finally,wediscussthepotentialbroaderimpactsofdeployingpowerfulcodegenerationtechnologies,coveringsafety,security,andeco-nomics.

1.Introduction

Scalablesequencepredictionmodels(Graves,2014;

Vaswanietal.,2017;Childetal.,2019)havebecomeageneral-purposemethodforgenerationandrepresentationlearninginmanydomains,includingnaturallanguagepro-cessing(Mikolovetal.,2013;Sutskeveretal.,2014;Dai&Le,2015;Petersetal.,2018;Radfordetal.,2018;Devlinetal.,2018),computervision(VanOordetal.,2016;Menick&Kalchbrenner,2018;Chenetal.,2020;Baoetal,2021),audioandspeechprocessing(Oordetal.,2016;2018;Dhari-waletal.,2020;Baevskietal.,2020),biology(Alleyetal.,2019;Rivesetal.,2021),andevenacrossmultiplemodali-ties(Dasetal.,2017;Luetal.,2019;Rameshetal.,2021;Zellersetal.,2021).Morerecently,languagemodelshavealsofueledprogresstowardsthelongstandingchallengeofprogramsynthesis(Simon,1963;Manna&Waldinger,

1971),spurredbythepresenceofcodeinlargedatasets(Husainetal.,2019;Gaoetal.,2020)andtheresultingpro-grammingcapabilitiesoflanguagemodelstrainedonthesedatasets(Wang&Komatsuzaki,2021).Popularlanguagemodelingobjectiveslikemaskedlanguagemodeling(Devlinetal.,2018)andspanprediction(Raffeletal.,2020)havealsobeenadaptedtotraintheirprogrammingcounterpartsCodeBERT(Fengetal.,2020)andPyMT5(Clementetal.,2020).

Equalcontribution

OpenAI,SanFrancisco,California,USA.

²AnthropicAI,SanFrancisco,California,USA.Workper-formedwhileatOpenAI.

³Zipline,SouthSanFrancisco,California,USA.Workper-formedwhileatOpenAI.

Correspondenceto:MarkChen<mark@>,JerryTworek<jt@>,HeewooJun<hee-woo@>,QimingYuan<qiming@>.

Similarly,ourearlyinvestigationofGPT-3(Brownetal.,2020)revealedthatitcouldgeneratesimpleprogramsfromPythondocstrings.Whilerudimentary,thiscapabilitywasexcitingbecauseGPT-3wasnotexplicitlytrainedforcodegeneration.Giventheconsiderablesuccessoflargelan-guagemodelsinothermodalitiesandtheabundanceofpubliclyavailablecode,wehypothesizedthataspecializedGPTmodel,calledCodex,couldexcelatavarietyofcodingtasks.ThispaperdescribesseveralearlyCodexmodels,whosedescendantspowerGitHubCopilotandtheCodexmodelsintheOpenAIAPI.

EvaluatingLargeLanguageModelsTrainedonCode

CodexandCodex-SPerformance

PasSrate

GPT-3pass@1

Codexpass@1

一Codex-Spass@1

—Codex-SmeanlogprerankingCodex-Soraclereranking

0.2

0.0

10⁵10⁶10⁷10⁸10⁹1010

0.8

0.6

0.4

Non-embeddingparameters

Figure1.PassratesofourmodelsontheHumanEvaldatasetasafunctionofmodelsize.Whenasinglesampleisgeneratedforeachproblem,GPT-12Bsolvesnoproblems,butCodex(fine-tunedoncode)solves28.8%oftheproblems,andCodex-S(furtherfine-tunedoncorrectlyimplementedstandalonefunctions)solves37.7%oftheproblems.Fromhere,furthergainscanberealizedbygenerating100samplesperproblemandselectingthesamplewiththehighestmeanlog-probability(44.5%solved)orbyselectingthesamplethatpassestheunittests(77.5%solved).Allsamplesaregeneratedwithtemperature0.8.

Inthiswork,wefocusonthetaskofgeneratingstan-dalonePythonfunctionsfromdocstrings,andevaluatethecorrectnessofcodesamplesautomaticallythroughunittests.Thisisincontrasttonaturallanguagegeneration,wheresamplesaretypicallyevaluatedbyheuristicsorbyhumanevaluators.Toaccuratelybenchmarkourmodel,wecreateadatasetof164originalprogrammingproblemswithunittests.Theseproblemsassesslanguagecompre-hension,algorithms,andsimplemathematics,withsomecomparabletosimplesoftwareinterviewquestions.Wereleasethisdataalongwithanevaluationframeworkat

/openai/human-eval.

Tosolveaprobleminourtestset,wegeneratemultiplesamplesfromthemodels,andcheckifanyofthempasstheunittests.Withjustasinglesample,a12BparameterCodexsolves28.8%oftheseproblems,anda300MparameterCodexsolves13.2%oftheseproblems.Incontrast,the6BparameterGPT-J(Wang&Komatsuzaki,2021)achieves

11.4%onthesamedataset,whileallGPTmodelsachievenear0%.Toimproveourmodel'sperformanceatthetaskoffunctionsynthesisfromdocstrings,wefine-tuneCodexonstandalone,correctlyimplementedfunctions.Theresultingmodel,Codex-S,solves37.7%ofproblemswithasinglesample.Figure2showcasesproblemsofvaryingdifficultyinourdataset,alongwithcorrectmodelgeneratedsolutions.

Real-worldprogrammingtasksofteninvolveiterationsofapproachesandbugfixes,whichisapproximatedbygener-atingmanysamplesfromourmodelsandselectingonethatpassesallunittests.Within100samples,Codex-Sisableto

generateatleastonecorrectfunctionfor77.5%oftheprob-lems.Thisresultsuggeststhataccuratecodesamplescanbeselectedviaheuristicrankinginsteadoffullyevaluatingeachsample,thelatterofwhichmaynotbepossibleorprac-ticalindeployment.Indeed,wefindthatthesamplewithhighestmeanlog-probabilitypassesunittestsfor44.5%oftheproblems.

WeconcludebydiscussingthelimitationsandpotentialbroaderimpactsoftheseCodexmodelsandofincreasinglypowerfulcodegeneratingmodelsmoregenerally.

2.EvaluationFramework

Inthissection,wediscussthedetailsofourevaluationframework.Webeginbydefiningthepass@kmetric,andexplainitsadvantagesoverstandardmatch-basedmetrics.Next,wedescribethedatasetofhand-writtenproblems,called"HumanEval,"whichwecreatedinordertobench-markourmodels.Finally,wediscussthesandboxenviron-mentweusedtosafelyexecutemodel-generatedcode.

2.1.FunctionalCorrectness

Generativemodelsforcodearepredominantlybenchmarkedbymatchingsamplesagainstareferencesolution,wherethematchcanbeexactorfuzzy(asinBLEUscore).How-ever,recentworkhassurfaceddeficienciesinmatch-basedmetricsforcode.Forinstance,Renetal.(2020)findsthatBLEUhasproblemscapturingsemanticfeaturesspecifictocode,andsuggestsseveralsemanticmodificationstothe

score.

Morefundamentally,match-basedmetricsareunabletoac-countforthelargeandcomplexspaceofprogramsfunction-allyequivalenttoareferencesolution.Asaconsequence,recentworksinunsupervisedcodetranslation(Lachauxetal.,2020)andpseudocode-to-codetranslation(Kulaletal.,2019)haveturnedtofunctionalcorrectnessinstead,whereasampleisconsideredcorrectifitpassesasetofunittests.Wearguethatthismetricshouldbeappliedtodocstring-conditionalcodegenerationaswell.

Perhapsthemostconvincingreasontoevaluatefunctionalcorrectnessisthatitisusedbyhumandeveloperstojudgecode.Aframeworkknownastest-drivendevelopmentdic-tatesthatsoftwarerequirementsbeconvertedintotestcasesbeforeanyimplementationbegins,andsuccessisdefinedbyaprogramthatpassesthesetests.Whilefeworganiza-tionsemployfulltest-drivendevelopment,integrationofnewcodeisusuallydependentoncreatingandpassingunittests.

Kulaletal.(2019)evaluatefunctionalcorrectnessusingthepass@kmetric,wherekcodesamplesaregeneratedperproblem,aproblemisconsideredsolvedifanysample

EvaluatingLargeLanguageModelsTrainedonCode

defincr_list(1:list):

"""Returnlistwithelementsincrementedby1.

>>>incr_list([1,2,3])

[2,3,4]

>>>incr_list([5,3,5,2,3,3,9,0,123])

[6,4,6,3,4,4,10,1,124]

return[i+1foriin1]

def

solution(lst):

"""Givenanon-emptylistofintegers,returnthesumofalloftheoddelementsthatareinevenpositions.

Examples

solution([5,8,7,1])=12

solution([3,3,3,3,3])=>9

solution([30,13,24,321])=→0

returnsum(lst[i]foriinrange(0,len(lst))ifi%2==0andlst[i]%2==1)

def

encode_cyclic(s:str):

returnsencodedstringbycyclinggroupsofthreecharacters.

def

#splitstringtogroups.Eachoflength3.

groups=[s[(3*i):min((3*i+3),len(s))]foriinrange((len(s)+2)//3)]#cycleelementsineachgroup.Unlessgrouphasfewerelementsthan3.

groups=[(group[1:]+group[0])iflen(group)==3elsegroupforgroupingroups]return"".join(groups)

decode_cyclic(s:str):

takesasinputstringencodedwithencode_cyclicfunction.Returnsdecodedstring.

#splitstringtogroups.Eachoflength3.

groups=[s[(3*i):min((3*i+3),len(s))]foriinrange((len(s)+2)//3)]

#cycleelementsineachgroup.

groups=[(group[-1]+group[:-1])iflen(group)==3elsegroupforgroupingroups]return"".join(groups)

Figure2.ThreeexampleproblemsfromtheHumanEvaldataset,wheretheprobabilitiesthatasinglesamplefromCodex-12Bpassesunittestsare0.9,0.17,and0.005.Thepromptprovidedtothemodelisshownwithawhitebackground,andasuccessfulmodel-generatedcompletionisshowninayellowbackground.Thoughnotaguaranteeforproblemnovelty,allproblemswerehand-writtenandnotprogrammaticallycopiedfromexistingsources.RandomproblemsandsamplescanbefoundinAppendixB.

passestheunittests,andthetotalfractionofproblemssolvedisreported.However,computingpass@kinthiswaycanhavehighvariance.Instead,toevaluatepass@k,wegeneraten≥ksamplespertask(inthispaper,weusen=200andk≤100),countthenumberofcorrectsamplesc≤nwhichpassunittests,andcalculatetheunbiasedestimator

(1)

Calculatingthisestimatordirectlyresultsinverylargenum-bersandnumericalinstability.InFigure3,weincludeanumericallystablenumpyimplementationthatsimplifiestheexpressionandevaluatestheproductterm-by-term.Onemaybetemptedtoestimatepass@kwith1-(1-p)kwherepistheempiricalestimateofpass@1,butweshowthatitisbiasedinAppendixA.

:paramn:totalnumberofsamples

:paramc:numberofcorrectsamples

:paramk:kinpass@$k$

ifn-c<k:return1.0

return1.0-d(1.0-k/

np.arange(n-c+1,n+1))

Figure3.Anumericallystablescriptforcalculatinganunbiasedestimateofpass@k.

Later,weprovideevidencethatBLEUscoremaynotbeareliableindicatoroffunctionalcorrectnessbyshowingthatfunctionallyinequivalentprogramsgeneratedbyourmodel(whichareguaranteedtodisagreewiththereferencesolutiononsomeinput)oftenhavehigherBLEUscoresthanfunctionallyequivalentones.

EvaluatingLargeLanguageModelsTrainedonCode

2.2.HumanEval:Hand-WrittenEvaluationSet

Weevaluatefunctionalcorrectnessonasetof164hand-writtenprogrammingproblems,whichwecalltheHu-manEvaldataset.Eachproblemincludesafunctionsig-nature,docstring,body,andseveralunittests,withanav-erageof7.7testsperproblem.Itisimportantforthesetaskstobehand-written,sinceourmodelsaretrainedonalargefractionofGitHub,whichalreadycontainssolutionstoproblemsfromavarietyofsources.Forexample,therearemorethantenpublicrepositoriescontainingsolutionstoCodeforcesproblems,whichmakeuppartoftherecentlyproposedAPPSdataset(Hendrycksetal.,2021).

ProgrammingtasksintheHumanEvaldatasetassesslan-guagecomprehension,reasoning,algorithms,andsimplemathematics.WereleasetheHumanEvaldatasetsothatotherscanevaluatefunctionalcorrectnessandmeasuretheproblem-solvingcapabilitiesoftheirmodels.Thedatasetcanbefoundat

/openai/human-eval.

2.3.SandboxforExecutingGeneratedPrograms

Sincepubliclyavailableprogramshaveunknownintentand

generatedprogramsareoftenincorrect

,

executing

theseprogramsposesasecurityrisk.Indeed,GitHubisknowntocontainmaliciousprogramsthatalterorchangetheirenvironments(Rokonetal.,2020).

Therefore,wedevelopedasandboxenvironmenttosafelyrununtrustedprogramsagainstunittests.Ourgoalsweretopreventtheseprogramsfrommodifying,gainingpersistenceon,accessingsensitiveresourceson,orexfiltratingdatafromahostornetwork.SinceOpenAI'straininginfrastructureisbuiltonKubernetesandcloudservices,wedesignedoursandboxtoaddressthelimitationsoftheseenvironmentswhileremainingidiomaticwiththeirpatternsofuse.

WeselectedthegVisorcontainerruntime(Lacasse,2018)asthemainhostprotectioncomponent.SincecontainerruntimeslikeDockercansharehostresourceswithcontain-ers,amaliciouscontainercouldpotentiallycompromiseahost.gVisorprotectsthehostbyemulatingitsresourcestointroduceasecurityboundarybetweenthehostanditscon-tainers.Network-adjacenthostsandservicesareprotectedbyeBPF-basedfirewallrulesthatpreventinboundandout-boundconnectionsexceptforthoserequiredforexperimentcontrol.

3.CodeFine-Tuning

Wefine-tuneGPTmodelscontainingupto12BparametersoncodetoproduceCodex.IncontrastwithGPT,Codexdisplaysnon-trivialperformanceontheHumanEvaldataset.Infact,CodexisabletosolvethemajorityoftheproblemsinHumanEvalifwegenerateandevaluate100samplesper

problem,andpickonethatpassesunittests.Whenlimitedtoabudgetofoneevaluationperproblem,producingmultiplesampleswithCodexandchoosingtheonewiththehighestmeanlog-probabilityprovidessignificantgains.

3.1.DataCollection

OurtrainingdatasetwascollectedinMay2020from54mil-lionpublicsoftwarerepositorieshostedonGitHub,contain-ing179GBofuniquePythonfilesunder1MB.Wefilteredoutfileswhichwerelikelyauto-generated,hadaveragelinelengthgreaterthan100,hadmaximumlinelengthgreaterthan1000,orcontainedasmallpercentageofalphanumericcharacters.Afterfiltering,ourfinaldatasettotaled159GB.

3.2.Methods

SinceCodexisevaluatedonnaturallanguageprompts,wehypothesizedthatitwouldbebeneficialtofine-tunefromtheGPT-3(Brownetal.,2020)modelfamily,whichalreadycontainsstrongnaturallanguagerepresentations.Surpris-ingly,wedidnotobserveimprovementswhenstartingfromapre-trainedlanguagemodel,possiblybecausethefine-tuningdatasetissolarge.Nevertheless,modelsfine-tunedfromGPTconvergemorequickly,soweapplythisstrategyforallsubsequentexperiments.

WetrainCodexusingthesamelearningrateasthecorre-spondingGPTmodel,witha175steplinearwarmupandcosinelearningratedecay.Wetrainforatotalof100billiontokens,usingtheAdamoptimizerwithβ₁=0.9,β₂=0.95,E=10-⁸,andaweightdecaycoefficientof0.1.

InordertomaximallyleveragetextrepresentationsfromGPT,webaseourcodelexerontheGPT-3texttokenizer.SincethedistributionofwordsinGitHubcodediffersfromthatofnaturaltext,thistokenizerisnotveryeffectiveforrepresentingcode.Thelargestsourceofinefficiencyarisesfromencodingwhitespace,soweaddanadditionalsetoftokensforrepresentingwhitespacerunsofdifferentlengths.Thisallowsustorepresentcodeusingapproximately30%fewertokens.

Tocomputepass@k,weassembleeachHumanEvalprob-lemintoapromptconsistingofaheader,asignature,andadocstring,whichisillustratedinFigure2.WesampletokensfromCodexuntilweencounteroneofthefollowingstopsequences:'\nclass',‘\ndef’,'\n#’,'\nif’,or

'\nprint’,sincethemodelwillcontinuegeneratingaddi-tionalfunctionsorstatementsotherwise.Weusenucleussampling(Holtzmanetal.,2020)withtopp=0.95forallsamplingevaluationinthiswork.

3.3.Results

InFigure4,weplottestlossonaheld-outvalidationsetagainstCodexmodelsize.Wefindthatjustaslanguage

EvaluatingLargeLanguageModelsTrainedonCode

esloss

CodexLossScaling

—(592+)-0.13

10°

6×10-1

10⁵10⁶10⁷10810⁹1010

2×10°

Non-embeddingparameters

Figure4.Modelcross-entropytestlossmeasuredonaheld-outsplitofourPythonGitHubcodecorpus.ThesmoothpowerlawscalingofperformancewithmodelsizeobservedinGPT-3appearstoholdevenaftercodefine-tuning.

modeltestlossfollowsapowerlawinmodelsize(Kaplanetal.,2020),testlossaftercodefine-tuningfollowsasimilar

powerlawwithfunctionalformwhereN

isthenumberofnon-embeddingparametersinthemodel.

Whenevaluatingpass@k,itisimportanttooptimizesam-plingtemperaturefortheparticularvalueofk.InFigure5,weplotpass@kagainstthenumberofsampleskandthesamplingtemperature.Wefindthathighertemperaturesareoptimalforlargerk,becausetheresultingsetofsampleshashigherdiversity,andthemetricrewardsonlywhetherthemodelgeneratesanycorrectsolution.

Inparticular,fora679Mparametermodel,theoptimaltem-peratureforpass@1isT*=0.2andtheoptimaltempera-tureforpass@100isT*=0.8.Withthesetemperatures,wefindthatpass@1andpass@100scalesmoothlyasafunctionofmodelsize(Figure6).

Pass@kcanalsobeinterpretedastheresultofevaluatingthebestoutofksamples,wherethebestsampleispickedbyanoraclewithpriorknowledgeoftheunittests.Fromapracticalperspective,wearealsointerestedintheset-tingwherewemustselectasinglesamplefromksampleswithouthavingaccesstoanoracle.Forinstance,whenthemodelisusedasanautocompletetoolwhereauserprovidesaprompt,wedonothaveunittests,butwouldliketoreturnonlyasinglecompletiontotheuserforevaluationsoastonotoverwhelmthem.

Inspiredbysimilarworkinlanguagemodeling,wefindthatchoosingthesamplewiththehighestmeantokenlogprobabilityoutperformsevaluatingarandomsample,whilechoosingthesamplebasedonsumlogprobabilitycanper-formslightlyworsethanpickingrandomly.Figure7demon-stratesthebenefitsofapplyingtheseheuristicstosamples(attemperature0.8)fromCodex-12B.

Pass@KvsK,Temperature

Pass@k

0.4

0.3

0.2

T=0.0T=0.2T=0.4T=0.6T=0.8T=1.0T=1.2

0.1

10°10¹10²

Numberofsamples(k)BestTemperaturevsK

Besttemperature

0.8

0.6

0.4

0.2

10⁰10¹10²

Numberofsamples(k)

Figure5.Inthetoppanel,weplotpass@kagainstthenumberofsamples(k)forvarioustemperaturesettings.Highertemperaturesarebetterwhenthenumberofsamplesislarge,likelyduetotheincreasedsamplediversity.Inthebottompanel,weplotthebesttemperaturesettingforeachk,obtainedbytakingtheupperhullofthetoppanel.

PassRatevsModelSize

Pass@k

pass@1(T*=0.2)

pass@100(T*=0.8)

0.5

0.2

0.1

0.0

10⁵10⁷10⁸10⁹1010

Non-embeddingparameters

Figure6.Usingtheoptimaltemperatures0.2and0.8forpass@1andpass@100,weplotthesetwometricsasafunctionofmodelsize.Performanceappearstoscalesmoothlyasasigmoidinlog-parameters.

EvaluatingLargeLanguageModelsTrainedonCode

Passrate

15

10

SampleRankingHeuristics

Oracle

Docstringbacktranslation

—Random

10°10¹10²

0.7

0.6

0.5

0.4

0.3

0.2

Numberofsamples(k)

Figure7.Modelperformanceinthesettingwherewecangeneratemultiplesamples,butonlyevaluateone.Wecandobetterthanran-domlyselectingasamplebychoosingthesolutionwiththehighestmeanlog-probability(red)orwiththehighestback-translationscore(orange)describedinSec.5.Thebluelinerepresentsthetheoreticalbestperformanceobtainedusinganoraclewithpriorknowledgeoftheunittests.

HumanEval/38

HumanEval/72

correct

wrong

0

0.00.10.20.3HumanEval/4

correct

wrong

10

0.0

20

15

0.2

0.4

HumanEval/21

correct

wrong

0

0.500.75

10.07.55.0

0.0

correct

wrong

0.00

0.25

BLEUscore

Figure8.BLEUscoreprobabilitydensitiesforcorrect(blue)andwrong(green)solutionsfromCodex-12Bfor4randomtasksfromHumanEval.Notethatthedistributionsarenotcleanlyseparable,suggestingthatoptimizingforBLEUscoreisnotequivalenttooptimizingforfunctionalcorrectness.

Finally,wecomputeBLEUscoresforallCodex-12BHu-manEvalsamples(attemperature0.8)againsttheirreferencesolutions.Foreachproblem,whenweplotthedistributionsofBLEUscoresforcorrectandincorrectsolutions,wenoticesignificantoverlap(Figure8).Sinceanincorrectsolutionisguaranteedtobefunctionallyinequivalenttothereferencesolution,weconcludethatimprovementsinBLEUscoremaynotindicateimprovedratesoffunctionalcorrectnessinpractice.

uatingattemperatures0.2,0.4,and0.8forGPT-Neo,andfromtemperatures0.2and0.8forGPT-J.DetailedresultsacrossmultiplemodelsizescanbefoundinTable1.

Finally,webenchmarkCodexagainstthelargestfreemodelfromTabnine,aleadingcodeautocompletesystem,whichachieves2.6%pass@1(atT=0.4)and7.6%pass@100(atT=0.8).ThisisroughlyequivalenttoCodex-12M,oneofthesmallestmodelsinoursuite.

3.4.ComparativeAnalysisofRelatedModelsandSystems

TworecentworkssimilarinspirittoCodexareGPT-Neo(Blacketal.,2021)andGPT-J(Wang&Komatsuzaki,2021),whicharetrainedonThePile(Gaoetal.,2020),adatasetcontainingtextfromavarietyofsourcesaswellas8%GitHubcode.ThebroaderresearchcommunityhasfoundthatthesemodelsoutperformexistingGPTsystemsinqualitativeprogrammingevaluations(Woolf,2021).

WeconfirmthesefindingsusingtheHumanEvaldataset,showingthatGPT-Neoachieves6.4%pass@1and21.3%pass@100,whileGPTmodelsofcomparablesizesachievenear0%onbothmetrics.Weseearemarkableprogressionincapabilities,withGPT-Neo-2.7BroughlyequivalenttoCodex-85M(30×fewerparameters).Similarly,GPT-J-6Bachieves11.6%pass@1and27.7%pass@100,whichisroughlyequivalenttoCodex-300M(20×fewerparameters).Passratesareobtainedbytakingthebestresultfromeval-

3.5.ResultsontheAPPSDataset

Recently,Hendrycksetal.(2021)introducedtheAPPSdatasettomeasurethecodingchallengecompetenceoflan-guagemodels.TheAPPSdatasetconsistsof5000trainingand5000testexamplesofcodingproblems,eachwithasetofunittestsand,forthetrainingdata,asetofcorrectsolu-tions.MostoftheAPPStestsproblemsarenotformulatedassingle-functionsynthesistasks,butratherasfull-programsynthesis,readinginputfromstdinandprintingoutputtostdout,incontrasttothemainCodextrainingdata.

InthepaperthatintroducesAPPS,theauthorsbenchmarkafewlanguagemodelsandreporttwometrics:thepercentageofproblemswherethemodelfindsacorrectsolution(calledthe“strictaccuracy”)andthepercentageofunittestspassed,evenifthesolutionisincorrect.Thelattermeasureisre-portedonlysoastoreducevarianceofthemeasurements,becausetheresultsonthefirstmetricweresolow.Weavoidthismetricandonlyfocuson“strictaccuracy”,and-asin

EvaluatingLargeLanguageModelsTrainedonCode

Table1.Codex,GPT-Neo,&TabNineevaluationsforHumanEval.WefindthatGPT-Jpass@1isbetweenCodex-85MandCodex-300Mperformance.

k=1

PASS@kk=10

k=100

GPT-NEO125M

0.75%

1.88%

2.97%

GPT-NEO1.3B

4.79%

7.47%

16.30%

GPT-NEO2.7B

6.41%

11.27%

21.37%

GPT-J6B

11.62%

15.74%

27.74%

TABNINE

2.58%

4.35%

7.59%

CODEX-12M

2.00%

3.62%

8.58%

CODEX-25M

3.21%

7.1%

12.89%

CODEX-42M

5.06%

8.8%

15.55%

CODEX-85M

8.22%

12.81%

22.4%

CoDEX-300M

13.17%

20.37%

36.27%

CoDEX-679M

16.22%

25.7%

40.95%

CoDEX-2.5B

21.36%

35.42%

59.5%

CoDEX-12B

28.81%

46.81%

72.31%

theprevioussections-wereportpass@knumbersforvari-ousk(Table2).Thereare2additionalfactors,well-knownfromcodingcompetitions,thatwetakeintoaccount:

·IncodingcompetitionsandintheAPPSdatasets,tasksareprovidedwith3input/outputexamplesincludedinthetaskdescription.Weutilizethisbysampling1000solutionsfromthemodelandfilteringoutonlythosethatpassthese3unittests(ifsuchsolutionsexist).Wethen

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论