版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
EvaluatingLargeLanguageModelsTrainedonCode
MarkChen*¹JerryTworek*¹HeewooJun*¹QimingYuan*¹HenriquePondedeOliveiraPinto*1
JaredKaplan*²HarriEdwards¹YuriBurda¹NicholasJoseph²GregBrockman¹AlexRay¹RaulPuri¹
GretchenKrueger¹MichaelPetrov¹HeidyKhlaaf³GirishSastry¹PamelaMishkin¹BrookeChan¹
ScottGray¹NickRyder¹MikhailPavlov¹AletheaPower¹LukaszKaiser¹MohammadBavarian¹
ClemensWinter¹PhilippeTillet¹FelipePetroskiSuch¹DaveCummings¹MatthiasPlappert¹
FotiosChantzis¹ElizabethBarnes¹ArielHerbert-Voss¹WilliamHebgenGuss¹AlexNichol¹AlexPaino¹
2021
NikolasTezak¹JieTang¹IgorBabuschkin¹SuchirBalaji¹ShantanuJain¹WilliamSaunders¹
ChristopherHesse¹AndrewN.Carr¹JanLeike¹JoshAchiam¹VedantMisra¹EvanMorikawa¹
AlecRadford¹MatthewKnight¹MilesBrundage¹MiraMurati¹KatieMayer¹PeterWelinder¹
BobMcGrew¹DarioAmodei²SamMcCandlish²IlyaSutskever¹WojciechZaremba¹
arXiv:2107.03374v2[cs.LG]14Jul
Abstract
WeintroduceCodex,aGPTlanguagemodelfine-tunedonpubliclyavailablecodefromGitHub,andstudyitsPythoncode-writingcapabilities.AdistinctproductionversionofCodexpowersGitHubCopilot.OnHumanEval,anewevalua-tionsetwereleasetomeasurefunctionalcorrect-nessforsynthesizingprogramsfromdocstrings,ourmodelsolves28.8%oftheproblems,whileGPT-3solves0%andGPT-Jsolves11.4%.Fur-thermore,wefindthatrepeatedsamplingfromthemodelisasurprisinglyeffectivestrategyforpro-ducingworkingsolutionstodifficultprompts.Us-ingthismethod,wesolve70.2%ofourproblemswith100samplesperproblem.Carefulinvestiga-tionofourmodelrevealsitslimitations,includingdifficultywithdocstringsdescribinglongchainsofoperationsandwithbindingoperationstovari-ables.Finally,wediscussthepotentialbroaderimpactsofdeployingpowerfulcodegenerationtechnologies,coveringsafety,security,andeco-nomics.
1.Introduction
Scalablesequencepredictionmodels(Graves,2014;
Vaswanietal.,2017;Childetal.,2019)havebecomeageneral-purposemethodforgenerationandrepresentationlearninginmanydomains,includingnaturallanguagepro-cessing(Mikolovetal.,2013;Sutskeveretal.,2014;Dai&Le,2015;Petersetal.,2018;Radfordetal.,2018;Devlinetal.,2018),computervision(VanOordetal.,2016;Menick&Kalchbrenner,2018;Chenetal.,2020;Baoetal,2021),audioandspeechprocessing(Oordetal.,2016;2018;Dhari-waletal.,2020;Baevskietal.,2020),biology(Alleyetal.,2019;Rivesetal.,2021),andevenacrossmultiplemodali-ties(Dasetal.,2017;Luetal.,2019;Rameshetal.,2021;Zellersetal.,2021).Morerecently,languagemodelshavealsofueledprogresstowardsthelongstandingchallengeofprogramsynthesis(Simon,1963;Manna&Waldinger,
1971),spurredbythepresenceofcodeinlargedatasets(Husainetal.,2019;Gaoetal.,2020)andtheresultingpro-grammingcapabilitiesoflanguagemodelstrainedonthesedatasets(Wang&Komatsuzaki,2021).Popularlanguagemodelingobjectiveslikemaskedlanguagemodeling(Devlinetal.,2018)andspanprediction(Raffeletal.,2020)havealsobeenadaptedtotraintheirprogrammingcounterpartsCodeBERT(Fengetal.,2020)andPyMT5(Clementetal.,2020).
Equalcontribution
OpenAI,SanFrancisco,California,USA.
²AnthropicAI,SanFrancisco,California,USA.Workper-formedwhileatOpenAI.
³Zipline,SouthSanFrancisco,California,USA.Workper-formedwhileatOpenAI.
Correspondenceto:MarkChen<mark@>,JerryTworek<jt@>,HeewooJun<hee-woo@>,QimingYuan<qiming@>.
Similarly,ourearlyinvestigationofGPT-3(Brownetal.,2020)revealedthatitcouldgeneratesimpleprogramsfromPythondocstrings.Whilerudimentary,thiscapabilitywasexcitingbecauseGPT-3wasnotexplicitlytrainedforcodegeneration.Giventheconsiderablesuccessoflargelan-guagemodelsinothermodalitiesandtheabundanceofpubliclyavailablecode,wehypothesizedthataspecializedGPTmodel,calledCodex,couldexcelatavarietyofcodingtasks.ThispaperdescribesseveralearlyCodexmodels,whosedescendantspowerGitHubCopilotandtheCodexmodelsintheOpenAIAPI.
EvaluatingLargeLanguageModelsTrainedonCode
CodexandCodex-SPerformance
PasSrate
GPT-3pass@1
Codexpass@1
一Codex-Spass@1
—Codex-SmeanlogprerankingCodex-Soraclereranking
0.2
0.0
10⁵10⁶10⁷10⁸10⁹1010
0.8
0.6
0.4
Non-embeddingparameters
Figure1.PassratesofourmodelsontheHumanEvaldatasetasafunctionofmodelsize.Whenasinglesampleisgeneratedforeachproblem,GPT-12Bsolvesnoproblems,butCodex(fine-tunedoncode)solves28.8%oftheproblems,andCodex-S(furtherfine-tunedoncorrectlyimplementedstandalonefunctions)solves37.7%oftheproblems.Fromhere,furthergainscanberealizedbygenerating100samplesperproblemandselectingthesamplewiththehighestmeanlog-probability(44.5%solved)orbyselectingthesamplethatpassestheunittests(77.5%solved).Allsamplesaregeneratedwithtemperature0.8.
Inthiswork,wefocusonthetaskofgeneratingstan-dalonePythonfunctionsfromdocstrings,andevaluatethecorrectnessofcodesamplesautomaticallythroughunittests.Thisisincontrasttonaturallanguagegeneration,wheresamplesaretypicallyevaluatedbyheuristicsorbyhumanevaluators.Toaccuratelybenchmarkourmodel,wecreateadatasetof164originalprogrammingproblemswithunittests.Theseproblemsassesslanguagecompre-hension,algorithms,andsimplemathematics,withsomecomparabletosimplesoftwareinterviewquestions.Wereleasethisdataalongwithanevaluationframeworkat
/openai/human-eval.
Tosolveaprobleminourtestset,wegeneratemultiplesamplesfromthemodels,andcheckifanyofthempasstheunittests.Withjustasinglesample,a12BparameterCodexsolves28.8%oftheseproblems,anda300MparameterCodexsolves13.2%oftheseproblems.Incontrast,the6BparameterGPT-J(Wang&Komatsuzaki,2021)achieves
11.4%onthesamedataset,whileallGPTmodelsachievenear0%.Toimproveourmodel'sperformanceatthetaskoffunctionsynthesisfromdocstrings,wefine-tuneCodexonstandalone,correctlyimplementedfunctions.Theresultingmodel,Codex-S,solves37.7%ofproblemswithasinglesample.Figure2showcasesproblemsofvaryingdifficultyinourdataset,alongwithcorrectmodelgeneratedsolutions.
Real-worldprogrammingtasksofteninvolveiterationsofapproachesandbugfixes,whichisapproximatedbygener-atingmanysamplesfromourmodelsandselectingonethatpassesallunittests.Within100samples,Codex-Sisableto
generateatleastonecorrectfunctionfor77.5%oftheprob-lems.Thisresultsuggeststhataccuratecodesamplescanbeselectedviaheuristicrankinginsteadoffullyevaluatingeachsample,thelatterofwhichmaynotbepossibleorprac-ticalindeployment.Indeed,wefindthatthesamplewithhighestmeanlog-probabilitypassesunittestsfor44.5%oftheproblems.
WeconcludebydiscussingthelimitationsandpotentialbroaderimpactsoftheseCodexmodelsandofincreasinglypowerfulcodegeneratingmodelsmoregenerally.
2.EvaluationFramework
Inthissection,wediscussthedetailsofourevaluationframework.Webeginbydefiningthepass@kmetric,andexplainitsadvantagesoverstandardmatch-basedmetrics.Next,wedescribethedatasetofhand-writtenproblems,called"HumanEval,"whichwecreatedinordertobench-markourmodels.Finally,wediscussthesandboxenviron-mentweusedtosafelyexecutemodel-generatedcode.
2.1.FunctionalCorrectness
Generativemodelsforcodearepredominantlybenchmarkedbymatchingsamplesagainstareferencesolution,wherethematchcanbeexactorfuzzy(asinBLEUscore).How-ever,recentworkhassurfaceddeficienciesinmatch-basedmetricsforcode.Forinstance,Renetal.(2020)findsthatBLEUhasproblemscapturingsemanticfeaturesspecifictocode,andsuggestsseveralsemanticmodificationstothe
score.
Morefundamentally,match-basedmetricsareunabletoac-countforthelargeandcomplexspaceofprogramsfunction-allyequivalenttoareferencesolution.Asaconsequence,recentworksinunsupervisedcodetranslation(Lachauxetal.,2020)andpseudocode-to-codetranslation(Kulaletal.,2019)haveturnedtofunctionalcorrectnessinstead,whereasampleisconsideredcorrectifitpassesasetofunittests.Wearguethatthismetricshouldbeappliedtodocstring-conditionalcodegenerationaswell.
Perhapsthemostconvincingreasontoevaluatefunctionalcorrectnessisthatitisusedbyhumandeveloperstojudgecode.Aframeworkknownastest-drivendevelopmentdic-tatesthatsoftwarerequirementsbeconvertedintotestcasesbeforeanyimplementationbegins,andsuccessisdefinedbyaprogramthatpassesthesetests.Whilefeworganiza-tionsemployfulltest-drivendevelopment,integrationofnewcodeisusuallydependentoncreatingandpassingunittests.
Kulaletal.(2019)evaluatefunctionalcorrectnessusingthepass@kmetric,wherekcodesamplesaregeneratedperproblem,aproblemisconsideredsolvedifanysample
EvaluatingLargeLanguageModelsTrainedonCode
defincr_list(1:list):
"""Returnlistwithelementsincrementedby1.
>>>incr_list([1,2,3])
[2,3,4]
>>>incr_list([5,3,5,2,3,3,9,0,123])
[6,4,6,3,4,4,10,1,124]
return[i+1foriin1]
def
solution(lst):
"""Givenanon-emptylistofintegers,returnthesumofalloftheoddelementsthatareinevenpositions.
Examples
solution([5,8,7,1])=12
solution([3,3,3,3,3])=>9
solution([30,13,24,321])=→0
returnsum(lst[i]foriinrange(0,len(lst))ifi%2==0andlst[i]%2==1)
def
encode_cyclic(s:str):
returnsencodedstringbycyclinggroupsofthreecharacters.
def
#splitstringtogroups.Eachoflength3.
groups=[s[(3*i):min((3*i+3),len(s))]foriinrange((len(s)+2)//3)]#cycleelementsineachgroup.Unlessgrouphasfewerelementsthan3.
groups=[(group[1:]+group[0])iflen(group)==3elsegroupforgroupingroups]return"".join(groups)
decode_cyclic(s:str):
takesasinputstringencodedwithencode_cyclicfunction.Returnsdecodedstring.
#splitstringtogroups.Eachoflength3.
groups=[s[(3*i):min((3*i+3),len(s))]foriinrange((len(s)+2)//3)]
#cycleelementsineachgroup.
groups=[(group[-1]+group[:-1])iflen(group)==3elsegroupforgroupingroups]return"".join(groups)
Figure2.ThreeexampleproblemsfromtheHumanEvaldataset,wheretheprobabilitiesthatasinglesamplefromCodex-12Bpassesunittestsare0.9,0.17,and0.005.Thepromptprovidedtothemodelisshownwithawhitebackground,andasuccessfulmodel-generatedcompletionisshowninayellowbackground.Thoughnotaguaranteeforproblemnovelty,allproblemswerehand-writtenandnotprogrammaticallycopiedfromexistingsources.RandomproblemsandsamplescanbefoundinAppendixB.
passestheunittests,andthetotalfractionofproblemssolvedisreported.However,computingpass@kinthiswaycanhavehighvariance.Instead,toevaluatepass@k,wegeneraten≥ksamplespertask(inthispaper,weusen=200andk≤100),countthenumberofcorrectsamplesc≤nwhichpassunittests,andcalculatetheunbiasedestimator
(1)
Calculatingthisestimatordirectlyresultsinverylargenum-bersandnumericalinstability.InFigure3,weincludeanumericallystablenumpyimplementationthatsimplifiestheexpressionandevaluatestheproductterm-by-term.Onemaybetemptedtoestimatepass@kwith1-(1-p)kwherepistheempiricalestimateofpass@1,butweshowthatitisbiasedinAppendixA.
:paramn:totalnumberofsamples
:paramc:numberofcorrectsamples
:paramk:kinpass@$k$
ifn-c<k:return1.0
return1.0-d(1.0-k/
np.arange(n-c+1,n+1))
Figure3.Anumericallystablescriptforcalculatinganunbiasedestimateofpass@k.
Later,weprovideevidencethatBLEUscoremaynotbeareliableindicatoroffunctionalcorrectnessbyshowingthatfunctionallyinequivalentprogramsgeneratedbyourmodel(whichareguaranteedtodisagreewiththereferencesolutiononsomeinput)oftenhavehigherBLEUscoresthanfunctionallyequivalentones.
EvaluatingLargeLanguageModelsTrainedonCode
2.2.HumanEval:Hand-WrittenEvaluationSet
Weevaluatefunctionalcorrectnessonasetof164hand-writtenprogrammingproblems,whichwecalltheHu-manEvaldataset.Eachproblemincludesafunctionsig-nature,docstring,body,andseveralunittests,withanav-erageof7.7testsperproblem.Itisimportantforthesetaskstobehand-written,sinceourmodelsaretrainedonalargefractionofGitHub,whichalreadycontainssolutionstoproblemsfromavarietyofsources.Forexample,therearemorethantenpublicrepositoriescontainingsolutionstoCodeforcesproblems,whichmakeuppartoftherecentlyproposedAPPSdataset(Hendrycksetal.,2021).
ProgrammingtasksintheHumanEvaldatasetassesslan-guagecomprehension,reasoning,algorithms,andsimplemathematics.WereleasetheHumanEvaldatasetsothatotherscanevaluatefunctionalcorrectnessandmeasuretheproblem-solvingcapabilitiesoftheirmodels.Thedatasetcanbefoundat
/openai/human-eval.
2.3.SandboxforExecutingGeneratedPrograms
Sincepubliclyavailableprogramshaveunknownintentand
generatedprogramsareoftenincorrect
,
executing
theseprogramsposesasecurityrisk.Indeed,GitHubisknowntocontainmaliciousprogramsthatalterorchangetheirenvironments(Rokonetal.,2020).
Therefore,wedevelopedasandboxenvironmenttosafelyrununtrustedprogramsagainstunittests.Ourgoalsweretopreventtheseprogramsfrommodifying,gainingpersistenceon,accessingsensitiveresourceson,orexfiltratingdatafromahostornetwork.SinceOpenAI'straininginfrastructureisbuiltonKubernetesandcloudservices,wedesignedoursandboxtoaddressthelimitationsoftheseenvironmentswhileremainingidiomaticwiththeirpatternsofuse.
WeselectedthegVisorcontainerruntime(Lacasse,2018)asthemainhostprotectioncomponent.SincecontainerruntimeslikeDockercansharehostresourceswithcontain-ers,amaliciouscontainercouldpotentiallycompromiseahost.gVisorprotectsthehostbyemulatingitsresourcestointroduceasecurityboundarybetweenthehostanditscon-tainers.Network-adjacenthostsandservicesareprotectedbyeBPF-basedfirewallrulesthatpreventinboundandout-boundconnectionsexceptforthoserequiredforexperimentcontrol.
3.CodeFine-Tuning
Wefine-tuneGPTmodelscontainingupto12BparametersoncodetoproduceCodex.IncontrastwithGPT,Codexdisplaysnon-trivialperformanceontheHumanEvaldataset.Infact,CodexisabletosolvethemajorityoftheproblemsinHumanEvalifwegenerateandevaluate100samplesper
problem,andpickonethatpassesunittests.Whenlimitedtoabudgetofoneevaluationperproblem,producingmultiplesampleswithCodexandchoosingtheonewiththehighestmeanlog-probabilityprovidessignificantgains.
3.1.DataCollection
OurtrainingdatasetwascollectedinMay2020from54mil-lionpublicsoftwarerepositorieshostedonGitHub,contain-ing179GBofuniquePythonfilesunder1MB.Wefilteredoutfileswhichwerelikelyauto-generated,hadaveragelinelengthgreaterthan100,hadmaximumlinelengthgreaterthan1000,orcontainedasmallpercentageofalphanumericcharacters.Afterfiltering,ourfinaldatasettotaled159GB.
3.2.Methods
SinceCodexisevaluatedonnaturallanguageprompts,wehypothesizedthatitwouldbebeneficialtofine-tunefromtheGPT-3(Brownetal.,2020)modelfamily,whichalreadycontainsstrongnaturallanguagerepresentations.Surpris-ingly,wedidnotobserveimprovementswhenstartingfromapre-trainedlanguagemodel,possiblybecausethefine-tuningdatasetissolarge.Nevertheless,modelsfine-tunedfromGPTconvergemorequickly,soweapplythisstrategyforallsubsequentexperiments.
WetrainCodexusingthesamelearningrateasthecorre-spondingGPTmodel,witha175steplinearwarmupandcosinelearningratedecay.Wetrainforatotalof100billiontokens,usingtheAdamoptimizerwithβ₁=0.9,β₂=0.95,E=10-⁸,andaweightdecaycoefficientof0.1.
InordertomaximallyleveragetextrepresentationsfromGPT,webaseourcodelexerontheGPT-3texttokenizer.SincethedistributionofwordsinGitHubcodediffersfromthatofnaturaltext,thistokenizerisnotveryeffectiveforrepresentingcode.Thelargestsourceofinefficiencyarisesfromencodingwhitespace,soweaddanadditionalsetoftokensforrepresentingwhitespacerunsofdifferentlengths.Thisallowsustorepresentcodeusingapproximately30%fewertokens.
Tocomputepass@k,weassembleeachHumanEvalprob-lemintoapromptconsistingofaheader,asignature,andadocstring,whichisillustratedinFigure2.WesampletokensfromCodexuntilweencounteroneofthefollowingstopsequences:'\nclass',‘\ndef’,'\n#’,'\nif’,or
'\nprint’,sincethemodelwillcontinuegeneratingaddi-tionalfunctionsorstatementsotherwise.Weusenucleussampling(Holtzmanetal.,2020)withtopp=0.95forallsamplingevaluationinthiswork.
3.3.Results
InFigure4,weplottestlossonaheld-outvalidationsetagainstCodexmodelsize.Wefindthatjustaslanguage
EvaluatingLargeLanguageModelsTrainedonCode
esloss
CodexLossScaling
—(592+)-0.13
10°
6×10-1
10⁵10⁶10⁷10810⁹1010
2×10°
Non-embeddingparameters
Figure4.Modelcross-entropytestlossmeasuredonaheld-outsplitofourPythonGitHubcodecorpus.ThesmoothpowerlawscalingofperformancewithmodelsizeobservedinGPT-3appearstoholdevenaftercodefine-tuning.
modeltestlossfollowsapowerlawinmodelsize(Kaplanetal.,2020),testlossaftercodefine-tuningfollowsasimilar
powerlawwithfunctionalformwhereN
isthenumberofnon-embeddingparametersinthemodel.
Whenevaluatingpass@k,itisimportanttooptimizesam-plingtemperaturefortheparticularvalueofk.InFigure5,weplotpass@kagainstthenumberofsampleskandthesamplingtemperature.Wefindthathighertemperaturesareoptimalforlargerk,becausetheresultingsetofsampleshashigherdiversity,andthemetricrewardsonlywhetherthemodelgeneratesanycorrectsolution.
Inparticular,fora679Mparametermodel,theoptimaltem-peratureforpass@1isT*=0.2andtheoptimaltempera-tureforpass@100isT*=0.8.Withthesetemperatures,wefindthatpass@1andpass@100scalesmoothlyasafunctionofmodelsize(Figure6).
Pass@kcanalsobeinterpretedastheresultofevaluatingthebestoutofksamples,wherethebestsampleispickedbyanoraclewithpriorknowledgeoftheunittests.Fromapracticalperspective,wearealsointerestedintheset-tingwherewemustselectasinglesamplefromksampleswithouthavingaccesstoanoracle.Forinstance,whenthemodelisusedasanautocompletetoolwhereauserprovidesaprompt,wedonothaveunittests,butwouldliketoreturnonlyasinglecompletiontotheuserforevaluationsoastonotoverwhelmthem.
Inspiredbysimilarworkinlanguagemodeling,wefindthatchoosingthesamplewiththehighestmeantokenlogprobabilityoutperformsevaluatingarandomsample,whilechoosingthesamplebasedonsumlogprobabilitycanper-formslightlyworsethanpickingrandomly.Figure7demon-stratesthebenefitsofapplyingtheseheuristicstosamples(attemperature0.8)fromCodex-12B.
Pass@KvsK,Temperature
Pass@k
0.4
0.3
0.2
T=0.0T=0.2T=0.4T=0.6T=0.8T=1.0T=1.2
0.1
10°10¹10²
Numberofsamples(k)BestTemperaturevsK
Besttemperature
0.8
0.6
0.4
0.2
10⁰10¹10²
Numberofsamples(k)
Figure5.Inthetoppanel,weplotpass@kagainstthenumberofsamples(k)forvarioustemperaturesettings.Highertemperaturesarebetterwhenthenumberofsamplesislarge,likelyduetotheincreasedsamplediversity.Inthebottompanel,weplotthebesttemperaturesettingforeachk,obtainedbytakingtheupperhullofthetoppanel.
PassRatevsModelSize
Pass@k
pass@1(T*=0.2)
pass@100(T*=0.8)
0.5
0.2
0.1
0.0
10⁵10⁷10⁸10⁹1010
Non-embeddingparameters
Figure6.Usingtheoptimaltemperatures0.2and0.8forpass@1andpass@100,weplotthesetwometricsasafunctionofmodelsize.Performanceappearstoscalesmoothlyasasigmoidinlog-parameters.
EvaluatingLargeLanguageModelsTrainedonCode
Passrate
15
10
SampleRankingHeuristics
Oracle
Docstringbacktranslation
—Random
10°10¹10²
0.7
0.6
0.5
0.4
0.3
0.2
Numberofsamples(k)
Figure7.Modelperformanceinthesettingwherewecangeneratemultiplesamples,butonlyevaluateone.Wecandobetterthanran-domlyselectingasamplebychoosingthesolutionwiththehighestmeanlog-probability(red)orwiththehighestback-translationscore(orange)describedinSec.5.Thebluelinerepresentsthetheoreticalbestperformanceobtainedusinganoraclewithpriorknowledgeoftheunittests.
HumanEval/38
HumanEval/72
correct
wrong
0
0.00.10.20.3HumanEval/4
correct
wrong
10
0.0
20
15
0.2
0.4
HumanEval/21
correct
wrong
0
0.500.75
10.07.55.0
0.0
correct
wrong
0.00
0.25
BLEUscore
Figure8.BLEUscoreprobabilitydensitiesforcorrect(blue)andwrong(green)solutionsfromCodex-12Bfor4randomtasksfromHumanEval.Notethatthedistributionsarenotcleanlyseparable,suggestingthatoptimizingforBLEUscoreisnotequivalenttooptimizingforfunctionalcorrectness.
Finally,wecomputeBLEUscoresforallCodex-12BHu-manEvalsamples(attemperature0.8)againsttheirreferencesolutions.Foreachproblem,whenweplotthedistributionsofBLEUscoresforcorrectandincorrectsolutions,wenoticesignificantoverlap(Figure8).Sinceanincorrectsolutionisguaranteedtobefunctionallyinequivalenttothereferencesolution,weconcludethatimprovementsinBLEUscoremaynotindicateimprovedratesoffunctionalcorrectnessinpractice.
uatingattemperatures0.2,0.4,and0.8forGPT-Neo,andfromtemperatures0.2and0.8forGPT-J.DetailedresultsacrossmultiplemodelsizescanbefoundinTable1.
Finally,webenchmarkCodexagainstthelargestfreemodelfromTabnine,aleadingcodeautocompletesystem,whichachieves2.6%pass@1(atT=0.4)and7.6%pass@100(atT=0.8).ThisisroughlyequivalenttoCodex-12M,oneofthesmallestmodelsinoursuite.
3.4.ComparativeAnalysisofRelatedModelsandSystems
TworecentworkssimilarinspirittoCodexareGPT-Neo(Blacketal.,2021)andGPT-J(Wang&Komatsuzaki,2021),whicharetrainedonThePile(Gaoetal.,2020),adatasetcontainingtextfromavarietyofsourcesaswellas8%GitHubcode.ThebroaderresearchcommunityhasfoundthatthesemodelsoutperformexistingGPTsystemsinqualitativeprogrammingevaluations(Woolf,2021).
WeconfirmthesefindingsusingtheHumanEvaldataset,showingthatGPT-Neoachieves6.4%pass@1and21.3%pass@100,whileGPTmodelsofcomparablesizesachievenear0%onbothmetrics.Weseearemarkableprogressionincapabilities,withGPT-Neo-2.7BroughlyequivalenttoCodex-85M(30×fewerparameters).Similarly,GPT-J-6Bachieves11.6%pass@1and27.7%pass@100,whichisroughlyequivalenttoCodex-300M(20×fewerparameters).Passratesareobtainedbytakingthebestresultfromeval-
3.5.ResultsontheAPPSDataset
Recently,Hendrycksetal.(2021)introducedtheAPPSdatasettomeasurethecodingchallengecompetenceoflan-guagemodels.TheAPPSdatasetconsistsof5000trainingand5000testexamplesofcodingproblems,eachwithasetofunittestsand,forthetrainingdata,asetofcorrectsolu-tions.MostoftheAPPStestsproblemsarenotformulatedassingle-functionsynthesistasks,butratherasfull-programsynthesis,readinginputfromstdinandprintingoutputtostdout,incontrasttothemainCodextrainingdata.
InthepaperthatintroducesAPPS,theauthorsbenchmarkafewlanguagemodelsandreporttwometrics:thepercentageofproblemswherethemodelfindsacorrectsolution(calledthe“strictaccuracy”)andthepercentageofunittestspassed,evenifthesolutionisincorrect.Thelattermeasureisre-portedonlysoastoreducevarianceofthemeasurements,becausetheresultsonthefirstmetricweresolow.Weavoidthismetricandonlyfocuson“strictaccuracy”,and-asin
EvaluatingLargeLanguageModelsTrainedonCode
Table1.Codex,GPT-Neo,&TabNineevaluationsforHumanEval.WefindthatGPT-Jpass@1isbetweenCodex-85MandCodex-300Mperformance.
k=1
PASS@kk=10
k=100
GPT-NEO125M
0.75%
1.88%
2.97%
GPT-NEO1.3B
4.79%
7.47%
16.30%
GPT-NEO2.7B
6.41%
11.27%
21.37%
GPT-J6B
11.62%
15.74%
27.74%
TABNINE
2.58%
4.35%
7.59%
CODEX-12M
2.00%
3.62%
8.58%
CODEX-25M
3.21%
7.1%
12.89%
CODEX-42M
5.06%
8.8%
15.55%
CODEX-85M
8.22%
12.81%
22.4%
CoDEX-300M
13.17%
20.37%
36.27%
CoDEX-679M
16.22%
25.7%
40.95%
CoDEX-2.5B
21.36%
35.42%
59.5%
CoDEX-12B
28.81%
46.81%
72.31%
theprevioussections-wereportpass@knumbersforvari-ousk(Table2).Thereare2additionalfactors,well-knownfromcodingcompetitions,thatwetakeintoaccount:
·IncodingcompetitionsandintheAPPSdatasets,tasksareprovidedwith3input/outputexamplesincludedinthetaskdescription.Weutilizethisbysampling1000solutionsfromthemodelandfilteringoutonlythosethatpassthese3unittests(ifsuchsolutionsexist).Wethen
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 《中国高血压防治指南》解读
- 2026秋新教材湘科版小学科学五年级上册第四单元《力与运动》分层作业(附答案)
- 2026小升初学生收心教育课件:收假收心逐梦前行
- 2026年羽毛球运动入门课件
- 2026 年秋季传染病预防多方协同宣讲
- 麻纺厂产品追溯管理制度
- 国家税收基础理论1章
- 医学院大学课件--胸部损伤
- 安信证券借壳上市案例广东中业矿业项目建议书
- 国家人力资源管理师人力资源规划
- 2026年国家电网职称考试(政工)中级题库
- 小学反诈骗工作制度
- 年加工400万吨选矿项⽬报告书
- 中国思想史马工程课件第秦汉篇
- 容诚事务所校园招聘笔试题a卷
- 2026年云南省事业单位行测真题及答案
- 2026江西三支一扶历年真题
- 新生儿和低体重新生儿的麻醉管理课件
- GB/T 20147.2-2026色度学第2部分:CIE标准照明体
- 中草药栽培技术专业介绍
- 水电水利工程覆盖层灌浆技术规范
评论
0/150
提交评论