外媒外刊-SemiAnalysis:开放模型是否正在迎头赶上?-Are Open Models Catching Up-20260821_第1页
外媒外刊-SemiAnalysis:开放模型是否正在迎头赶上?-Are Open Models Catching Up-20260821_第2页
外媒外刊-SemiAnalysis:开放模型是否正在迎头赶上?-Are Open Models Catching Up-20260821_第3页
外媒外刊-SemiAnalysis:开放模型是否正在迎头赶上?-Are Open Models Catching Up-20260821_第4页
外媒外刊-SemiAnalysis:开放模型是否正在迎头赶上?-Are Open Models Catching Up-20260821_第5页
已阅读5页,还剩15页未读 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

1/15

(1)AreOpenModelsCatchingUp?

/p/are-open-models-catching-up

EvanCloutier,MaxKan,JordanNanos,DylanPatelAugust21,2026

ThepasttwomonthshavebeenabreakoutperiodforopensourceAI.Yes,therewasthe"DeepSeekmoment"backinJanuary2025,butnooneactuallyusedR1todoany

economicallyvaluablework.Incontrast,modelslikeGLM5.3andKimiK3aregenuinelycapableofmanyofthesamecodingandagentictasksthatrocketedAnthropicto

$65B+ARR.UnlikeotherswhoinflatedARR,ourfiguresweremuchclosertoreality.

Anopenmodeleverytendays

Majoropen-weightsreleases,June-August2026

DeepSeekV4Flash

MiniMaxM³

MuseSpark1.2Qwen3.8-27B

Nemotron3Ultra

Source:SemiAnalysis

2/15

【价值目录】网整理:

Itisanexcitingtimetobeatokenconsumer.Competitionisheatingup,usageresetsarebeingdoledout,andthebattleforyourtokensnowextendsbeyondtheOpenAI-Anthropicduopoly.Fireworksaloneisprocessingover40Ttokensperday—2xtheOpenAIAPI’s

volumeattheendofMarch.

However,majorFUDhasalsoemergedasaresultofopenmodelsuccess:ifopenmodelsstaycapableenoughrelativetotheclosedfrontieratafractionofthecost,won'tthemodellayerbecomecommoditized?Thisoutcomewouldobviouslybedisastrousforfrontierlabmargins.ForfulldetailsonAnthropicandOpenAI’sfinancials,seeour

TokenomicsModel.

Toprojecthowtheopenvsclosedcapabilitygapwillprogressinthefuture,wefirstneedtomeasurethepast.Naively,youmightpickasinglesetofbenchmarkstomeasureall

historicalmodels,butthisisamistake.Everybenchmarkisaproductofaparticularera.Whensomeonecreatesanewbenchmark,theirgoalistodiscerndifferencesin

modelcapabilitiesatthetime.Ifthey’resuccessful,themodelmakerswillclimbsaid

benchmarkuntilitbecomessaturated.Oncethathappens,everyonestopscaringaboutthebenchmark,andthecyclerepeats.

TherehavebeenthreeerasthusfarinthehistoryofLLMs:earlyscaling,reasoning,andagentic.Eacherarepresentedastep-functionincreaseinmodelutility,andrather

thantryingtoplotasinglecontinuoustrend,webelieveit’sbettertoevaluatethemodelsandbenchmarksfromeacheraindividually.

Whenviewedthisway,itbecomesclearthattheopenvs.closedgapmovesincycles.Atthestartofeachera,afrontierlabcompletessomepromisingresearch,trainsan

impressivemodel,deploysitatscaletotheirusers,andjumpsahead.Then,otherlabs

identifythekeyadvances,reverse-engineerwhatthefrontierlabisdoing,replicatethemintheirownmodels,andclosethegap.Nothingstayssecretforever—especiallywhenyoufactorindistillation.It’sjustaquestionofhowlongittakes.

Toanswerthisquestion,wetookalltherelevantmodelsfromeacheraandranacuratedsetofbenchmarkstogetacompositecapabilityscore.Theresultisacleartrend:witheachgeneration,open-sourcemodelstakehalfaslongtocatchuptothefirst

closed-sourcemodeloftheera.

3/15

Catch-uptimehalveseveryera

Thegapeacheraopenedwith,andhowlongitsfoundingclosedflagshipstood-frompublicdebut-beforeanopenmodelsurpassedit

GAPWHENTHEERAOPENED·POINTSTIMETOCLOSEIT·MONTHS

Era1·Earlyscaling

35.8pts

Era2·Reasoning

Era3·Agentic

witlinechera.theerasbestreultcncechbenchmark=100.Cloekstatateachfoundlngfaghip'spubledobut(ChaGPTNov'22,01-provowSop'24.0pus4.5Now'25)

Source:SemiAnalysis

Ofcourse,benchmarksdon'ttellthefullstory,andwe'llhighlightalltherelevantcaveatsbelow.Finally,we'llextendthisanalysisintothefuture,andexplainwhyit'slessbearishfrontiermodelsthanyoumightinitiallythink.

Howwemeasured

Here'sanoverviewofthemodelsandbenchmarksweselectedforeachera:

4/15

Thefullslate

Twelvebenchmarksacrossthreeeras-eachera'smodelsscoredonlyontheteststhatdefneditEra1·EarlyScaling2022-2024Wordprobems,inglefunctlons,andmutplecholce-ntols,nharnesses

GSM⁸KHumanEval

Trvlaansweredclosed-bokfrommodel

TheUSMathChympladqualferIntegerShortfactualquestlonsenglneeredtopunish

Era3·Agentic2025-todayLong-horizonagentwork:aterminalaresearchcorpus,asupportdesk,andacodebase

Source:SemiAnalysis

PickingasingleSOTAclosedandopenmodelataparticulartimeissubjective,butour

selectionsreflectthegeneralconsensusamongAlexperts.Incaseswherethere'sdebate—e.g.Fable5vsGPT5.6today—wewereconservativeandtestedboth.

Forbenchmarks,wereliedonacombinationofpersonaltasteandpopularity.HumanitiesLastExam(HLE),forexample,isknowntohavelotsofissues,butwasalsotrulyoneofthedefiningbenchmarksofthereasoningerawithnoclosesubstitutes.SWE-benchPro,ontheotherhand,issimilarlypopularandproblematic,butalsocloselyapproximatedbyDeepSWE.

MostofthebenchmarkscoreshereweranourselvesusingPrimeIntellect'sevaluation

stack,specificallytheirenvironmentshubandtheevalsharnessincludedinPrime-RL.TherestcomefromrunsbyourfriendsatArtificialAnalysisandDatacurve'sDeepSWE

leaderboard.Openmodelswereservedthewaytheywouldhavebeenatrelease:vLLMversions,hardwarethatwasinuseatthetime,andsamplingsettingsfromthemodelcard.Forclosedmodels,weranagainsttheirpinnedAPIversions.Whereournumbersshareachartwiththird-partyvalues,wematchedtheirrulesets.

We'dliketogiveahugethankyoutoFlorianBrand(@xeophon)fromPrimeIntellect

forhelpinguspickbenchmarks/models,implementevals,andcheckforcorrectness.

Era1|Earlyscaling(2022-2024)

5/15

It'sJune2023.TheworldisreckoningwithChatGPT,andMarkZuckerbergjustagreedtofightElonMuskattheColosseum.ButwhileZuckistrainingjiu-jitsuanddoingMurphs,hiscompanyisdoingsometrainingoftheirown.FAIRisabouttopushpasttheMistralexodusandotherdrama,andsuccessfullyshipLlama-2-70B.Thefirstopenmodelthat

approachedthefrontier.

Howfarbehindthefrontierwasit?FourbenchmarkshelpeddefineSOTAatthetime:

GSM8K,HumanEval,TriviaQA,andMMLU-Pro:

TheEra1slate

Fourbenchmarksthatdefnedearlyscaling:wordproblems,singlefunctions,andmultiplechoice-noharnesses,notools

GSM8K

Grade-schoolwordproblemssolvedstepbystep.Notools,notricks:read,

HumanEval

Handwrittensinge-functionprogrammingtasks.Themodelwritesthebody;unittestsgradeit.

MMLU-Pro

Ten-optionmultiplechoiceacross14disciplines,lawtophysics.ThehardersuccessortothesaturatedMMLU.

Source:SemiAnalysis

Thesebenchmarksarerepresentativeofwhatthefrontiermodelswerecapableofatthetime.Simplemultiplechoicequestions,wordproblems,andprogrammingproblems

scopedtosinglefunctions.Howtimeshavechanged!HereishowLlama-2stackedupagainstGPT-3.5Turboinacagematchoftheirown:

6/15

GPT-3.5vs.Llama-2-70B

Theera'sfirstclosedflagshipagainstthefrstopenchallenger,atLlama-2'sJuly2023releasenormalizedscore,erabestperbenchmark=100

78.8

GSM8KHumanEvalMMLU-ProTriviaQA

Source:SemiAnalysis

Toaccountfordifferencesinbenchmarkdifficulty,wenormalizedthescores.Eachera'sbestresultissetto100,andeveryothermodelisscoredrelativetothat.Thecompositescorerepresentstheequal-weightaverageofthefour:75.7forGPT-3.5Turboonthe

frontier,and39.9forLlama-2-70B.Ameasured,butconsiderablegap.Thisinitiallag

createsthestorylinefortherestoftheera:Mixtral-8x7BreleasedinDecember2023

createdmomentumtowardsGPT-4capability,onlytohaveGPT-4TurboandGPT-4oraceahead:

Source:SemiAnalysis

7/15

IttookuntiltheLlama-3.1-405BreleaseinJuly2024foropenmodelstoclosetheGPT-3.5Turbogap,withacompositescoreof86.Thelastfrontiermodel,GPT-4o,wasmatchedincapabilitybyDeepSeekV3inDecember2024,scoring95.5and94.1respectively.

Qwen2.5-72BlandedwithinstrikingdistanceofGPT-4oatasixthofthe405Bparametercount,pre-trainedon18Ttokens.

Thisisthefirstinstanceofthegapclosing.Throughthisera,wedidn'tseethefrontierrisemuchbeyondthecapabilitiesofGPT-4,butthiswasmostlyduetopriorities:Turboand4owerebuilttomakeGPT-4cheaperandfaster,notsmarter.

Meanwhile,OpenAlhadbeenworkingtowardsadifferentkindofmodel.Theprocess-rewardpaperandNoamBrownhirebothpointedatreasoning,andbymid-2024every.majorlabwaspublishingtest-time-computeresearch.Sevenweeksafter405B,on

September122024,OpenAlshippedo1-preview:amodelthatsparkedaneweraofinnovation.

Era2|Reasoning(2024-2025)

01resetthechoiceofbenchmarks,alongwiththegap.TheelementaryevalsfromEra1werenolongerdifficultenoughtotesto¹'sfullabilities.GradeschoolmathproblemswerereplacedbytheAIME.ScaleAlcollectedsomeofthemostesoteric,PhD-levelmultiple

choicequestionsintheworldandprovocativelycalleditHumanity'sLastExam.

TheEra2slate

Fourbenchmarksthatdefnedreasoning:harderhomework,builttorewardthinkingoverrecall-stillnotools,noharnesses

GPQA-Diamond5CIENCE

Graduate-levelmultiplechoicewrittenbydomainexperts,builttobeGoogle-proof:skilednon-expertsgetthemwrongevenwiththirty

NYU,Cohere&Anthropileresearchers·No2023-198-questionDlamondsubset

SimpleQAVerified

Shortfactualquestionsengineeredtopunishconfdentguessing.ArebuiltsubsetofOpenA'sSimpleQAwiththenoisylabelsandredundancy

GoogleDeepMind·Sep2025-1,000prompts

BenchmarkprotocolsasrunintheSemlAnayseraevalualonssieseflecttheevalutedspls.

AIME2026

ThequalifebetweentheAMCandtheUSMathOlympiad.Thirty

problems,everyansweranintegerfromOto999,nopartialcredit.The2026paperreplacesthe2024and2025sets,whichhadleakedinto

MA,compledbyMathArenaETHZurch)-Feb2026-30problems

Humanity'sLastExam

Frontler-dificultyquestionscrowdsourcedfromroughly1,000subject-

ScaleA&CAIS·Jan2025·2,15Btext-onlyquestions

Source:SemiAnalysis

It'shardtooverstatehowimportantthereleaseofo1wasintechcircles.Manyconsiderittheday“weknewforsurewe'dgetAGI.”However,unlikeLlama-2-70BvsGPT-4,thegap

8/15

betweenopenvsclosedsourcestartedoffmuchsmallerduringEra2.Theculprit?AlittleknownmodelcalledDeepSeekR1.

ERAEVAL5·ERA2REASONING

DeepSeek-R1arrives

R1vs.olacrossthefourreasoning-erabenchmarks,January2025·normalizedscore,erabestperbenchmark=100

93.4

AIME2026HLE(notools)SimpleQAVerifed

oucesSemuscesoerymoPoioesourosescosomulsesecomsthesbetroutoehbecmok-10comouewswesedwohtmemsthnomsiodseosR-9820mnessdocummentedlength-stoprecoveryre-rurs;TrnityLargeThinkngdatedatitsFeb192026technkcal-reportpublcation.

Source:SemiAnalysis

A12.1pointgapvs35.8atthestartofthepreviousera.Themarketpukedinresponse.Fortunately,theAlcapextradequicklyrecovered,andthe"we'resoback"openmodelmomentumestablishedbyR1wassoonsquashedbyMeta'sLlama-4Maverick.

Era2:thegapnarrows

Bestcompositescoreavalableateachdate·equal-weightmeanofnormalizedAIME2026,GPQA-Diamond,HLE,andSimpleCAVerifed(erabestperbenchmark=

ouceSomonesetearymoPomohoaesocuonteucomsoomlodpobocdmuktheomsbotreoateoechbecmek-10ComotaedoghmensotonomznodseoA-G⁶270necadocumentedlength-toprecoveryre-run;:oldtedtthea-reiewobut(Sep122024ecefromtheol-202-12-17napahotLnetrackteruraingbetcompeate;dtmarmodeielozes

Source:SemiAnalysis

9/15

Gemini2.5Proando3continuedtopushthereasoningfrontier,andtheR1-0528

checkpointclosedtheinitialgapinMay2025withascoreof78.An8.5monthwindowtoclosea12.1pointgap:

ERAEVAL5·THECYCLE

Era2closedsooner

Thegapeacheraopenedwith,andhowlongitsfoundingclosedflagshipstood-frompublicdebut-beforeanopenmodelsurpassedit

Era1·Earlyscaling

35.8pts

Era2·Reasoning

oesomeseterymoPoioeseodurtoaeceneomubedcomeonopontsihehetoensboetrecenedhberesoak-1oCesnatechfonanfoadtppeledbuchacerNO

Source:SemiAnalysis

NotablyabsentfromthechartssofarisAnthropic.Theirmodelcardsreportedthese

benchmarkslikeeveryoneelse's,buttheyneverfoughtforthetopoftheleaderboardinthisera.WhileOpenAlandGoogletradedcrowns,AnthropicwasturningClaudeintothedefaultcodingagent.Thissetthetermsforthenextera:thebenchmarksthatnowmatterruninaterminal.

Era3|Agentic(2025-Today)

PriortoClaudeCode,agentshadtheirmoments(likeCognition'sviraldemoofDevinin

March2024),butAnthropicwasthefirsttonailamodel+harnessproduct—anditpaidoff.SincethegeneralreleaseofClaudeCodeinMay2025,Anthropichasaddednorthof$65BinARR.Forin-depthAnthropicandOpenAlARRprojections,seeourTokenomicsModel.

Withagentscameyetanothernewsetofbenchmarks.Fancymathproblemswereno

longerthebesttestofmodelcapabilities.Instead,peoplewantedtoknowhowwellmodelscouldwritecode,dowebsearch,andgenerallyuseacomputerlikeahuman.

10/15

TheEra3slate

Fourbenchmarksthatmeasureagents:aterminal,aresearchcorpus,asupportdesk,andacodebase

BrowseComp-Plus

830hard-to-fndquestionsansweredagainstafxedcorpusof100,195webdocuments.Multi-hopretrieval,andknowingwhentheevidenceactually

supportsananswer.

TooLsBM25searchoverthecorpus,top5documentsperquery

HARNEsSFixedsearchloopwitha100-turmncap,identicalforeverymodelTevatronproject·Aug2025·830questions

DeepSWE

113engineeringtasksacros⁵91activeopen-sourcerepos,writtenfrom

scratchandnevermergedupstream.Long-horizonworkinrealcodebases,gradedagainsthiddentests.

TooLsBashonlyInsdetherepo'scontalner

HARME⁵5minl-SWE-agentforeverymodel;patchesgradedlnapristheontalnerDatacurve·May2026·113tasks

BenchmarkprotocolsarunintheSemlAnslyeseraevsbluatlensanbyArtlfcalAnalyulst²-Barnkingv.0L¹:szereflectthecvalutedsets.

T³-Banking

97bankingsupporttaskswheretheagentworksbackendtoolswhiletalkingtoasimulatedcustomer.Gradedonthefinaldatabasestateandpolicy

compliance,nottheconversation.

Terminal-Bench2.1

Command-linetaskssolvedend-to-end,eachinitsownprebuiltDocker

container.Theagentworkstheshel;ahiddenveriferchecksthefinalstate,

TooLsAfllshellinsidethetaskcontainer

HARNESsEachmodelsonCUagent(ClaudeCode,Codex,KimiCode,pi);eleTerminus2

Sera'sT-benchlineage·97tasks-bankingdomaln

Source:SemiAnalysis

Terminal-Bench2.1,BrowseComp-Plus,T³-banking,andDeepSWEcoverthelong-horizonworkagentsareusedfortoday:softwareengineering,deepresearch,andknowledge

work.Wealsopickedbenchmarksthatskewnewerbydesigntolimitmemorization.

Theagenticerastartingline

Closedflagshipsvs.thebestopencodingmodeloftheDecember2025cohort·normalizedscore,erabestperbenchmark=100

58.1

37.6

Terminal-Bench2.1BrowseComp-PlusDeepSWE

Opus4.5andGPT-5.2areourfist-partyrursontheeariertaskset,whichsoreslower.Composlteistheequal-weightmeanofthefourbenchmarks.Scoresnormalzedperbenchmark:theera'sbestresutoneachbenchemark=100;composites

Source:SemiAnalysis

MostAlexpertsconsiderOpus4.5theofficialstartoftheagenticeraduetothereliabilityofthemodel.Interestingly,GPT-5.2(OpenAI'sflagshipatthetime)performedbetteronour

11/15

benchmarksuite,butthisdidn'tcorrespondtoabetteruserexperience.Thefullagentic

product(model+harness)wasnowwhatmattered,andAnthropichadbeenlaser-focusedoniteratingtowardsaharnessthatexcelledatgeneralagenticwork.Codex,incontrast,wascomparativelycrude,andOpenAlwassimultaneouslypursuingsidequestslikewebbrowsers.

Modelreleasesalsocompressedbetweenthetwofrontierlabs.OpenAlandAnthropic

createdtheirduopolybyreleasingamodelevery51daysonaveragethroughoutthisera.Comparedtothe213and120dayreleaseaveragesthroughoutEra1andEra2,

respectively,thisisamassivespeedup.

Theflagshiprefreshcyclecollapsed

Daysbetweensuccessivefrontier-labflagshipreleases,erabyera

■OpenAlAnthropic

ERA1·EARLYSCALINGavg213days

GPT-4→GPT-4Turbo

132dayso1→03

avg51days

GPT-5.2→GPT-5.3-Codex

GPT-5.4→GPT-5.5

SoucsbcommocomoihdAu7.202EaandzhowonoudmotugbopomeodudmtflrosoaD⁰5,2024AmoenoshmoomoumoeomoiosS3robsmounstous4aEa

3labaveragesaremeansofeachlab'sIrntervalsshown(OpenAIn=4,Anthroplcn=4).

2·REASONING

3·AGENTIC

ERA

ERA

Source:SemiAnalysis

Yetdespitethemassiveexplosionineconomicvaluecreatedbyfrontiermodels,thegapclosedfasterinEra3thaneithererabeforeit.KimiK2.6surpassedOpus4.5withascoreof56.3in4.8months,andGLM-5.2clearedGPT-5.2withascoreof72.4in6

months.Thetrendoftheclosingtimehalvingwitheachsubsequenteraisremarkablyconsistent.

12/15

Source:SemiAnalysis

Lookingforward

So,whatdoesthisallmeanforthefutureofclosedvsopensourcemodels?

First,wewanttoacknowledgethatbenchmarksarenottheendallbeall.KimiK3mayscorehigherthanFable5onourcuratedcomposite,butwestillpreferusingFableat

SemiAnalysisforourdaytodaywork.ThisispartlybecauseAnthropichasdoneabetterjobproductizingtheirmodelviathingslikeClaudeCodeandClaudeTag,butalsolargely

becausebenchmarksaren'taperfectproxyforrealwork.Thisisespeciallytruefor

publicbenchmarks,whichmodelmakerscaneasilyhillclimbbysimplycreatingabunchofRLenvironmentsthatcloselymimicthebenchmarktasks.

Second,youmayarguethattheclosingtimeforEra3isartificiallydeflateddueto

AnthropicandOpenAlspendingmoretimeonsafetytestingthanMoonshotandZhipu,butthisisnotanewphenomenon.GPT-4,forexample,finishedtraining218daysbefore

release.EvenifweassumeMythosfinishedtraininginmidFebruary,that'sstillonlya114daydelaybeforetheFablerelease.

TheUpcomingEra

Webelieveweareonthecuspofanotherstepfunctionimprovementinclosedsourcemodelcapabilities.Itwilldramaticallyre-widenthegapvsopensourceandrequireanentirelynewsetofbenchmarkstomeasureproperly.

13/15

【价值目录】网整理:

WethinkthekeybreakthroughforthiserawillbeAImodelsthatcanrunautonomouslyformultipledaysatatime,andcollaboratewithmanycopiesofeachothertosolveextremelydifficultlong-horizontasks.Yes,theleadingclosed-sourcemodelsarealreadycapableofthistosomedegree,butwethinkit’sabouttogetsignificantlybetter.

WegotatasteofwhatthiswilllooklikeinJuly,whenanunreleasedOpenAImodel,alongwithGPT-5.6,brokeintoHuggingFace.ThemodelswerebeingtestedonExploitGym,abenchmarkthatasksamodeltoturnaknownvulnerabilityintoaworkingexploit.As

HuggingFacereports,themodelwentlookingfortheanswerkeyratherthanattemptingtosolvetheproblemitself,ultimatelyescapingOpenAI’sevaluationsandboxthroughazero-dayinitspackage-registryinfrastructure,exploitingvulnerabilitiesinHuggingFace’s

dataset-processingpipeline,andthenusingmisconfiguredKubernetespermissionstogaincontrolofproductionnodesandmovedeeperintoHuggingFace’sinfrastructure.Crucially,

itwasn’tjustasingleinstanceofthemodelthatdiscoveredtheseexploits,but

rathermanycopiesofthemodelworkingtogetherformultipleweeks.Thisled

OpenAItoannouncetheyhadpausedRLtra

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论