版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
NotforredistributionwithoutwrittenconsentofMorganStanley
MorganstanleyRESEARCH
October5,202606:27AMGMT
Semiconductors|NorthAmerica
HowcantheAIecosystemworkaroundmemorybottlenecks?
WeexpectthememoryshortagetobeaconstraintonAIbuildsthroughthedurationofthisAIcycle–butAIwon'twaitfor
DRAMfabstobebuilt.Welookatwaystoworkaroundshortages.
KeyTakeaways
Intensityofmemoryshortagesmayfluctuate,butAIwillbasicallyuseallsupplyfortheforeseeablefuture
CEOJensenHuanghassaidthattheindustryneedstothinkdifferentlytoworkaroundtheseshortages.Wecontemplatewhathemightbeconsidering
AmongthepotentialbeneficiarieswouldbeMRVLandALABforCXL&largerscale-up,CBRSfordisaggregation
RemainOWMUandSNDKaswedon'tseetheshortageending
Whatwasthecatalystforthisnote?Inrecentconversationswithcompute
companieswehaveheardagreatdealabouttheneedtoworkaroundmemory
bottlenecks-NVIDIACEOJensenHuangamongthem.Thiswasnotanegative
memoryconversation-it'sveryclearthattheindustryanticipatesmultipleyearsofshortages.
Butifmemorysupplycannotkeepupwithtokengrowth,theindustrywillneedtobecomemorememoryefficient.NVIDIAhastalkedaboutusingitscontrolover
compute,networking,andstoragetomitigatethosebottlenecks.
Inthisnote,welookatsomeofthewaysthattheindustrymighttrytoworkaroundmemoryshortages,notably:
De-speccing,reducingtherackscalecentralmemory(LPDDR5),andselectivelyreducinghighbandwidthmemory.Thisisinmanywaystheleastfavorable,butsometimesneeded,approach.Ourviewisthattherewillcontinuetobeproductsthatmaximizememorycontent,butthattherewillbeselectivede-speccing.
Potentialbeneficiaries:Nonedirectly,butallowscomputecompaniestomaketheirnumbers-andshiftssomeofthebottleneckfromDRAMtoNAND.Atthemargin,de-speccingpushesforalargerscaleupdomain,whichwouldhelpscaleupplayssuchasALABandMRVL.
DisaggregationofAIworkloads.ThisinvolvesbreakingupAItasksintomemoryintensiveandlessmemoryintensiveportions,whereareassuchasprefillwhichdonotrequireasmuchmemory.
WealsonotethatasDRAMpricesrise,someofthelowerlatencysolutionsfrom
IDEA
MoRGANSTANLEy&Co.LLCJosephMoore
EquityAnalyst
Joseph.Moore@
+1212761-7516
EllaTulchinsky
ResearchAssociate
Ella.Tulchinsky@
MasonWayne
ResearchAssociate
+1212761-2222
Mason.Wayne@
CateFolan
+1212761-6012
ResearchAssociate
Cate.M.Folan@
NicoleKozhukhov
ResearchAssociate
+1212296-3520
Nicole.Kozhukhov@
ShaneBrett
EquityAnalyst
+1212761-1636
Shane.Brett@
+1212761-1022
SEmiconductoRs
NorthAmerica
IndustryViewAttractive
MorganStanleydoesandseekstodobusinesswith
companiescoveredinMorganStanleyResearch.Asaresult,investorsshouldbeawarethatthefirmmayhaveaconflictofinterestthatcouldaffecttheobjectivityofMorganStanley
Research.InvestorsshouldconsiderMorganStanley
Researchasonlyasinglefactorinmakingtheirinvestmentdecision.
Foranalystcertificationandotherimportantdisclosures,refertotheDisclosureSection,locatedattheendofthisreport.
2
Morganstanley
RESEARCH
IDEA
companiessuchasGroq(nowwithinNVIDIA)orCerebrasorothersuseonchip
SRAM,whichhashistoricallybeenmuchmoreexpensivethanDRAM,butcurrentlyisnot.
Potentialbeneficiaries:computesolutionsthatcanbeafactorindisaggregatedworkloads,suchasCerebras,andNVIDIAthroughGroq
CXLmemorycentralization-CXLisaswitchedaccesstoDRAMwhichallowsforsubstantialefficiencyadvantagestocentralizedmemory"pooling",allowingseveralprocessorstoaccessthesamememory.Thishasgenerallybeenconsidereda
generalpurposecomputetechnology,butrecentcommentsfromchipsuppliersarepointingtolargeopportunitieswithinAIaswell.
Potentialbeneficiaries:ALAB,MRVL.
Isn'tthisallbadforDRAM?Intermsofneartermearningsmaximization,yes,a
little.AnenvironmentinwhichtheAIecosystemcomestoahaltandstopsbuildingbecauseoftheDRAMshortagesprobablyleadstothebestneartermpricing
outcomes.
Butthatwasneverrealistic,andwehaveconsistentlyarguedforduration,not
amplitude,astheprimarydriverofthecyclefromhere.Andwewouldmakethe
argumentthatreducingmemoryspecsbecauseofsupplyconstraints-importantly,notbecauseofpricing-leadstoasituationwheremoresupplywillbeveryeasilyabsorbedaswerespectohigherlevels.
Whilewedon'ttakethisforgranted-wedoextensivemonthlychecksintomemoryconditionsatthispoint-ourworkinghypothesisisthatbecauseofthispent-up
demand,therisktothememorycycleisthesameastherisktothecomputespace,thatisadecelerationinAIitself.(Weshouldintroduceoneriskfactorhere-thatif
AIslows,notbecauseofdemand,butbecauseofland/power/shellbottlenecks,suchapausecouldcausememorydisruptionthatwewouldnotseewiththecompute
names.
Forthatreason,wecontinuetobepositiveonMUandSNDKinourcoverage.
Morganstanley
RESEARCH
IDEA
TheBruteForcereductioninmemoryspecifications
Summary:Memoryisakeyaspecttocomputeperformance,andgenerallythereis
reluctancetolowerspecifications.ButtheworldisnotgoingtowaitforDRAMfabstogetbuilt,sothereareoptimizationsaroundlowermemorylevels.Wewouldnotethatsofareveryreductioninspecificationscausesmemorybottleneckstosurfaceinotherpartsofthecomputespectrum.
WeareseeingNvidiaofferreducedmemorycontentskewsforrubin,includingbothHBMandmainmemory,inanattempttooffsetrisingpricesandenablecontinuedunitgrowthbutthat'slessofanegativeformemorycompaniesthanitappears.IncreasingmemorydemandforAIaseculartrend,andreducingmemorycontentatonelevelofthememoryhierarchyputsincrementalpressureontheremainingtiers,shiftingthememorybottleneckmorethaneliminatingit.
What'shappeningwithde-speccingtoday?Thecurrentshortagesandrisingpricesof
bothDRAMandNANDhasputsignificantpressureonmemorybuyers,andinresponsecompaniesthroughouttheecosystemarelookingforwaystohelpeasethoseconstraints.Themoststraightforwardresponse,especiallyforDRAM,istoreducetheamountof
memoryperdeviceorperserver.We'reseeingthishappenoutofnecessityinmostcases,wherereducingtheamountofDRAMisameasuretakentoenablehigherunitshipments.WebelieveNvidiaisofferingadditionalrubinskewstocustomers,withLPDDR5perrackcontentof28TBvstheoriginal54TBsbyaskingDRAMvendorstosupply96GB
SOCAMM2modulesvs192GBs.ForHBMRubinwasoriginallyexpectedtobe288GBofHBMperGPU(8stacksof12hiHBM4),butwenowexpectskewswith192GB(8stacksof8hiHBM4)tobeoffered.ForRubinultracontentwasoriginallydebutedat1TBHBM4eper4dieaccelerator,themovetokeeprubinultraata2diedesignmakes512GBthe
relevantcomparisonvstheoriginallyannouncedspecs.Versusthat512wecouldseeNvidiastayatHBM4ratherthanmoveto4e,aswellasoffer8hivs12hialternativesloweringoverallcapacitytothe192GB-384GBrange.
WhatdrivestheneedformorememorycapacityinAIworkloads?Before
understandingwhat'slostwhenmemorycapacityisreduced,wethinkitsimportantto
lookatwhytherehasbeenapersistentpushtoaddmorememoryovertime.Wehighlightthreefactorsbelow:
Modelsize:Overdecadesofmachinelearninghistorytherehasbeenaconsistentdesireforlargerandlargermodels,asresearchhasconsistentlyshownthelargerthemodelthebetteritsperformance.ModelweightsarealmostalwaysstoredinHBM,increasingtheneedforlargerHBMcapacityovertime.
MoRGANSTANLEyREsEARcH3
4
Morganstanley
RESEARCH
IdEa
Exhibit1:FrontierAImodelsizesaredoublingevery6months
Parameters(logscale)
LLMeratrend:4.0x/yr,doublesevery6mo
Pre-LLMeratrend:1.3x/yr,doublesevery31mo
195519601965197019751980198519901995200020052010201520202025
Publicationdate
Pre-LLMeraLLMeraPre-LLMtrendLLMeratrend
1E+13
1E+12
1E+11
1E+10
1E+9
1E+8
1E+7
1E+6
1E+5
1E+4
1E+3
1E+2
1E+1
1E+0
1950
FrontierAIModelSizeOverTime
Source:
EPOC.ai
,MorganStanleyResearch
Contextlength:LLMscanonlytakeintoaccountacertainnumberoftokensatonetime,includingtheprompt,uploadeddocuments,conversationhistory,andthemodeloutput.Thatlimitiscalledthecontextlengthorcontextwindow,withmostleadingmodelstodayatabout1Mtokens.AsAIusecasesincreasinglyrevolvearoundlonghorizonagentictasksthecontextwindowisacriticalbottleneckfortheamountofworkanagentcan
accomplish.
Exhibit2:LLMcontextwindowsaregoing~5xannually
Contextwindow(tokens,logscale)
1,000,000
100,000
10,000
1,000
Mar2023Jul2023Nov2023Mar2024Jul2024Nov2024Mar2025Jul2025
Releasedate
ClosedmodelsOpenmodelsClosedtrendOpentrend
10,000,000
Closed:~4.8x/yr(+383%)
Open:~6.1x/yr(+508%)
ContextWindowsOverTime:ClosedvsOpenModels
Source:
EPOC.ai
,MorganStanleyResearch
Inferenceconcurrency:ThenumberofrequestsanAIsystemservesatthesametime.
LLMinferencehastwosteps,prefillanddecode.Prefillisusuallylimitedbycompute,
whiledecodeisusuallylimitedbymemory.Moresimultaneousrequestsrequiremore
memorycapacity,becauseeachsessionstoresitsowncontext(theKVcache),andmorememorybandwidth,becauseeverydecodestepneedstoprocessdatafromtheKVcache.
MoRGANSTANLEyREsEARcH5
Morganstanley
RESEARCH
IDEA
WeexpectovertimeLLMswillbecomesmarter(modelsize),needmorein-contextinformation(contextlength),andservemoreusers/agentsovertime(inference
concurrency).Thatmeansde-speccingmemoryislikelytoonlyprovetemporaryaseachofthosevectorscontinuetorequiregreateramountsofmemoryacrossthehierarchy.
Whencapacityisreducedatonelevelofthememoryhierarchyitputspressureontheremainingtiers,de-speccingHBM/LPDDRiscreatingopportunitiesfornetworked
DRAM(CXL)andNAND.GiventhefactorsoutlinedabovecreatingincrementaldemandformemorycapacityforAI,whenonetierisreduced(i.eHBM)thedataneedstomove
elsewherecreatingnewpainpointsinmemoryandIO.Weareseeingthisinrealtime,asNvidiamovestode-specLPDDR5racksweareseeingincrementaldemandforNANDasKVcachestoragedemandsincreaseforthatlevelofthememorytier.WearealsoseeinglowerDRAMcontentputtingincreasingweightonnetworking,asbitsthatwerestoredinHBMneedtobespreadacrossmoreGPUs(puttingpressureonscale-upbandwidth),orreadfrommainmemory(puttingpressureonCPUIO).
Exhibit3:De-speccingHBM/LPDDRshouldopenupopportunitieslowerinthememoryhierarchy
Source:TrendForce,MorganStanley
Therewillalsobeopportunitiesthatemergefornewtiersinthememoryhierarchy;we
discussCXLandSRAMbasedarchitecturesbelowbutgoingforward,opticallyconnectedmemoryisaninterestingavenuefortheindustryaswellasnewtechnologiessuchasHBF.
IscuttingHBMcapacitya“freelunch”forbandwidthboundworkloads?IfHBM
contentisbeingreducedbyloweringHBMstackheights,liketherubinde-spec(movingfrom12hito8hi),bandwidthisn'taffected,asHBMbandwidthisfunctionoftheinterfaceandpinspeeds.Thatcanworkwellforsmallcontextandsmallermodels,wherethereislowercapacityrequirementsbutitcreatesrisks.Iftrendsoutlinedabove(longcontext,
agenticworkloads)continuethenumberofworkloadsthatfallintothatcategoryislikelytoshrink,requiringmoreGPUsand/orincreasinglatencyasmoredataifoff-loadedto
lowermemorytiers.ThatsaidbandwidthisakeybottleneckforLLMdecode,and
loweringstackheightsisastraightforwardwaytoreducecosts.GoingforwardthekeytowatchifHBMvendorscanchargeapremiumforhigherbandwidthasHBMgenerations
progress-thathasbeentruehistorically.However,ascustombasedie'sareintroduced,thedriverofHBMbandwidth(theDRAMPHY)canbedesignedbythecomputevendor,
6
Morganstanley
RESEARCH
IDEA
whichmayshiftthevalueawayfromDRAMvendors.
Thescale-updomainisalsoafactortoconsiderwithRubinUltraandbeyond.An
effectivescale-upmechanismeffectivelycreatesarackthatactsasasingleGPU,byusingamultiplexedswitchingmechanismtoconnectalltoallwithintherack.Wedon'tknowwhatRubinUltralookslikeatthispoint-thecompanyhassaidtheyhavealteredthe
Kyberformfactorandwillhaveahigherscaleupdomain-iemoreGPUsperrack-andthiswouldallowthetotalmemorycontentattherackleveltostayconstantwithlessmemoryperGPU.
Exhibit4:NvidiacustomHBM4eoffers30%morebandwidthvsJDECstandardHBM4e
Source:Nvidia丿MorganStanley
Exhibit5:NvidiacustomHBMreducesPHYandsupportarea67%vsJDECstandardHBM4e
Source:Nvidia丿MorganStanley
MoRGANSTANLEyREsEARcH7
Morganstanley
RESEARCH
IDEA
Noteverysteprequiressomuchmemory:Disaggregation
Theopportunityindisaggregationcomesfromoptimizinghardwarearoundthe
differentrequirementsofeachstageofinference.Notallpartsofinferenceplacethe
samedemandsoncomputeandmemory.Prefillprocessestheinputpromptinparallel
andisrelativelycomputeintensive.Decodegeneratesoutputonetokenatatimeand
repeatedlyaccessesmodelweightsandthegrowingKVcache,makingmemorybandwidthandcapacitymoreimportant.Thesameacceleratoristhereforebeingaskedtohandletwoworkloadswithverydifferentresourcerequirements.
Disaggregationseparatesthesestagessoeachcanrunonhardwarebettersuitedtothetask.Computedensesystemscanhandleprefill,whilearchitecturesoptimizedfor
memorybandwidthandlowlatencycanhandledecode.Thisallowsoperatorstoscalethetworesourcesindependentlyandimproveutilization,ratherthanprovisioningthesamemixofcomputeandmemoryforeverystageofinference.
Fromamemoryperspective,disaggregationallowstheindustrytousedifferenttypesofmemorywheretheyaremostvaluable.Memorybandwidthintensiveportionsof
inferencecanmovetoarchitecturesbuiltaroundhighbandwidthSRAMorother
approachesthatreducedatamovement,whileHBMandlargercapacitymemorycan
remainfocusedonworkloadsthatneedtheircapacityandflexibility.Disaggregationdoesnoteliminatememorydemand,butheterogeneousarchitecturescanimprovehow
efficientlyscarceandexpensivememoryisusedasinferencescales.
Exhibit6:AIinferenceconsistsoftwodistinctstages:prefillanddecode
Source:AdaptedfromPateletal.,“Splitwise:EfficientGenerativeLLMInferenceUsingPhaseSplitting,”ISCA2024;MorganStanleyResearch.
Whynow?Thememorydemandsofinferenceareincreasingasmodelsprocesslongercontextandgeneratemoreoutputtokens.Reasoningmodels,codingagentsandotheragenticapplicationscanrequiresubstantiallymoreinferencetimethantraditional
questionandanswerworkloads,increasingtheimportanceofbothmemorycapacityandbandwidth.Atthesametime,fasterinterconnectsandimprovementsininference
softwarearemakingitmorepracticaltomovetheKVcachebetweenseparateprefillanddecodeengines,helpingaddressoneofthemainimplementationchallengesof
disaggregation.
Interestindisaggregatedinferenceacceleratedsharplyaroundtheendof2025.NVIDIA
acquiredGroqaroundthistime,AWSsubsequentlyannouncedadisaggregated
architecturepairingTrainiumforprefillwithCerebrasfordecodeoverEFA,andCerebrashassinceannouncedasimilarapproachwithAMD.Meanwhile,dMatrixhasmovedCorsair
8
Morganstanley
RESEARCH
IDEA
intovolumeproduction.Takentogether,thesedevelopmentspointtoabroadermovetowardheterogeneousinference,wheredifferentprocessorsareoptimizedfordifferentpartsoftheworkload.
Theemergingsolutionstakedifferentapproachestothememoryproblem,ranging
fromreplacingexternalmemorywithhighbandwidthSRAMtocombiningSRAM,HBMandDDRwithinthesameinferencesystem:
•CBRS:CerebrasusesitswaferscaleprocessortocombinealargeamountofSRAMandcomputeonthesamepieceofsilicon,reducingtheneedtomovedata
betweenprocessorsandexternalmemory.Thearchitectureprovidesveryhigh
memorybandwidth,whichisparticularlyusefulduringdecodewhenmodel
weightsmustbeaccessedrepeatedlyforeachgeneratedtoken.Cerebrascanrunbothprefillanddecode,whileitsdisaggregatedarchitectureswithAMDandAWSuseGPUsorTrainiumforthemorecomputeintensiveprefillstageandCerebrasforthememorybandwidthintensivedecodestage.
•NVIDIAGroq:Groq'sLPUuseshighbandwidthonchipSRAMandacompiler
scheduledarchitecturedesignedforlowlatencyinference.NVIDIAisintegratingGroqtechnologyalongsideitsVeraRubinplatform,combiningHBMbasedGPUswithSRAMbasedLPUs.RubinGPUshandleprefillandtheattentionportionofdecode,includingworkloadsthatrequirelargeKVcaches,whileGroqhandlesthefeedforwardandmixtureofexpertsportionsofdecode.NVIDIADynamo
coordinatestheworkloadandtransfersactivationsbetweenthetwosystems.
•DMatrix:dMatrixusesadigitalinmemorycomputearchitecturethatplaces
computedirectlyalongsidehighbandwidthonchipSRAM,allowingmodelweightstobeprocessedwheretheyarestoredratherthanrepeatedlymovedbetween
separatememoryandcomputeresources.Thisisdesignedtoreducethedata
movementbottleneckthatbecomesparticularlyimportantduringdecode,wheremodelweightsareaccessedrepeatedlyforeachgeneratedtoken.dMatrix
combinesthishighbandwidthlocalmemorywithlargercapacitymemoryfordatathatdoesnotneedtoremainonchip,andispositioningitsacceleratorsalongsideGPUsinheterogeneoussystemswhereGPUshandleprefillanddMatrixhandlesdecode.
•SambaNova:SambaNovausesareconfigurabledataflowarchitecturethatmapsthemodel'scomputationdirectlyontotheprocessor,allowingdatatoflowfromoneoperationtothenextwithoutrepeatedlyreturningtomemorybetweeneachstep.ThearchitectureispairedwithathreetiermemoryhierarchyspanningonchipSRAM,HBMandDDR,withthefastestmemoryusedforthehottestlocal
dataandlargermemorytiersprovidingcapacityformodelweights,KVcacheandotherlesslatencysensitivedata.SambaNovaisusingthisarchitecturein
heterogeneousinferencesystemswhereGPUshandleprefill,itsRDUshandledecodeandCPUsmanageorchestrationandagentictoolexecution.
WeexpectinferencespendingtogrowfasterthantrainingasAIusageexpandsand
eachinteractionconsumesmorecompute.Disaggregatedinferenceshouldbe
particularlyrelevantforworkloadswhereprefillanddecodeplaceverydifferentdemandsontheunderlyinghardware.Codingassistantsandagenticapplicationsaregood
examples:theycaningestlargeamountsofcontext,suchascodebases,documentsor
priorinteractions,andthengeneratesubstantialoutputacrossrepeatedmodelcalls.This
MoRGANSTANLEyREsEARcH9
Morganstanley
RESEARCH
IDEA
createsamixofcomputeintensiveprefillandmemorybandwidthintensivedecode,
makingitmoreattractivetoseparatethetwostagesandscaleeachindependently.Longcontextworkloadscanalsobenefitbecausealargeprefillrequestcanotherwiseinterferewithtokengenerationforotherusers.Astheseworkloadsbecomealargershareof
inference,disaggregationprovidesanotherwayfortheecosystemtoscaleAIoutput
withoutrequiringeverystagetobeprovisionedwiththesamemixofcomputeandhighperformancememory.
Exhibit7:TotalAIinfrastructurespendforecast
Source:BBG,MorganStanleyResearch
PotentialStockBeneficiaries:CBRS,NVDA
WeseeCerebrasasoneoftheclearestpureplaybeneficiariesofgreateradoptionofdisaggregatedinference.ThecompanyisalreadyworkingwithbothAMDandAWSon
systemsthatpairthirdpartyprocessorsforprefillwithCerebrasforthememory
bandwidthintensivedecodestage.WithAMD,HelioshandlespromptprocessingandlargecontextwindowswhiletheCerebrasWaferScaleEnginehandlesdecode,withthe
solutionexpectedtoenterproductionin4Q26.AWSistakingasimilarapproachwith
TrainiumforprefillandCerebrasfordecode,withthecombinedsolutionexpectedto
reachAmazonBedrockin1Q27.ThesepartnershipsgiveCerebrasapathtoparticipateinheterogeneoussystemsevenwhereanotheracceleratorremainstheprimarycomputeengine
DisaggregationcanalsomateriallyimprovetheeconomicsofCerebrasCloud.CerebrasandAMDhavedisclosedthattheircombinedsystemcanincreasethroughputbyupto5xwhilemaintainingCerebrasinferencespeed,allowingeachCerebrassystemtoproduce
substantiallymoretokensfromthesameunderlyingCerebrascapacity.BecauseCerebrasoperatesitsowncloudinfrastructure,lowercostpertokencanaccruedirectlytothe
companythroughhighergrossmargin,lowercustomerpricing,orsomecombinationofthetwo.ThismakesdisaggregationmorethananincrementalhardwareopportunityforCerebras,itcanincreasetheamountofrevenuegeneratedfromeachdeployedsystemwhileloweringthecostofproducingeachtoken.
Otherpotentialbeneficiariesincludecomputesolutionsthatcanparticipatein
disaggregatedinferenceworkloads.NVIDIAisincorporatingGroqtechnologyintoVera
10
Morganstanley
RESEARCH
IDEA
Rubin,pairingHBMbasedGPUswithSRAMbasedLPUsfordifferentportionsofinference.Asdisaggregatedinferencebecomesmorecommon,weseeanopportunityforabroadersetofprocessorstoparticipateinAIinfrastructurebyspecializingaroundtheportionsoftheworkloadwheretheirarchitectureismostefficient.
MoRGANSTANLEyREsEARcH11
Morganstanley
RESEARCH
IdeA
CXL:Usingnetworking/interconnecttechnologiestooptimizememoryusage
TheCXLOpportunity
CXL,orComputeExpressLink,isahighspeedinterconnectdesignedtogiveprocessors
moreflexibleaccesstomemory.Itusesthesameunderlyingphysicalinfrastructureas
PCIe,butaddsspecializedmemoryandcacheprotocolsthatallowexternalmemoryto
behavemuchmorelikememorydirectlyattachedtotheprocessor.ThekeybenefitisthatmemorynolongerneedstoberigidlytiedtoaspecificCPUorserver,allowingdatacenteroperatorstouseavailableDRAMmoreefficientlyandaddcapacitywithoutnecessarily
addingmorecompute.
Weseethreemainusecases:Expansionallowsaservertoaddmemorybeyondthe
capacitysupportedbyitslocalmemorychannels.Sharingallowsmultipleprocessorstoaccessthesamememoryresources,reducingduplicationandimprovingoverallmemoryutilization.Poolingcreatesacommonpoolofmemorythatcanbedynamicallyallocatedacrossmultipleprocessorsorservers,reducingstrandedcapacityandoverprovisioning.
Exhibit8:MemoryExpansion:Addsadditionalmemory
capacitytoasingleprocessorbeyonditslocallyattachedDRAM.
Source:AsteraLabs
Exhibit9:MemorySharing:Allowsmultipleprocessorstoaccessthesamememoryresourceanddata.
Source:AsteraLabs
12
Morganstanley
RESEARCH
IDEA
Exhibit10:MemoryPooling:Createsacommonmemorypoolthatcanbedynamicallyallocatedacrossmultipleprocessors.
Source:AsteraLabs
WhyCXLforAI?
Inthepast,CXLhasbeenusedprimarilyfortraditionalCPUbasedcompute,where
workloadsoftenbenefitmorefromadditionalmemorycapacitythanfromoptimizingfortheabsolutelowestlatencyandhighestbandwidth.CXLintroducesanadditionalhop
betweentheprocessorandmemory,makingCXLattachedDRAMslowerthanlocallyattachedDRAM.ThishashistoricallylimitedCXL'sroleinthemostlatencyand
bandwidthsensitiveAIworkloads,whereGPUsrelyonHBMforextremelyhighbandwidth.
Thebroaderbackdropisthatcomputeperformancehasscaledmuchfasterthanmemoryandinterconnectbandwidth,asshowninExhibit11.Thiswideninggapmakesmemoryplacementandutilizationincreasinglyimportanttooverallsystemperformance.Italsoincreasesthevalueofatieredmemoryarchitecture,whereHBMisreservedforthemostbandwidthsensitivedatawhilelargerpoolsofDRAMcansupportworkloadsthatare
morecapacityconstrained.
Morerecently,CXLhasattractedgrowinginterestforAIinferenceaslongercontext
windows,agenticworkloads,andgrowingKVcachescreatemuchlargermemorycapacityrequirements,whilemuchofthatdatadoesnotnecessarilyneedtoresideinHBMatalltimes.CXLattachedDRAMcanthereforeserveasalowercost,highercapacity
memorytier,withthemostperformancesensitivedataremainingclosetotheacceleratorandlesssensitivedatamovedintoexternalmemory.
MoRGANSTANLEyREsEARcH
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 杭州医药产业发展趋势研究报告范文
- 婚庆个人租车合同范本
- 小学五年级道德与法治“9 中国有了共产党”革命历史主题教案
- 2026中国工业元宇宙概念落地场景与投资可行性研究报告
- 2026医疗健康产业科技创新趋势及市场增长潜力与投资战略研究报告
- 2026中国智能家居市场消费者偏好与产品升级趋势
- 2026区块链技术商业化路径与投资价值评估报告
- 2026住宅室内空气质量检测仪精度提升解决方案与市场需求对比分析评估报告
- 2026中国慢性肾病透析设备市场供需状况与投资回报预测
- 2026中国智慧医院建设标准与实施路径研究报告
- 高盛-变革中的中国:中企出海浩荡征程从边缘市场走向核心腹地(摘要)-20260914
- 中国鼻内镜检查操作规范(2025版)
- 2026年中国中煤能源秋招面试题及答案
- 2026 年夏季高校宿舍反诈知识普及主题教育班会
- 2026年山东省考《申论》真题及答案解析(B卷)
- 中国广电山东网络有限公司2026年度市县公司招聘(145个)笔试历年常考点试题专练附带答案详解
- 2026北京急救中心第一批招聘备考考试题库含答案解析
- ICU危重患者呼吸机管理
- 社会语言学讲稿
- 乡统计站工作制度
- 托育食品安全课件
评论
0/150
提交评论