摩根士丹利-如何让AI生态系统绕开内存瓶颈?-How can the AI ecosystem work around memory bottlenecks-20261005_第1页
摩根士丹利-如何让AI生态系统绕开内存瓶颈?-How can the AI ecosystem work around memory bottlenecks-20261005_第2页
摩根士丹利-如何让AI生态系统绕开内存瓶颈?-How can the AI ecosystem work around memory bottlenecks-20261005_第3页
摩根士丹利-如何让AI生态系统绕开内存瓶颈?-How can the AI ecosystem work around memory bottlenecks-20261005_第4页
摩根士丹利-如何让AI生态系统绕开内存瓶颈?-How can the AI ecosystem work around memory bottlenecks-20261005_第5页
已阅读5页,还剩27页未读, 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

NotforredistributionwithoutwrittenconsentofMorganStanley

MorganstanleyRESEARCH

October5,202606:27AMGMT

Semiconductors|NorthAmerica

HowcantheAIecosystemworkaroundmemorybottlenecks?

WeexpectthememoryshortagetobeaconstraintonAIbuildsthroughthedurationofthisAIcycle–butAIwon'twaitfor

DRAMfabstobebuilt.Welookatwaystoworkaroundshortages.

KeyTakeaways

Intensityofmemoryshortagesmayfluctuate,butAIwillbasicallyuseallsupplyfortheforeseeablefuture

CEOJensenHuanghassaidthattheindustryneedstothinkdifferentlytoworkaroundtheseshortages.Wecontemplatewhathemightbeconsidering

AmongthepotentialbeneficiarieswouldbeMRVLandALABforCXL&largerscale-up,CBRSfordisaggregation

RemainOWMUandSNDKaswedon'tseetheshortageending

Whatwasthecatalystforthisnote?Inrecentconversationswithcompute

companieswehaveheardagreatdealabouttheneedtoworkaroundmemory

bottlenecks-NVIDIACEOJensenHuangamongthem.Thiswasnotanegative

memoryconversation-it'sveryclearthattheindustryanticipatesmultipleyearsofshortages.

Butifmemorysupplycannotkeepupwithtokengrowth,theindustrywillneedtobecomemorememoryefficient.NVIDIAhastalkedaboutusingitscontrolover

compute,networking,andstoragetomitigatethosebottlenecks.

Inthisnote,welookatsomeofthewaysthattheindustrymighttrytoworkaroundmemoryshortages,notably:

De-speccing,reducingtherackscalecentralmemory(LPDDR5),andselectivelyreducinghighbandwidthmemory.Thisisinmanywaystheleastfavorable,butsometimesneeded,approach.Ourviewisthattherewillcontinuetobeproductsthatmaximizememorycontent,butthattherewillbeselectivede-speccing.

Potentialbeneficiaries:Nonedirectly,butallowscomputecompaniestomaketheirnumbers-andshiftssomeofthebottleneckfromDRAMtoNAND.Atthemargin,de-speccingpushesforalargerscaleupdomain,whichwouldhelpscaleupplayssuchasALABandMRVL.

DisaggregationofAIworkloads.ThisinvolvesbreakingupAItasksintomemoryintensiveandlessmemoryintensiveportions,whereareassuchasprefillwhichdonotrequireasmuchmemory.

WealsonotethatasDRAMpricesrise,someofthelowerlatencysolutionsfrom

IDEA

MoRGANSTANLEy&Co.LLCJosephMoore

EquityAnalyst

Joseph.Moore@

+1212761-7516

EllaTulchinsky

ResearchAssociate

Ella.Tulchinsky@

MasonWayne

ResearchAssociate

+1212761-2222

Mason.Wayne@

CateFolan

+1212761-6012

ResearchAssociate

Cate.M.Folan@

NicoleKozhukhov

ResearchAssociate

+1212296-3520

Nicole.Kozhukhov@

ShaneBrett

EquityAnalyst

+1212761-1636

Shane.Brett@

+1212761-1022

SEmiconductoRs

NorthAmerica

IndustryViewAttractive

MorganStanleydoesandseekstodobusinesswith

companiescoveredinMorganStanleyResearch.Asaresult,investorsshouldbeawarethatthefirmmayhaveaconflictofinterestthatcouldaffecttheobjectivityofMorganStanley

Research.InvestorsshouldconsiderMorganStanley

Researchasonlyasinglefactorinmakingtheirinvestmentdecision.

Foranalystcertificationandotherimportantdisclosures,refertotheDisclosureSection,locatedattheendofthisreport.

2

Morganstanley

RESEARCH

IDEA

companiessuchasGroq(nowwithinNVIDIA)orCerebrasorothersuseonchip

SRAM,whichhashistoricallybeenmuchmoreexpensivethanDRAM,butcurrentlyisnot.

Potentialbeneficiaries:computesolutionsthatcanbeafactorindisaggregatedworkloads,suchasCerebras,andNVIDIAthroughGroq

CXLmemorycentralization-CXLisaswitchedaccesstoDRAMwhichallowsforsubstantialefficiencyadvantagestocentralizedmemory"pooling",allowingseveralprocessorstoaccessthesamememory.Thishasgenerallybeenconsidereda

generalpurposecomputetechnology,butrecentcommentsfromchipsuppliersarepointingtolargeopportunitieswithinAIaswell.

Potentialbeneficiaries:ALAB,MRVL.

Isn'tthisallbadforDRAM?Intermsofneartermearningsmaximization,yes,a

little.AnenvironmentinwhichtheAIecosystemcomestoahaltandstopsbuildingbecauseoftheDRAMshortagesprobablyleadstothebestneartermpricing

outcomes.

Butthatwasneverrealistic,andwehaveconsistentlyarguedforduration,not

amplitude,astheprimarydriverofthecyclefromhere.Andwewouldmakethe

argumentthatreducingmemoryspecsbecauseofsupplyconstraints-importantly,notbecauseofpricing-leadstoasituationwheremoresupplywillbeveryeasilyabsorbedaswerespectohigherlevels.

Whilewedon'ttakethisforgranted-wedoextensivemonthlychecksintomemoryconditionsatthispoint-ourworkinghypothesisisthatbecauseofthispent-up

demand,therisktothememorycycleisthesameastherisktothecomputespace,thatisadecelerationinAIitself.(Weshouldintroduceoneriskfactorhere-thatif

AIslows,notbecauseofdemand,butbecauseofland/power/shellbottlenecks,suchapausecouldcausememorydisruptionthatwewouldnotseewiththecompute

names.

Forthatreason,wecontinuetobepositiveonMUandSNDKinourcoverage.

Morganstanley

RESEARCH

IDEA

TheBruteForcereductioninmemoryspecifications

Summary:Memoryisakeyaspecttocomputeperformance,andgenerallythereis

reluctancetolowerspecifications.ButtheworldisnotgoingtowaitforDRAMfabstogetbuilt,sothereareoptimizationsaroundlowermemorylevels.Wewouldnotethatsofareveryreductioninspecificationscausesmemorybottleneckstosurfaceinotherpartsofthecomputespectrum.

WeareseeingNvidiaofferreducedmemorycontentskewsforrubin,includingbothHBMandmainmemory,inanattempttooffsetrisingpricesandenablecontinuedunitgrowthbutthat'slessofanegativeformemorycompaniesthanitappears.IncreasingmemorydemandforAIaseculartrend,andreducingmemorycontentatonelevelofthememoryhierarchyputsincrementalpressureontheremainingtiers,shiftingthememorybottleneckmorethaneliminatingit.

What'shappeningwithde-speccingtoday?Thecurrentshortagesandrisingpricesof

bothDRAMandNANDhasputsignificantpressureonmemorybuyers,andinresponsecompaniesthroughouttheecosystemarelookingforwaystohelpeasethoseconstraints.Themoststraightforwardresponse,especiallyforDRAM,istoreducetheamountof

memoryperdeviceorperserver.We'reseeingthishappenoutofnecessityinmostcases,wherereducingtheamountofDRAMisameasuretakentoenablehigherunitshipments.WebelieveNvidiaisofferingadditionalrubinskewstocustomers,withLPDDR5perrackcontentof28TBvstheoriginal54TBsbyaskingDRAMvendorstosupply96GB

SOCAMM2modulesvs192GBs.ForHBMRubinwasoriginallyexpectedtobe288GBofHBMperGPU(8stacksof12hiHBM4),butwenowexpectskewswith192GB(8stacksof8hiHBM4)tobeoffered.ForRubinultracontentwasoriginallydebutedat1TBHBM4eper4dieaccelerator,themovetokeeprubinultraata2diedesignmakes512GBthe

relevantcomparisonvstheoriginallyannouncedspecs.Versusthat512wecouldseeNvidiastayatHBM4ratherthanmoveto4e,aswellasoffer8hivs12hialternativesloweringoverallcapacitytothe192GB-384GBrange.

WhatdrivestheneedformorememorycapacityinAIworkloads?Before

understandingwhat'slostwhenmemorycapacityisreduced,wethinkitsimportantto

lookatwhytherehasbeenapersistentpushtoaddmorememoryovertime.Wehighlightthreefactorsbelow:

Modelsize:Overdecadesofmachinelearninghistorytherehasbeenaconsistentdesireforlargerandlargermodels,asresearchhasconsistentlyshownthelargerthemodelthebetteritsperformance.ModelweightsarealmostalwaysstoredinHBM,increasingtheneedforlargerHBMcapacityovertime.

MoRGANSTANLEyREsEARcH3

4

Morganstanley

RESEARCH

IdEa

Exhibit1:FrontierAImodelsizesaredoublingevery6months

Parameters(logscale)

LLMeratrend:4.0x/yr,doublesevery6mo

Pre-LLMeratrend:1.3x/yr,doublesevery31mo

195519601965197019751980198519901995200020052010201520202025

Publicationdate

Pre-LLMeraLLMeraPre-LLMtrendLLMeratrend

1E+13

1E+12

1E+11

1E+10

1E+9

1E+8

1E+7

1E+6

1E+5

1E+4

1E+3

1E+2

1E+1

1E+0

1950

FrontierAIModelSizeOverTime

Source:

EPOC.ai

,MorganStanleyResearch

Contextlength:LLMscanonlytakeintoaccountacertainnumberoftokensatonetime,includingtheprompt,uploadeddocuments,conversationhistory,andthemodeloutput.Thatlimitiscalledthecontextlengthorcontextwindow,withmostleadingmodelstodayatabout1Mtokens.AsAIusecasesincreasinglyrevolvearoundlonghorizonagentictasksthecontextwindowisacriticalbottleneckfortheamountofworkanagentcan

accomplish.

Exhibit2:LLMcontextwindowsaregoing~5xannually

Contextwindow(tokens,logscale)

1,000,000

100,000

10,000

1,000

Mar2023Jul2023Nov2023Mar2024Jul2024Nov2024Mar2025Jul2025

Releasedate

ClosedmodelsOpenmodelsClosedtrendOpentrend

10,000,000

Closed:~4.8x/yr(+383%)

Open:~6.1x/yr(+508%)

ContextWindowsOverTime:ClosedvsOpenModels

Source:

EPOC.ai

,MorganStanleyResearch

Inferenceconcurrency:ThenumberofrequestsanAIsystemservesatthesametime.

LLMinferencehastwosteps,prefillanddecode.Prefillisusuallylimitedbycompute,

whiledecodeisusuallylimitedbymemory.Moresimultaneousrequestsrequiremore

memorycapacity,becauseeachsessionstoresitsowncontext(theKVcache),andmorememorybandwidth,becauseeverydecodestepneedstoprocessdatafromtheKVcache.

MoRGANSTANLEyREsEARcH5

Morganstanley

RESEARCH

IDEA

WeexpectovertimeLLMswillbecomesmarter(modelsize),needmorein-contextinformation(contextlength),andservemoreusers/agentsovertime(inference

concurrency).Thatmeansde-speccingmemoryislikelytoonlyprovetemporaryaseachofthosevectorscontinuetorequiregreateramountsofmemoryacrossthehierarchy.

Whencapacityisreducedatonelevelofthememoryhierarchyitputspressureontheremainingtiers,de-speccingHBM/LPDDRiscreatingopportunitiesfornetworked

DRAM(CXL)andNAND.GiventhefactorsoutlinedabovecreatingincrementaldemandformemorycapacityforAI,whenonetierisreduced(i.eHBM)thedataneedstomove

elsewherecreatingnewpainpointsinmemoryandIO.Weareseeingthisinrealtime,asNvidiamovestode-specLPDDR5racksweareseeingincrementaldemandforNANDasKVcachestoragedemandsincreaseforthatlevelofthememorytier.WearealsoseeinglowerDRAMcontentputtingincreasingweightonnetworking,asbitsthatwerestoredinHBMneedtobespreadacrossmoreGPUs(puttingpressureonscale-upbandwidth),orreadfrommainmemory(puttingpressureonCPUIO).

Exhibit3:De-speccingHBM/LPDDRshouldopenupopportunitieslowerinthememoryhierarchy

Source:TrendForce,MorganStanley

Therewillalsobeopportunitiesthatemergefornewtiersinthememoryhierarchy;we

discussCXLandSRAMbasedarchitecturesbelowbutgoingforward,opticallyconnectedmemoryisaninterestingavenuefortheindustryaswellasnewtechnologiessuchasHBF.

IscuttingHBMcapacitya“freelunch”forbandwidthboundworkloads?IfHBM

contentisbeingreducedbyloweringHBMstackheights,liketherubinde-spec(movingfrom12hito8hi),bandwidthisn'taffected,asHBMbandwidthisfunctionoftheinterfaceandpinspeeds.Thatcanworkwellforsmallcontextandsmallermodels,wherethereislowercapacityrequirementsbutitcreatesrisks.Iftrendsoutlinedabove(longcontext,

agenticworkloads)continuethenumberofworkloadsthatfallintothatcategoryislikelytoshrink,requiringmoreGPUsand/orincreasinglatencyasmoredataifoff-loadedto

lowermemorytiers.ThatsaidbandwidthisakeybottleneckforLLMdecode,and

loweringstackheightsisastraightforwardwaytoreducecosts.GoingforwardthekeytowatchifHBMvendorscanchargeapremiumforhigherbandwidthasHBMgenerations

progress-thathasbeentruehistorically.However,ascustombasedie'sareintroduced,thedriverofHBMbandwidth(theDRAMPHY)canbedesignedbythecomputevendor,

6

Morganstanley

RESEARCH

IDEA

whichmayshiftthevalueawayfromDRAMvendors.

Thescale-updomainisalsoafactortoconsiderwithRubinUltraandbeyond.An

effectivescale-upmechanismeffectivelycreatesarackthatactsasasingleGPU,byusingamultiplexedswitchingmechanismtoconnectalltoallwithintherack.Wedon'tknowwhatRubinUltralookslikeatthispoint-thecompanyhassaidtheyhavealteredthe

Kyberformfactorandwillhaveahigherscaleupdomain-iemoreGPUsperrack-andthiswouldallowthetotalmemorycontentattherackleveltostayconstantwithlessmemoryperGPU.

Exhibit4:NvidiacustomHBM4eoffers30%morebandwidthvsJDECstandardHBM4e

Source:Nvidia丿MorganStanley

Exhibit5:NvidiacustomHBMreducesPHYandsupportarea67%vsJDECstandardHBM4e

Source:Nvidia丿MorganStanley

MoRGANSTANLEyREsEARcH7

Morganstanley

RESEARCH

IDEA

Noteverysteprequiressomuchmemory:Disaggregation

Theopportunityindisaggregationcomesfromoptimizinghardwarearoundthe

differentrequirementsofeachstageofinference.Notallpartsofinferenceplacethe

samedemandsoncomputeandmemory.Prefillprocessestheinputpromptinparallel

andisrelativelycomputeintensive.Decodegeneratesoutputonetokenatatimeand

repeatedlyaccessesmodelweightsandthegrowingKVcache,makingmemorybandwidthandcapacitymoreimportant.Thesameacceleratoristhereforebeingaskedtohandletwoworkloadswithverydifferentresourcerequirements.

Disaggregationseparatesthesestagessoeachcanrunonhardwarebettersuitedtothetask.Computedensesystemscanhandleprefill,whilearchitecturesoptimizedfor

memorybandwidthandlowlatencycanhandledecode.Thisallowsoperatorstoscalethetworesourcesindependentlyandimproveutilization,ratherthanprovisioningthesamemixofcomputeandmemoryforeverystageofinference.

Fromamemoryperspective,disaggregationallowstheindustrytousedifferenttypesofmemorywheretheyaremostvaluable.Memorybandwidthintensiveportionsof

inferencecanmovetoarchitecturesbuiltaroundhighbandwidthSRAMorother

approachesthatreducedatamovement,whileHBMandlargercapacitymemorycan

remainfocusedonworkloadsthatneedtheircapacityandflexibility.Disaggregationdoesnoteliminatememorydemand,butheterogeneousarchitecturescanimprovehow

efficientlyscarceandexpensivememoryisusedasinferencescales.

Exhibit6:AIinferenceconsistsoftwodistinctstages:prefillanddecode

Source:AdaptedfromPateletal.,“Splitwise:EfficientGenerativeLLMInferenceUsingPhaseSplitting,”ISCA2024;MorganStanleyResearch.

Whynow?Thememorydemandsofinferenceareincreasingasmodelsprocesslongercontextandgeneratemoreoutputtokens.Reasoningmodels,codingagentsandotheragenticapplicationscanrequiresubstantiallymoreinferencetimethantraditional

questionandanswerworkloads,increasingtheimportanceofbothmemorycapacityandbandwidth.Atthesametime,fasterinterconnectsandimprovementsininference

softwarearemakingitmorepracticaltomovetheKVcachebetweenseparateprefillanddecodeengines,helpingaddressoneofthemainimplementationchallengesof

disaggregation.

Interestindisaggregatedinferenceacceleratedsharplyaroundtheendof2025.NVIDIA

acquiredGroqaroundthistime,AWSsubsequentlyannouncedadisaggregated

architecturepairingTrainiumforprefillwithCerebrasfordecodeoverEFA,andCerebrashassinceannouncedasimilarapproachwithAMD.Meanwhile,dMatrixhasmovedCorsair

8

Morganstanley

RESEARCH

IDEA

intovolumeproduction.Takentogether,thesedevelopmentspointtoabroadermovetowardheterogeneousinference,wheredifferentprocessorsareoptimizedfordifferentpartsoftheworkload.

Theemergingsolutionstakedifferentapproachestothememoryproblem,ranging

fromreplacingexternalmemorywithhighbandwidthSRAMtocombiningSRAM,HBMandDDRwithinthesameinferencesystem:

•CBRS:CerebrasusesitswaferscaleprocessortocombinealargeamountofSRAMandcomputeonthesamepieceofsilicon,reducingtheneedtomovedata

betweenprocessorsandexternalmemory.Thearchitectureprovidesveryhigh

memorybandwidth,whichisparticularlyusefulduringdecodewhenmodel

weightsmustbeaccessedrepeatedlyforeachgeneratedtoken.Cerebrascanrunbothprefillanddecode,whileitsdisaggregatedarchitectureswithAMDandAWSuseGPUsorTrainiumforthemorecomputeintensiveprefillstageandCerebrasforthememorybandwidthintensivedecodestage.

•NVIDIAGroq:Groq'sLPUuseshighbandwidthonchipSRAMandacompiler

scheduledarchitecturedesignedforlowlatencyinference.NVIDIAisintegratingGroqtechnologyalongsideitsVeraRubinplatform,combiningHBMbasedGPUswithSRAMbasedLPUs.RubinGPUshandleprefillandtheattentionportionofdecode,includingworkloadsthatrequirelargeKVcaches,whileGroqhandlesthefeedforwardandmixtureofexpertsportionsofdecode.NVIDIADynamo

coordinatestheworkloadandtransfersactivationsbetweenthetwosystems.

•DMatrix:dMatrixusesadigitalinmemorycomputearchitecturethatplaces

computedirectlyalongsidehighbandwidthonchipSRAM,allowingmodelweightstobeprocessedwheretheyarestoredratherthanrepeatedlymovedbetween

separatememoryandcomputeresources.Thisisdesignedtoreducethedata

movementbottleneckthatbecomesparticularlyimportantduringdecode,wheremodelweightsareaccessedrepeatedlyforeachgeneratedtoken.dMatrix

combinesthishighbandwidthlocalmemorywithlargercapacitymemoryfordatathatdoesnotneedtoremainonchip,andispositioningitsacceleratorsalongsideGPUsinheterogeneoussystemswhereGPUshandleprefillanddMatrixhandlesdecode.

•SambaNova:SambaNovausesareconfigurabledataflowarchitecturethatmapsthemodel'scomputationdirectlyontotheprocessor,allowingdatatoflowfromoneoperationtothenextwithoutrepeatedlyreturningtomemorybetweeneachstep.ThearchitectureispairedwithathreetiermemoryhierarchyspanningonchipSRAM,HBMandDDR,withthefastestmemoryusedforthehottestlocal

dataandlargermemorytiersprovidingcapacityformodelweights,KVcacheandotherlesslatencysensitivedata.SambaNovaisusingthisarchitecturein

heterogeneousinferencesystemswhereGPUshandleprefill,itsRDUshandledecodeandCPUsmanageorchestrationandagentictoolexecution.

WeexpectinferencespendingtogrowfasterthantrainingasAIusageexpandsand

eachinteractionconsumesmorecompute.Disaggregatedinferenceshouldbe

particularlyrelevantforworkloadswhereprefillanddecodeplaceverydifferentdemandsontheunderlyinghardware.Codingassistantsandagenticapplicationsaregood

examples:theycaningestlargeamountsofcontext,suchascodebases,documentsor

priorinteractions,andthengeneratesubstantialoutputacrossrepeatedmodelcalls.This

MoRGANSTANLEyREsEARcH9

Morganstanley

RESEARCH

IDEA

createsamixofcomputeintensiveprefillandmemorybandwidthintensivedecode,

makingitmoreattractivetoseparatethetwostagesandscaleeachindependently.Longcontextworkloadscanalsobenefitbecausealargeprefillrequestcanotherwiseinterferewithtokengenerationforotherusers.Astheseworkloadsbecomealargershareof

inference,disaggregationprovidesanotherwayfortheecosystemtoscaleAIoutput

withoutrequiringeverystagetobeprovisionedwiththesamemixofcomputeandhighperformancememory.

Exhibit7:TotalAIinfrastructurespendforecast

Source:BBG,MorganStanleyResearch

PotentialStockBeneficiaries:CBRS,NVDA

WeseeCerebrasasoneoftheclearestpureplaybeneficiariesofgreateradoptionofdisaggregatedinference.ThecompanyisalreadyworkingwithbothAMDandAWSon

systemsthatpairthirdpartyprocessorsforprefillwithCerebrasforthememory

bandwidthintensivedecodestage.WithAMD,HelioshandlespromptprocessingandlargecontextwindowswhiletheCerebrasWaferScaleEnginehandlesdecode,withthe

solutionexpectedtoenterproductionin4Q26.AWSistakingasimilarapproachwith

TrainiumforprefillandCerebrasfordecode,withthecombinedsolutionexpectedto

reachAmazonBedrockin1Q27.ThesepartnershipsgiveCerebrasapathtoparticipateinheterogeneoussystemsevenwhereanotheracceleratorremainstheprimarycomputeengine

DisaggregationcanalsomateriallyimprovetheeconomicsofCerebrasCloud.CerebrasandAMDhavedisclosedthattheircombinedsystemcanincreasethroughputbyupto5xwhilemaintainingCerebrasinferencespeed,allowingeachCerebrassystemtoproduce

substantiallymoretokensfromthesameunderlyingCerebrascapacity.BecauseCerebrasoperatesitsowncloudinfrastructure,lowercostpertokencanaccruedirectlytothe

companythroughhighergrossmargin,lowercustomerpricing,orsomecombinationofthetwo.ThismakesdisaggregationmorethananincrementalhardwareopportunityforCerebras,itcanincreasetheamountofrevenuegeneratedfromeachdeployedsystemwhileloweringthecostofproducingeachtoken.

Otherpotentialbeneficiariesincludecomputesolutionsthatcanparticipatein

disaggregatedinferenceworkloads.NVIDIAisincorporatingGroqtechnologyintoVera

10

Morganstanley

RESEARCH

IDEA

Rubin,pairingHBMbasedGPUswithSRAMbasedLPUsfordifferentportionsofinference.Asdisaggregatedinferencebecomesmorecommon,weseeanopportunityforabroadersetofprocessorstoparticipateinAIinfrastructurebyspecializingaroundtheportionsoftheworkloadwheretheirarchitectureismostefficient.

MoRGANSTANLEyREsEARcH11

Morganstanley

RESEARCH

IdeA

CXL:Usingnetworking/interconnecttechnologiestooptimizememoryusage

TheCXLOpportunity

CXL,orComputeExpressLink,isahighspeedinterconnectdesignedtogiveprocessors

moreflexibleaccesstomemory.Itusesthesameunderlyingphysicalinfrastructureas

PCIe,butaddsspecializedmemoryandcacheprotocolsthatallowexternalmemoryto

behavemuchmorelikememorydirectlyattachedtotheprocessor.ThekeybenefitisthatmemorynolongerneedstoberigidlytiedtoaspecificCPUorserver,allowingdatacenteroperatorstouseavailableDRAMmoreefficientlyandaddcapacitywithoutnecessarily

addingmorecompute.

Weseethreemainusecases:Expansionallowsaservertoaddmemorybeyondthe

capacitysupportedbyitslocalmemorychannels.Sharingallowsmultipleprocessorstoaccessthesamememoryresources,reducingduplicationandimprovingoverallmemoryutilization.Poolingcreatesacommonpoolofmemorythatcanbedynamicallyallocatedacrossmultipleprocessorsorservers,reducingstrandedcapacityandoverprovisioning.

Exhibit8:MemoryExpansion:Addsadditionalmemory

capacitytoasingleprocessorbeyonditslocallyattachedDRAM.

Source:AsteraLabs

Exhibit9:MemorySharing:Allowsmultipleprocessorstoaccessthesamememoryresourceanddata.

Source:AsteraLabs

12

Morganstanley

RESEARCH

IDEA

Exhibit10:MemoryPooling:Createsacommonmemorypoolthatcanbedynamicallyallocatedacrossmultipleprocessors.

Source:AsteraLabs

WhyCXLforAI?

Inthepast,CXLhasbeenusedprimarilyfortraditionalCPUbasedcompute,where

workloadsoftenbenefitmorefromadditionalmemorycapacitythanfromoptimizingfortheabsolutelowestlatencyandhighestbandwidth.CXLintroducesanadditionalhop

betweentheprocessorandmemory,makingCXLattachedDRAMslowerthanlocallyattachedDRAM.ThishashistoricallylimitedCXL'sroleinthemostlatencyand

bandwidthsensitiveAIworkloads,whereGPUsrelyonHBMforextremelyhighbandwidth.

Thebroaderbackdropisthatcomputeperformancehasscaledmuchfasterthanmemoryandinterconnectbandwidth,asshowninExhibit11.Thiswideninggapmakesmemoryplacementandutilizationincreasinglyimportanttooverallsystemperformance.Italsoincreasesthevalueofatieredmemoryarchitecture,whereHBMisreservedforthemostbandwidthsensitivedatawhilelargerpoolsofDRAMcansupportworkloadsthatare

morecapacityconstrained.

Morerecently,CXLhasattractedgrowinginterestforAIinferenceaslongercontext

windows,agenticworkloads,andgrowingKVcachescreatemuchlargermemorycapacityrequirements,whilemuchofthatdatadoesnotnecessarilyneedtoresideinHBMatalltimes.CXLattachedDRAMcanthereforeserveasalowercost,highercapacity

memorytier,withthemostperformancesensitivedataremainingclosetotheacceleratorandlesssensitivedatamovedintoexternalmemory.

MoRGANSTANLEyREsEARcH

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论