2026AI基础设施真实成本研究报告:本地部署与云部署总拥有成本TCO分析_第1页
2026AI基础设施真实成本研究报告:本地部署与云部署总拥有成本TCO分析_第2页
2026AI基础设施真实成本研究报告:本地部署与云部署总拥有成本TCO分析_第3页
2026AI基础设施真实成本研究报告:本地部署与云部署总拥有成本TCO分析_第4页
2026AI基础设施真实成本研究报告:本地部署与云部署总拥有成本TCO分析_第5页
已阅读5页,还剩51页未读 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

TheRealCostofAIInfrastructure

ATCOAnalysisof

On-PremisesvsCloud

May2026

May2026TheRealCostofAIInfrastructure2of31

ExecutiveSummary

There’sbeenatippingpointoverthelastcoupleofyearsasAImovedfrom

experimentstopilotprojectstothepointwhereorganizationshavedecidedthat

thereareenoughbenefitsfromAItojustifyspendingbigmoneyonit.That’swhenAIbecomesinfrastructure.

Overthenextseveralyears,webelievethatAIwon’tbediscussedasaseparate“thing.”AIfeatureslikemachinelearning,deeplearning,andgenerativeAIwillbefoldedintoanyapplicationorworkflowwhereitmakessense(andplentywhereitprobablydoesn’tmakesenseaswell).

In2025alone,theglobalAIinfrastructuremarket(AIhardware,notfacilities)was

$318billionaccordingtoIDC,morethandoublethe$153billionspentin2024.Thisisexpectedtocontinueunabatedfortheforeseeablefuture.

Theproblemthatmanyorganizationsarewrestlingwithrightnowisnotthe“if”they’regoingtoneedanAIinfrastructure,butthe“what”and“where”questions.

The“what”questionishighlydependentonuniqueorganizationalfactorslikewhattheyseeasthelowhangingAIfruitfortheirbusinessandhowbigamovetheywanttomake.

Inthisreport,weareaddressingthe“where”questionintermsofon-premisesvs.publiccloudoptionsandtherespectivecosts.

Wecomparedtheestimatedcostofdeployingandoperatingaproduction-scaleAIinfrastructureon-premisesversusinthreemajorcloudplatforms:AmazonWeb

Services(AWS),GoogleCloudPlatform(GCP),andOracleCloudInfrastructure

(OCI).Thecomparisonisapples-to-apples,basedonpubliclyavailablepricingandclearlydefinedworkloadassumptions.Themodelassumessteady-state,24×7

operationofa248-GPUenterpriseAIcluster.

Thescenarioexaminesamature,diversifiedmanufacturingorganizationthathasalreadycompletedextensiveAIpilotprogramsacrossmultiplebusinessunits.

Thoseeffortshavedemonstratedmeasurableoperationalimprovementsinareas

suchasdefectreduction,scheduling,supplychainoptimization,andAI-augmentedHPCworkloads.Basedonvalidateddemand,theorganizationisevaluatinga

centralized,sharedAlinfrastructuretosupportongoingmodeltraining,testing,anddevelopment.

Cloudconfigurationsareprovisionedtodelivercomparableperformanceandcontinuousenterprise-scalecapacityundercommittedpricing.Allcostsarederivedfrompubliclyavailablepricesourcesatthetimeofanalysis.

1-Year,3-Year,5-YearCumulativeCosts,(USDmillions)

$180

$160

$140

$120

$100

$80

$60

$40

$20

$0

OCI(Oracle)AWS(Amazon)GCP(Google)OriginAl(Penguin)

■Year5

Undertheseconditions,thecostdifferenceisconsiderable.Overathree-year

period,thecloudoptionsareestimatedtocostanaverageofover3.5×morethantheon-premisesdeployment.Overfiveyears,thatdifferentialwidensfurther.

Thecostspreadisanillustrationofrentvsbuyeconomics.Simplyput,ifyouuse

somethingalotandit'sasignificantcost,inmostcasesyou'rebetteroffpurchasingitthanrentingit.

Ifyouchangetherequirementsandassumptions,thenumberswillchangeandsometimesthedecisionwillchangeaswell.Inthisreport,weclearlystatetherequirements/assumptions,configureon-premisesandCloudServiceProvider(CSP)environmentstobestsatisfytheconditions,thencostthemout.

May2026TheRealCostofAlInfrastructure3of31

May2026TheRealCostofAIInfrastructure4of31

TableofContents

WhyThisReport,WhyNow?

5

BuildingtheModel

6

GlobalAssumptionsandModelingBoundaries

7

Continuous24×7steady-stateoperation

ClusterUtilization

HighPerformanceInfrastructureConfiguration

NoSpotorPreemptibleCapacity

Three-YearCommitmentPricingWhereAvailable

WhatThisModelDoesNotAttempttoCapture

AICrrseline

Ii

9

StoragePerformanceBaseline

NetworkingRequirements

PerformanceEquivalence

On-PremisesConfiguration:PenguinSolutionsOriginAI

12

OriginAICostSummary(PenguinSolutions)15

AmazonWebServices(AWS)ConfigurationSummary

18

AWSCostSummary19

GoogleCloudPlatform(GCP)ConfigurationSummary

20

GCPCostSummary21

OracleCloudInfrastructure(OCI)ConfigurationSummary

22

OCICostSummary24

WhattheNumbersShow

25

AppendixA:DetailedOn-PremisesConfiguration&Costs

27

AppendixB:DetailedAWSConfiguration&Costs

29

AppendixB:DetailedGCPConfigurations&Costs

30

AppendixB:DetailedOCIConfigurations&Costs

31

May2026TheRealCostofAIInfrastructure5of31

WhyThisReport,WhyNow?

Everyonewhohasevencasuallykeptupwithbusinessortechnologyknowsthatwe’reinthemidstofahugemovetoaddAIfunctionalitytonearlyeveryaspectofbusiness.

Thekeyquestionthatmanyorganizationsarewrestlingwithnowishowtoaddthe

necessarytechnologytoenabletheirAIinitiatives.Dotheyaddnewservers,GPUs,andinfrastructuretotheirexistingdatacenter?Ordotheyrentthecapacityinthecloud?Andwhatcantheyexpecttopayineithercase?

I’vebeencuriousforawhileaboutwhatitreallycoststorunseriousIT(andnowAI)infrastructureon-premisesversusinthepubliccloud.

Cloudpricingispublic.AnyonecangointoAmazonWebServices(AWS)orGoogleCloudpricingtoolsandmodeldetailedconfigurations.Butit’smuchhardertoobtainafullybuilt-out,enterprise-sizedon-premiseswithrealpricingattached.Withoutthat,anapples-to-applescomparisonislargelyspeculation.

PenguinSolutionswasinterestedinthesamequestion.ThesurgeinAIinvestmenthas

createdawaveofinfrastructuredecisionpoints,andorganizationsareactivelydebatingwheretheirlong-termAIbackboneshouldreside.Penguinfundedthetimerequiredformetoconductadetailedcostanalysisofa248-GPUon-premisesclusterandcompareit

againstAWS,GoogleCloud,andOracleCloud.

Tobeclear,Penguinprovidedadetailedhardwareconfigurationandassociatedpricingfortheon-premisessystem.Ibuiltthecloudmodelsbyusingpubliclyavailablepricingtools.Idefinedtheworkloadassumptions,utilizationmodel,storageandnetworkrequirements,andranthecostcomparisons.Themethodologyandconclusionsinthisreportaremine.

PriorworkinHPCandinfrastructureeconomicshasgivenmeageneralsenseofhowthis

mightturnout.Theunderlyingeconomicsaren’tmysterious.Ifyou’revisitingacityfora

week,youbookahotel.Ifyou’restayingforayear,yourentanapartment.Ifyou’resettlinginforthelongterm,youbuyahouse.Thelongerandsteadierthedemand,thestrongerthe

economiccaseforownership.

Thatdoesn’tmeanhotelsarewrong.Weabsolutelyneedhotels.Wealsoneedapartmentsandhouses.Butonemodeldoesn’tfiteverysituation.ThesameistrueforAIinfrastructure.

Butanalogiesaren’tanalysis.Thepointofthisreportistomovebeyondintuitionandrunthenumbersunderclearlydefinedreal-worldconditions.

May2026TheRealCostofAIInfrastructure6of31

BuildingtheModel

Weapproachedthiscomparisonthesamewaywewouldapproachanyinfrastructureevaluation:definetheworkloadfirst,thenbuildsystemstosupportit.

Themodelassumespersistentusageandhighutilizationplus24×7operationofa

centralized248-GPUAIclusterservingmultiplebusinessunitswithinamature,globalmanufacturingorganization.Utilizationismodeledatapproximately80%,reflectingaggregatedemandacrossarangeofongoingAIinitiatives.

CloudconfigurationswerebuilttomirrorthesameGPUcount,storagecapacity,andoverallperformanceenvelope.Nospotinstancesorpreemptiblecapacityareused.

IselectedAmazonAWS,GoogleCloudPlatform,andOracleCloudInfrastructurebecausetheyhadinstanceswithx86CPUscoupledwithNVIDIAB200GPUs.MicrosoftAzuredidnothavetheseinstancesavailableduringourresearchperiod.Wealsodidnotanalyzethe

smallerAIspecialistcloudprovidersduetotheirrelativelysmallsizeandnewnesstothemarket.

Wherepublishedthree-yearcommitmentpricingwasnotavailable(AWSB200andOCI),weappliedconservativediscountassumptionsbasedonhistoricalprecedentandtypical

enterprisenegotiationpatterns.Thoseassumptionsarecalledoutexplicitlyintherelevantsections.

Five-yeartotalsarecalculatedbyextendingmonthlypricinglinearly.Whyfiveyearsinsteadofthree?BecausemajorCSPshaveextendedtheirownserverdepreciationschedulesto

fiveyears,andthePenguinSolutionson-premisesconfigurationincludesfive-yearbreak/fixservicecoverage.

Noassumptionsaremadeaboutrenegotiation,futurepricereductions,orhardwarerefreshcycles.Vendorsnegotiatepricingallthetimebasedondealsize,timing,andinternal

revenuetargets.That’sreal,butit’soutsidethescopeofthismodel.

May2026TheRealCostofAIInfrastructure7of31

GlobalAssumptionsandModelingBoundaries

Continuous24×7steady-stateoperation:

LargeglobalorganizationswithmultiplelinesofbusinesstendtohavemanyAIefforts

underwayatthesametime.Someareearlypilots,othersaretestingphases,andsomearefull-scaletrainingorretrainingworkloads.ThisclusterisdesignedprimarilyforAI

developmentandtraining.Itcouldhandleinferenceworkloadsaswell,butinferencedemandvarieswidelydependingonthemodelanddeploymentpattern,sowe’renotlookingatitspecificallyhere.

ClusterUtilizationAssumption:

Utilizationismodeledattheaggregateclusterlevel,notperindividualnode.Weareusing80%astheaverageclusterutilizationrate.Theideaistoreflecttotaldemandacross

multipleconcurrentinitiativesratherthanassumeeveryGPUisfullysaturatedatalltimes.IncentralizedHPCandenterpriseAIenvironments,sustainedutilizationinthe75–85%

rangeisn’tunusualwheninfrastructureissizedtothevalidateddemandinsteadofspeculativegrowth.Thismodelshowsoperatingreality.

HighPerformanceInfrastructureConfiguration:

AItrainingworkloadsarecomputationallyintensiveandrequiretightlyintegrated,high-performanceinfrastructure—essentiallysupercomputing-classsystems.Equivalentconfigurationsareavailablefrombothon-premisesvendorsandCSPs,andthecloudmodelswerebuiltaccordingly.

NoSpotorPreemptibleCapacity:

Theon-premisessystemisavailabletousers24×7,sotheCSPconfigurationsareheldtothesamestandard.Manyoftheworkloadsmodeledhererunfordaysorevenweeks.

Allowinginterruptionswouldrequirefrequentcheckpointing,whichaddsoverheadand

cost.Spotmarketpricingandavailabilityalsofluctuatesignificantly,makingthembettersuitedforshort-termorburstworkloadsratherthanproductionenvironmentswithongoinghighutilization.

May2026TheRealCostofAIInfrastructure8of31

Three-YearCommitmentPricingWhereAvailable:

Committedpricingimprovescostpredictabilityandreducescloudexpense.Insomecases,high-demandGPUinstancesareonlyofferedunderreservedorcommittedagreements,sothisassumptionmimicshowlargeenterprisesactuallyprocurecloudinfrastructure.

WhatThisModelDoesNotAttempttoCapture

•Softwareorapplicationlicensingbeyondbaselinesystemmanagementtools.

•Managedservices.Thesearewidelyavailableandcustomizable,buttheyare

outsidethescopeofthisinfrastructurecostcomparison.Break/fixserviceis

includedfortheon-premisessystemsinceCSPinfrastructureinherentlyincludeshardwaresupport.

•Financialtreatment(CapExvsOpEx).Depreciationschedules,taxtreatment,andcapitalstructuringvarybyorganizationandaren’tpresentedindepthhere.

•Futurepricingchanges.Allpricingisfrompubliclyavailabledataasof1Q2026.

May2026TheRealCostofAIInfrastructure9of31

AIInfrastructureConfigurationBaseline

Beforecomparingspecificvendorimplementations,it’simportanttodefinethecommonconfigurationbaselineusedacrossallscenarios.Thegoalisnottomodelahypotheticalhyperscaleenvironmentorasmallexperimentalcluster,butaserious,large-scale

enterprisedeploymentsizedtothevalidateddemand.

Thisconfigurationassumesalarge,matureorganizationconsolidatingAIdevelopmentandtrainingworkloadsintoacentralizedsharedresource.Itisintentionallylargeenoughto

exposemeaningfuleconomicdifferences,butnotsolargeastorepresenthyperscaler-scaleinfrastructure.

ClusterSizeandGPUGeneration

OnPremises&CSPConfigurationSummary

Specification

OriginAI(PenguinSolutions)

AWS

GoogleCloud

OCI

ComputePlatform

CustomGPUcluster

p6-b200.48xlarge

a4-highgpu-8g

BM.GPU.B200.8

#ComputeNodes

31

31

31

31

GPUs(Total/Node)

248/8

248/8

248/8

248/8

GPUType

NVIDIAB200

NVIDIAB200

NVIDIAB200

NVIDIAB200

CPUperNode

IntelXeon(128cores)

Intel(192vCPUs)

Intel(224vCPUs)

AMDEPYC(256cores)

MemoryperNode(GB)

2048

2048

3968

2048

LocalStorageperNode

30.72TBNVMe

30.72TBSSD

28TBSSD

54TBSSD

Interconnect

400Gb/sNDRInfiniBand

3200Gb/s(EFA,RDMA)

3600Gb/s(RDMAfabric)

3200Gb/s(RoCEv2RDMA)

SharedStorage

VASTData,2.5PBusable

FSxforLustre,2.5PBusable

ManagedLustre,2.5PBusable

OCILustre,2.5PBusable

Admin/MgmtNodes

8×dualIntel6767P

8×i4i.32xlarge

8×m3-megamem-128

8×VM.Standard.E5.Flex

DeploymentModel

On-premcluster

Persistent(24x7)

Persistent(24x7)

Persistent(24x7)

(*detailedconfigurationsincludedbelow)

Theconfigurationisbuiltaround248NVIDIAB200GPUs,deployedas31nodeswith8GPUspernode.

TheB200generationwasselectedbecauseitisawidelyavailable,production-provenGPUarchitectureofferedacrosstheselectedcloudprovidersandtheon-premisesvendor.

Atthetimeofmodeling,B300-classsystemswereemergingbutnotyetbroadlyavailable

acrossallprovidersincomparableconfigurations.UsingB200GPUsallowsforaconsistent,apples-to-applescomparison.

FutureGPUgenerationswillimproveperformanceperwattandperdollar.That

improvementbenefitsbothrentalandownershipmodels.Thefocushereisthecoststructureundersustainedutilization,notgenerationalbenchmarking.

The248-GPUscalewaschosentorepresentasubstantialbutrealisticenterprise

deployment.ItislargeenoughtosupportmultipleconcurrentAIinitiativesacrossbusiness

May2026TheRealCostofAIInfrastructure10of31

unitsandtostresstheeconomicmodelmeaningfully,butitdoesnotapproachhyperscaledatacentermagnitude.

StoragePerformanceBaseline

AItrainingworkloadsrequirehigh-throughput,parallelstoragecapableofsustaininglargedataflowstoandfromGPUnodes.StorageneedstodeliverenoughdatatokeepGPUs

highlyutilizedandavoididletime.Thebaselineconfigurationassumesstorage

performanceconsistentwithapproximately250MB/sperTiBinthecloudenvironments,alignedwithenterprise-classmanagedparallelfilesystemtiers.

On-premisesstorageisconfiguredtodelivercomparableorhigheraggregateread/writethroughputusingascale-outarchitecture.Theintentisnottooptimizeforanysingle

vendor’speakperformance,buttoensurethatallconfigurationsmeetaconsistentperformanceenvelopeappropriateforlarge-scaleAItrainingworkloads.

Capacityandthroughputweresizedtoavoidbottlenecksineitherenvironment.Storagecapacitywassetat2.5petabytesofusablespace.

ThisbaselineintentionallyfocusesonAIdevelopmentandtrainingworkloads.Inference

environmentsvarysignificantlydependingonmodelsize,deploymentpattern,andlatencyrequirements,makingthemhighlyworkload-specificandlesssuitableforgeneralizedcostmodelingatthisscale.

NetworkingRequirements

Large-scaleAItrainingdependsonhigh-bandwidth,low-latencyinterconnectsbetweenGPUnodes.Thebaselineassumesahigh-performancefabricappropriatefordistributedtrainingworkloads.

On-premisesconfigurationsutilizemodernhigh-speedinterconnecttechnologiescommoninHPC-classdeployments.Cloudconfigurationswereselectedtoprovidecomparable

intra-clusternetworkingcapabilitieswithinasingleregion.

Allscenariosassumethatthefullclusteroperateswithinasingleregionordatacenterenvironmenttoavoidcross-regionlatencypenaltiesandtomaintaintrainingefficiency.

Thisconfigurationisoptimizedfordistributed,multi-nodetrainingworkloadsratherthan

single-nodefine-tuningtasks.Whilesmallerfine-tuningjobscanbeaccommodated,the

clusterissizedandinterconnectedtosupportlarge-scalesynchronizedtrainingrunsacrossmanyGPUs.

May2026TheRealCostofAIInfrastructure11of31

PerformanceEquivalence

Allvendorconfigurationsareconstructedtomeetthesamegeneralperformanceandcapacitytargets.Thegoalisnottoclaimarchitecturalsuperiorityforanydeploymentmodel,buttodefineaconsistentbaselinesothatcostcomparisonsreflecteconomicstructureratherthanconfigurationdifferences.

Thisbaselineisarealistic,enterprise-scaleAIinfrastructureforamaturecompanywhereAImodelswillneedtobetestedandtrainedatscale.

Configuringequallycapablesystemsacrosson-premisesandcloudenvironmentsisnottrivial.Theterminology,specifications,andpackagingdiffersignificantlybetweenphysicalinfrastructureandcloud-definedinstances.EvenamongCSPs,thereislimitedcommongroundinhowinstancesandperformancecharacteristicsaredescribed.

Forthatreason,thefollowingsectionsexplicitlyoutlinehoweachconfigurationwasconstructed.Whiletheimplementationsdifferindetail,eachisdesignedtomeet

comparableperformanceexpectations.

May2026TheRealCostofAIInfrastructure12of31

On-PremisesConfiguration:PenguinSolutionsOriginAIInfrastructure

Theon-premconfigurationinthisanalysisisbasedonadetailedsystemspecificationandpricingprovidedbyPenguinSolutions.Thesystemisacentralized248-GPUAIcluster

intendedtosupportsustainedtrainingworkloadsacrossmultiplebusinessunits.Thespecificationsbelowsummarizethecompute,storage,networking,andsupportinginfrastructurecomponentsincludedinthequotedconfiguration.

ComputeInfrastructure(Summarized)

Component

Specification

GPUNode

Configuration

31nodes;248GPUstotal(8×NVIDIAB200(192GB)pernode);DualIntelXeonCPUs;2TBRAMpernode;30.73TBNVMepernode;NVIDIA

400Gb/sInfiniBand

Theclusterconsistsof31GPUnodesconfiguredfordistributedmulti-nodetraining

workloads.EachnodecombineseightB200GPUswithbalancedCPU,memory,andhigh-speedinterconnectresources.

DetailedsystemspecificationsareprovidedinAppendixA.

StorageConfiguration(Summarized)

StorageConfiguration

Component

Specification

Architecture

VASTDatascale-outstorage

Configuration

64VASTCboxes;10VASTDboxes

UsableCapacity

~2.5PB

AggregateThroughput384GB/swrite;600GB/sreadSoftwareSubscription$354,000peryear

May2026TheRealCostofAIInfrastructure13of31

Storagecapacityandthroughputweresizedtomeetthecloudstorageperformancetieratapproximately250MB/sperTiB.Theconfigurationsupportssustainedtrainingworkloadsandlargedatasetiterationcycles.

DetailedstoragespecificationsareprovidedinAppendixA.

NetworkingFabric(Summarized)

Component

Specification

Back-EndNetwork

NDRInfiniBand(400Gb/spernode);non-blocking,fullyredundant

Front-EndNetwork

400GbEthernet;non-blockingGPU/storageconnectivity;redundantarchitecture

Out-of-BandNetwork

DedicatedGigabitEthernetmanagementnetwork

ThearchitectureseparatesGPUinterconnecttraffic,storageandin-bandcommunications,andmanagementfunctions.TheGPUfabricoperatesinafullyredundant,non-blocking

topologyoptimizedfordistributedtraining.

InfiniBandisusedby55%ofthesystemsonthe

Top500

listoflargestsupercomputersand68%ofthetop100supercomputers.

DetailednetworktopologyandswitchspecificationsareprovidedinAppendixA.

Integration,Support,andClusterManagement

Theon-premisespurchasepriceincludesallrackhardware,cabling,andrelated

infrastructurerequiredtodeliveracompleteandfullyfunctioningsystem.Thequotedconfigurationisanintegrated,deployableclusterratherthanapartialbillofmaterials.

Integrationanddeliveryservicesareincludedaspartofthesystemprice.Thiscovers

systemassembly,configuration,validation,anddeploymentwithinthetargetdatacenterenvironment.

Five-yearExtendedPlatformSupportisincludedandprovides:

•24×7technicalsupportavailability

•Severity1:24×7response,one-hourtarget

May2026TheRealCostofAIInfrastructure14of31

•Severity2:9×5response,four-hourtarget

•Severity3:9×5response,eight-hourtarget

Thissupportcoveragespansthefullfive-yearmodelinghorizonusedinthisanalysis.The

configurationalsoincludesPenguin’sClusterWareAIsoftwareplatform,whichprovides

centralizedclustermonitoring,management,andoperationaltoolingacrosscompute,

storage,andnetworkingcomponents.ClusterWareAIsoftwareisbilledat$128,000peryear(invoicedmonthly)andisincludedinthetotalcostofownershipcalculationsacrossthe

five-yearperiod.

Allintegration,support,andmanagementsoftwarecostsareincludedinthetotalcostofownershipmodelandarenottreatedasoptionaladd-ons.Detailedsupportterms,servicescope,andsoftwareinclusionsareprovidedinAppendixA.

EnergyandFacilitiesAssumptions

Penguinprovidedmeasuredpowerdatafortheclusterundermaximumsystemload:538kWofITload.Forthecontinuoususagemodel,weappliedour80%averageutilizationassumptiontoconvertpeaktestpowerintoanestimatedaverageoperatingload.

AnnualEnergyCalculation(BaseCase)

•PeaktestedITload:538kW

•ModeledaverageITload(80%):538×0.80=430kW

•MarginalPUE(air-cooled):1.30

•Incrementalfacilityload:430×1.30=559kW

•Annualhours:8,760

•Annualenergy:559×8,760=4,896,840kWh

•Electricityrate:$0.12/kWh(UScommercialrateaverage,Ohio)

•Estimatedadditionalannualelectricitycost:$587,621

ThisapproachtreatstheclusterasanincrementalloadwithinanexistingdatacenterandappliesamarginalPUEtocapturecoolingandpoweroverheadattributabletotheaddedITload.

Onthefacilitiesside,thismodelassumesthatthereissufficientelectricalserviceanddatacenterfloorspacetoaccommodatetheAIinfrastructurecluster.Thisisn’ttrueinall

situations,ofcourse.Buttherearesomeroutestopursuebeforelookingforaco-location

May2026TheRealCostofAIInfrastructure15of31

deal.Largedatacentersnearlyalwayshave‘ghost’systemsthatarepoweredon,takeupfloorspace,andyethavefewornousers.Adatacenterassessmentuncoversthese

systemssotheycanbedecommissioned,andtheirpower/footprintdevotedtonewerandmoreefficientsystems.

Ifthisdoesn’tfreeupenoughresourcesforthenewcluster,thenco-locationisthebestoption,andthereisadizzyingarrayofco-locationvendorsandplansavailabletoday.

PersonnelAssumptions

Operatingacentralized248-GPUAIclusterrequiresdedicatedinfrastructureoversight.Themodelassumesthreeincrementalfull-timeequivalents(FTEs)associatedwithongoing

clusteroperations.

Thisincludes:

•Oneseniorcluster/AIarchitectatanannualfullyloadedcostof$260,000

•Twoclustersupportengineersat$125,000eachperyear

TheserolescoverAIinfrastructurearchitecture,systemmanagement,administration,andusersupport.Theadditionalstaffwillhelpsupportacentralizedsharedenvironment

servingmultiplebusinessunitsandsupportingconcurrenttrainingworkloadsacrosstheorganization.

Totalincrementalpersonnelcostis$510,000peryear,includedinthetotalcostofownershipcalculationsandextendedacrossthefive-yearmodelinghorizon.

May2026TheRealCostofAIInfrastructure16of31

On-PremisesCostSummary:PenguinSolutionsOriginAIInfrastructure

CapitalInvestmentComponent

Cost

GPUNodeInfrastructure(31nodes,248GPUs)

$18,833,366

AdministrativeNodes(8nodes)

$336,940

StorageHardware(VASTConfiguration)

$2,138,000

TotalCapitalInvestment

$21,308,306

Thecapitalinvestmentincludesallnetworkinginfrastructure,racks,cabling,integrationservices,andfive-yearextendedplatformsupport.DetailedspecificationsareprovidedinAppendixA.

AnnualOperatingCostsComponent

AnnualCost

Energy(Incremental)

$587,621

VASTSoftwareSubscription

$354,000

ClusterWareAI

$128,000

Personnel(3FTE)

$510,000

TotalAnnualOperatingCost$1,579,621

TotalCostofOwnership(OpEx+CapEx)3-YearTCO$26,047,168

5-YearTCO$29,206,410

May2026TheRealCostofAIInfrastructure17of31

AnnualCostswithDepreciationandCostShare:

AnnualCosts,3-YearStraightLineDepreciation

AnnualCosts,5-YearStraightLineDepreciation

Cost

%Total

Cost

%Total

Purchase($21,308,306)

$7,102,769

82%

$4,261,661

73%

VastStorageSWcost

$354,000

4%

$354,000

6%

ClusterWareAISoftware

$128,000

1%

$128,000

2%

Power&Coolingcosts

$587,621

7%

$587,621

10%

AdditionalPersonnelcosts

$510,000

6%

$510,000

9%

$8,682,389.47

$5,841,282.00

It’sinterestingtonotethathardwaredepreciationisbyfarthelargestportionofannual

costsbothonathree-andfive-yeardepreciationbasis.We’veheardalotinthepressaboutelectricitydemandradicallyincreasingwiththeadventof

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论