版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
DVIDIA
TheChatGPTMoment
forAutonomousDriving
MarcoPavone,Sr.DirectorofAutonomousVehicleResearchAlpamayoSummit@CVPR2026|June4,2026
2nVIDIA
AVrepresentsoneofthebestexamplesofPhysicalAltodate
ProgressinAVhasbeenmind-blowing
2005
2026
2nVIDIA
AVrepresentsoneofthebestexamplesofPhysicalAltodate
ProgressinAVhasbeenmind-blowing
2005
2026
KeyPillarsofNVIDIAAutonomousVehicleSolutions
AfullstackprogramtargetingL4autonomy|Billionsofinvestmentover10years
OrinThor
Scalability/OneChipArchitecture
Al/DLInfrastructuretoEdge
DriveSDK-OS,Tools,DriveWorks
Simulation
DRIVEAVStack&Fleet
AlEngineering
FunctionalSafety
AVResearch
nVIDIA
FoundationModelsforPhysicalAl
Afoundationmodelisanymodelthatistrainedonbroaddata(usingself-supervisionatinternetscale)andthatcanbeadaptedtoawiderangeofdownstreamtasks
CosmosWorldFoundationModels(WFMs)
Embed|Predict|Transfer|Reason
CosmosPredict
PhysicalAlSynthesis
CosmosTransfer
PhysicalAl
Understanding
CosmosEmbed
Cosmos
Reason
"Acarovertakingatruck"015-021
GeneratefutureworldstatesfrommultimodalinputsForpost-trainingtodevelopspecializedphysical
Almodels
GeneratePhysics-groundedphotorealoutputsconditionedonstructuralinputsorgroundtruthfromNVIDIAOmniverse
ForsyntheticdatagenerationwithNVIDIAOmniverse
Ajointvideo-textembeddertailoredforphysicalAlFortext-to-videoretrieval,inversevideosearch,
semanticdeduplication,anddatacuration
Chain-of-thoughtreasoningtoprovidethebestresponse
Forout-of-theboxphysicalAlreasoningandpost-trainingforembodiedAlmodels
nVIDIA
FoundationModelsDrivingEnd-to-EndPhysicalAIDevelopment
CosmosWorldFoundationModels(WFMs)Embed|Predict|Transfer|Reason
Long-tailgeneralizationviacommon-sensereasoning
On-vehicleAVstack
FM-empoweredE2Estacks
Interactivevehicleinterface
Explainability:fromvisiontoaction
Run-timemonitoring
AVdevelopmentprocesses
Datageneration
Datacuration
Sensorandtrafficsimulation
Rateestimation&safetyvalidation
Large,intelligentmodelsrunninginthecloudSmall,efficientmodelsrunningonboard
TheEmergenceof
EmbodiedReasoning
Background:AlpamayoVision-Action(VA)Model
FoundationforAlpamayoproduction(NDASE2Eprogram)
Behavior
Cloning
ReconstructionLosses
Action:AutoregressiveTrajectoryDecoder
TrajectoryDecoder
VideoDecoder
(training-only)
·Next-token-predictiontrainingpipelineonthefullmultimodalsequence(images,actions)
TokenProcessing:LLMBackbone
Llama2Backbone
·Trainedoverlarge-scaleAVdata
·Runtime-friendlysizeforembeddeddeploymentbydesign
ViTEncoder
Vision:SpatiotemporalContextEncoder
·Efficientmulti-camera,multi-frametokenizationtoreducetokensequencelengths
Egohistory
Multi-camera,
multi-frameimages
nVIDIA
Wu,AcceleratethefutureofAl-definedvehiclesandautonomousdriving,GTCTalk2025
nVIDIA
AlpamayoVAinAction
Videos2xspeed
Unprotectedleft,withVRUcrossing,
thennudgingforVRUgettingintovehicle
withopendoor
Creepingandyieldingatminortomajoroccludedleftandstopsign
NudgingforDPV,slowingforspeedbump,nudgingfortrashcan
Nudgingforaparkingcar,slowing/waiting
foranoncomingbusinourlane,stopping
atastopsignandtakingoff
Lanechangewithdeadlineintostopandgotrafficforupcomingconstructionzone
NudgeforVRU,yieldbehindparkingtruck,andnudgeintooncominglane
FromAlpamayoVAtoAlpamayoVLA
Assessingthebenefitsofinternet-scalepre-training
SupervisedFine-Tuning
ReconstructionLosses
BehaviorCloning
Action:MultimodalTrajectoryDecoder
TrajectoryDecoder
·Trainingwithautolabeledmeta-actionstoimprove
trajectorymulti-modalityandinterpretability
VideoDecoder
(training-only)Meta-Actions
Language:Internet-PretrainedLLMBackbone
●Pretrainedoverinternet-scaledata
·Distilledintoaruntime-friendlysizeforembeddeddeployment
VILALLMBackbone
(distilled)
Text
Encoder
ViTEncoder
Vision:SpatiotemporalContextEncoder
Egohistory
·Efficientmulti-sensortokenizationtoreducetokensequencelengths
·Nativescalingtohigherresolutionsandsensorcounts
"In400feet,turnright”
Multi-camera,Navigation
multi-frameimages
AlpamayoVLAConvergesFasterandDrivesBetter
Keyresult:Languagepre-trainingsynergizeswellwiththetaskofdriving
·Languagepre-trainingprovidesfasterconvergenceandleadstoabettersteadystatecomparedtotrainingfromscratch,evaluatedon25,600scenariosacrosstheUSandEU*
·Closed-loopevaluationcorroboratestheresults
1.61.51.41.3
1.2
1.1
1
NoLLMpre-training
Step
LLMpre-trained
20k40k60k80k
Onthevalueoflanguagepre-training
*Notthefirsttimewehaveseensucharesult:Mao,Ye,Qian,Pavone,Wang.ALanguageAgentforAutonomousDriving.COLM2024
AlpamayoReasoningVLA:ScalableArchitectureforReasoningVLAs
Action:Reasoning-informedTrajectoryDiffusion
·Efficientaction-experttrajectorydecodingvialightweightconditionalflow-matching
·RLpost-trainingtoimprovereasoning-actionalignmentandactionquality
Reasoning:Internet-PretrainedReasoningBackbone
·Pretrainedtoreasonoverinternet-scaledata
·RLwithverifiablerewardsimprovecausalreasoning
Vision:EfficientContextEncoder
·Handlesmultipleinputmodalities(cameras,text)
·Efficientmulti-camera,multi-timesteptokenizationtoreducetokensequencelengths
·Nativescalingtohigherresolutionsandsensorcounts
TrainingSignals
(IL,SFT,RL)
TrajectoryDecoder
Cosmos-ReasonBackbone
VisionEncoderTextEncoder
Ego
history
"Stopover"In400feet,
there”turnright”
UserNavigationCommands
Imagesfrommultiplecamerasandmultipletimesteps
CoTReasoning
Meta-Actions
nVIDIA
2025-10-28
nVIDIA.
Alpamayo-R1:BridgingReasoningandActionPredictionforGeneralizableAutonomousDrivingintheLongTail
NVIDIA¹
Abstract
End-to-endarchitecturestrainedviaimitationlearninghaveadvancedautonomousdrivingbyscal-ingmodelsizeanddata,yetperformanceremainsbrittleinsafety-criticallong-tailscenarioswheresupervisionissparseandcausalunderstandingislimited.WeintroduceAlpamayo-R1(AR1),avi-sion-language-actionmodelthatintegratesChainofCausationreasoningwithtrajectoryplanningtoenhancedecision-makingincomplexdrivingscenarios.Ourapproachisbuiltonthreekeyinnovations:
NVIDIA,Alpamayo-R1:BridgingReasoningandActionPredictionforGeneralizableAutonomousDrivingintheLongTail,2025;
/publication/2025-10alpamayo-rl
NVIDIA
Alpamayo1
Chain-of-ThoughtReasoning
Images,OverTime
Cosmos-Reason
TrainingSignals
MetaActions
Backbone
TrajectoryDecoder
Navigation
Reasoning
Stopduetopedestriansinthecrosswalk.
Alpamayo1DrivinginAlpaSim
PhysicalAlAVDataset
AlpamayoDeployedonCar
nVIDIA
Pavone,Liu,FromResearchtoProduction:HowAlpamayoAcceleratesAutonomousVehicleDevelopment,GTCTalk2026
/en-us/on-demand/session/gtc26-s81779/
NVIDIAAlpamayo
AnOpenEcosystemDesignedtoAccelerateReasoning-BasedAutonomousVehicleDevelopment
NVIDIAAlpamayoOpenPlatform
t=0.5s
Reasoning:
Slowdownduetotheleadvehicleahead
↑ContinuerightontoSjosavägenin10m
AlpamayoModels
Open10B-parameterReasoningVLAs
·Steerablebehaviorvianavigationandtextprompts
·Flexiblemulti-camerasupport
·Post-trainingscripts
AlpaSim
Opensimulationframeworkforclosed-looptestingandvalidation
·Extensiblepluginsystem
·LaunchinganAlpaSim-basedClosed-LoopE2EDrivingBenchmark
Entrypoints:
16
·
https://huggingface.co/blog/drmapavone/nvidia-alpamayo
·
https://huggingface.co/blog/drmapavone/nvidia-alpamayo-1-5
PhysicalAIDataset
MostdiverseopenAVdataset
·~1700hoursofmulti-sensordatafrom2500citiesin25countries
·HumanQA'dreasoninglabels
·CoCauto-labelingpipeline
·LaunchingaReasoningBenchmarkfocusedonchallengingscenarios
NVIDIA
EcosystemRecognition
ThankYOUforengagingwithAlpamayo!
Computex2026BestChoiceAward
AlpamayoOpenPlatformrecognizedfor:
·EmpoweringmanufacturerstobuildAVswithhuman-likeperception,reasoning,anddecision-makingcapabilities.
·Enablingclosed-looptestingacrossdiversetrafficscenarios,weatherconditions,andedgecases.
·Fosteringarobustopen-sourceAlecosystem.
17Zero-shotapplicationtoJapanesedata!
(TierIV)
400,000+downloadsofAlpamayoModels
2,300,000+downloadsofthePhysicalAlAVDataset
3,000+starsacrossallAlpamayoGitHubrepos
2026ChallengewInnERSLeaderboardSubmitAbout
Method
Team
MMS
MMS(selected)
MMS(heavyrai.
MMS(constructio...
TTVLM-BaseQwen3.52BMD+BV0.25
STLA-MINES
5.15
4.52
6.44
5.89
4.45
4.84
KinematicsV2
A1.5
Rangers
KE:SAI
4.31
4.31
4.18
4.21
4.39
4.50
Zero-shot3rdplaceonKITScenesLongTailChallenge!②nVIDIA(KE:SAITeam)
NVIDIAAlpamayoOpenPlatform
Chain-of-CausationReasoningLabels
100%manually-verifiedlabelstobootstrapreasoningmodelR&D,evaluation,andbenchmarks
·Explorereasoningimmediately,fromtrainingtoevaluation.
ienieiaotheefesie.snteointhe
center.Unlesstheegoischanginglane8,butthemissionis"KeepLane7;sovahice1mugntnothecritical
eieceiforhnaedsteceteoisaeceleratin,but
vehicle3;
Thisvehicleisontherightside.Unlesstheegowas
chngnglanesbutthemisionisTKepLane,somaybe
sehiaveticte4unhevonue4sthntnegi
vehicle4
egosfane,Lhascuudbeapoterislconfict.
byect
ReasoningAuto-LabelingPipeline
Toolstogeneratereasoninglabelsonyourowndata
·Generatemeta-actions,identifykeyframes,and/orproducechain-of-causationreasoninglabels.
·ExtensibleAPls;useyourfavoriteserverorlocalVLM/NVlabs/alpamavo-coc-autolabeler
Recipes
EachrecipefoldercontainsitsownREADMEwithinstallationandtraininginstructions.
Recipe
Description
recipes/alpamayo1sft/
Alpamayo1supervisedfine-tuning(HuggingFaceTrainer+DeepSpeed)
recipes/alpamayo1_5_sft/
Alpamayo1.5SFT(HuggingFaceTrainer+DeepSpeed)
recipes/alpamayo1xrl/
Alpamayo1and1.5RLpost-training(Cosmos-RL/GRPO)
UtilityScripts
Script
Purpose
scripts/curate_p
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 煤层气录井安全规范深度解析
- 大数的认识重点单元核心归纳与易错警示
- FPGA上的嵌入式系统设计实例 作者 赵峰第7章
- 创建电力工程策划与控制
- 证券公司个人年终总结12篇
- 《托福考试简介》课件
- 雨课堂学堂在线学堂云《健身理论与指导(湖北科技学院)》单元测试考核答案
- 银行电话客服工作总结13篇
- 整合证券公司个人年终总结12篇
- 副族元素性质归纳及解题分析
- 湖南省2027届高三九校联盟第一次联考语文试卷(含答案及解析)
- 【方案】2026AI 智慧工厂解决方案
- 中国银河资产2027年“新苗计划”校园招聘笔试模拟试题及答案解析
- 2026全国中小学生天文知识竞赛(小学组)历年参考题库含答案详解
- 2026年团校考试入团考试题库(含答案)
- 下肢深静脉血栓的预防和护理
- 超粗晶WC-Co硬质合金:制备工艺与高温性能的深度解析
- 1-轨道工程施工方案-八局一-新建铁路临沂临港疏港铁路工程
- 建筑安全党课:风险与防控
- DB23∕T 3534-2023 水稻耐盐碱性鉴定技术规程
- 中医阴阳五行学说课件
评论
0/150
提交评论