英伟达-自动驾驶的Chat GPT 时刻 The ChatGPT Moment for Autonomous Driving 2026_第1页
英伟达-自动驾驶的Chat GPT 时刻 The ChatGPT Moment for Autonomous Driving 2026_第2页
英伟达-自动驾驶的Chat GPT 时刻 The ChatGPT Moment for Autonomous Driving 2026_第3页
英伟达-自动驾驶的Chat GPT 时刻 The ChatGPT Moment for Autonomous Driving 2026_第4页
英伟达-自动驾驶的Chat GPT 时刻 The ChatGPT Moment for Autonomous Driving 2026_第5页
已阅读5页,还剩45页未读, 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

DVIDIA

TheChatGPTMoment

forAutonomousDriving

MarcoPavone,Sr.DirectorofAutonomousVehicleResearchAlpamayoSummit@CVPR2026|June4,2026

2nVIDIA

AVrepresentsoneofthebestexamplesofPhysicalAltodate

ProgressinAVhasbeenmind-blowing

2005

2026

2nVIDIA

AVrepresentsoneofthebestexamplesofPhysicalAltodate

ProgressinAVhasbeenmind-blowing

2005

2026

KeyPillarsofNVIDIAAutonomousVehicleSolutions

AfullstackprogramtargetingL4autonomy|Billionsofinvestmentover10years

OrinThor

Scalability/OneChipArchitecture

Al/DLInfrastructuretoEdge

DriveSDK-OS,Tools,DriveWorks

Simulation

DRIVEAVStack&Fleet

AlEngineering

FunctionalSafety

AVResearch

nVIDIA

FoundationModelsforPhysicalAl

Afoundationmodelisanymodelthatistrainedonbroaddata(usingself-supervisionatinternetscale)andthatcanbeadaptedtoawiderangeofdownstreamtasks

CosmosWorldFoundationModels(WFMs)

Embed|Predict|Transfer|Reason

CosmosPredict

PhysicalAlSynthesis

CosmosTransfer

PhysicalAl

Understanding

CosmosEmbed

Cosmos

Reason

"Acarovertakingatruck"015-021

GeneratefutureworldstatesfrommultimodalinputsForpost-trainingtodevelopspecializedphysical

Almodels

GeneratePhysics-groundedphotorealoutputsconditionedonstructuralinputsorgroundtruthfromNVIDIAOmniverse

ForsyntheticdatagenerationwithNVIDIAOmniverse

Ajointvideo-textembeddertailoredforphysicalAlFortext-to-videoretrieval,inversevideosearch,

semanticdeduplication,anddatacuration

Chain-of-thoughtreasoningtoprovidethebestresponse

Forout-of-theboxphysicalAlreasoningandpost-trainingforembodiedAlmodels

nVIDIA

FoundationModelsDrivingEnd-to-EndPhysicalAIDevelopment

CosmosWorldFoundationModels(WFMs)Embed|Predict|Transfer|Reason

Long-tailgeneralizationviacommon-sensereasoning

On-vehicleAVstack

FM-empoweredE2Estacks

Interactivevehicleinterface

Explainability:fromvisiontoaction

Run-timemonitoring

AVdevelopmentprocesses

Datageneration

Datacuration

Sensorandtrafficsimulation

Rateestimation&safetyvalidation

Large,intelligentmodelsrunninginthecloudSmall,efficientmodelsrunningonboard

TheEmergenceof

EmbodiedReasoning

Background:AlpamayoVision-Action(VA)Model

FoundationforAlpamayoproduction(NDASE2Eprogram)

Behavior

Cloning

ReconstructionLosses

Action:AutoregressiveTrajectoryDecoder

TrajectoryDecoder

VideoDecoder

(training-only)

·Next-token-predictiontrainingpipelineonthefullmultimodalsequence(images,actions)

TokenProcessing:LLMBackbone

Llama2Backbone

·Trainedoverlarge-scaleAVdata

·Runtime-friendlysizeforembeddeddeploymentbydesign

ViTEncoder

Vision:SpatiotemporalContextEncoder

·Efficientmulti-camera,multi-frametokenizationtoreducetokensequencelengths

Egohistory

Multi-camera,

multi-frameimages

nVIDIA

Wu,AcceleratethefutureofAl-definedvehiclesandautonomousdriving,GTCTalk2025

nVIDIA

AlpamayoVAinAction

Videos2xspeed

Unprotectedleft,withVRUcrossing,

thennudgingforVRUgettingintovehicle

withopendoor

Creepingandyieldingatminortomajoroccludedleftandstopsign

NudgingforDPV,slowingforspeedbump,nudgingfortrashcan

Nudgingforaparkingcar,slowing/waiting

foranoncomingbusinourlane,stopping

atastopsignandtakingoff

Lanechangewithdeadlineintostopandgotrafficforupcomingconstructionzone

NudgeforVRU,yieldbehindparkingtruck,andnudgeintooncominglane

FromAlpamayoVAtoAlpamayoVLA

Assessingthebenefitsofinternet-scalepre-training

SupervisedFine-Tuning

ReconstructionLosses

BehaviorCloning

Action:MultimodalTrajectoryDecoder

TrajectoryDecoder

·Trainingwithautolabeledmeta-actionstoimprove

trajectorymulti-modalityandinterpretability

VideoDecoder

(training-only)Meta-Actions

Language:Internet-PretrainedLLMBackbone

●Pretrainedoverinternet-scaledata

·Distilledintoaruntime-friendlysizeforembeddeddeployment

VILALLMBackbone

(distilled)

Text

Encoder

ViTEncoder

Vision:SpatiotemporalContextEncoder

Egohistory

·Efficientmulti-sensortokenizationtoreducetokensequencelengths

·Nativescalingtohigherresolutionsandsensorcounts

"In400feet,turnright”

Multi-camera,Navigation

multi-frameimages

AlpamayoVLAConvergesFasterandDrivesBetter

Keyresult:Languagepre-trainingsynergizeswellwiththetaskofdriving

·Languagepre-trainingprovidesfasterconvergenceandleadstoabettersteadystatecomparedtotrainingfromscratch,evaluatedon25,600scenariosacrosstheUSandEU*

·Closed-loopevaluationcorroboratestheresults

1.61.51.41.3

1.2

1.1

1

NoLLMpre-training

Step

LLMpre-trained

20k40k60k80k

Onthevalueoflanguagepre-training

*Notthefirsttimewehaveseensucharesult:Mao,Ye,Qian,Pavone,Wang.ALanguageAgentforAutonomousDriving.COLM2024

AlpamayoReasoningVLA:ScalableArchitectureforReasoningVLAs

Action:Reasoning-informedTrajectoryDiffusion

·Efficientaction-experttrajectorydecodingvialightweightconditionalflow-matching

·RLpost-trainingtoimprovereasoning-actionalignmentandactionquality

Reasoning:Internet-PretrainedReasoningBackbone

·Pretrainedtoreasonoverinternet-scaledata

·RLwithverifiablerewardsimprovecausalreasoning

Vision:EfficientContextEncoder

·Handlesmultipleinputmodalities(cameras,text)

·Efficientmulti-camera,multi-timesteptokenizationtoreducetokensequencelengths

·Nativescalingtohigherresolutionsandsensorcounts

TrainingSignals

(IL,SFT,RL)

TrajectoryDecoder

Cosmos-ReasonBackbone

VisionEncoderTextEncoder

Ego

history

"Stopover"In400feet,

there”turnright”

UserNavigationCommands

Imagesfrommultiplecamerasandmultipletimesteps

CoTReasoning

Meta-Actions

nVIDIA

2025-10-28

nVIDIA.

Alpamayo-R1:BridgingReasoningandActionPredictionforGeneralizableAutonomousDrivingintheLongTail

NVIDIA¹

Abstract

End-to-endarchitecturestrainedviaimitationlearninghaveadvancedautonomousdrivingbyscal-ingmodelsizeanddata,yetperformanceremainsbrittleinsafety-criticallong-tailscenarioswheresupervisionissparseandcausalunderstandingislimited.WeintroduceAlpamayo-R1(AR1),avi-sion-language-actionmodelthatintegratesChainofCausationreasoningwithtrajectoryplanningtoenhancedecision-makingincomplexdrivingscenarios.Ourapproachisbuiltonthreekeyinnovations:

NVIDIA,Alpamayo-R1:BridgingReasoningandActionPredictionforGeneralizableAutonomousDrivingintheLongTail,2025;

/publication/2025-10alpamayo-rl

NVIDIA

Alpamayo1

Chain-of-ThoughtReasoning

Images,OverTime

Cosmos-Reason

TrainingSignals

MetaActions

Backbone

TrajectoryDecoder

Navigation

Reasoning

Stopduetopedestriansinthecrosswalk.

Alpamayo1DrivinginAlpaSim

PhysicalAlAVDataset

AlpamayoDeployedonCar

nVIDIA

Pavone,Liu,FromResearchtoProduction:HowAlpamayoAcceleratesAutonomousVehicleDevelopment,GTCTalk2026

/en-us/on-demand/session/gtc26-s81779/

NVIDIAAlpamayo

AnOpenEcosystemDesignedtoAccelerateReasoning-BasedAutonomousVehicleDevelopment

NVIDIAAlpamayoOpenPlatform

t=0.5s

Reasoning:

Slowdownduetotheleadvehicleahead

↑ContinuerightontoSjosavägenin10m

AlpamayoModels

Open10B-parameterReasoningVLAs

·Steerablebehaviorvianavigationandtextprompts

·Flexiblemulti-camerasupport

·Post-trainingscripts

AlpaSim

Opensimulationframeworkforclosed-looptestingandvalidation

·Extensiblepluginsystem

·LaunchinganAlpaSim-basedClosed-LoopE2EDrivingBenchmark

Entrypoints:

16

·

https://huggingface.co/blog/drmapavone/nvidia-alpamayo

·

https://huggingface.co/blog/drmapavone/nvidia-alpamayo-1-5

PhysicalAIDataset

MostdiverseopenAVdataset

·~1700hoursofmulti-sensordatafrom2500citiesin25countries

·HumanQA'dreasoninglabels

·CoCauto-labelingpipeline

·LaunchingaReasoningBenchmarkfocusedonchallengingscenarios

NVIDIA

EcosystemRecognition

ThankYOUforengagingwithAlpamayo!

Computex2026BestChoiceAward

AlpamayoOpenPlatformrecognizedfor:

·EmpoweringmanufacturerstobuildAVswithhuman-likeperception,reasoning,anddecision-makingcapabilities.

·Enablingclosed-looptestingacrossdiversetrafficscenarios,weatherconditions,andedgecases.

·Fosteringarobustopen-sourceAlecosystem.

17Zero-shotapplicationtoJapanesedata!

(TierIV)

400,000+downloadsofAlpamayoModels

2,300,000+downloadsofthePhysicalAlAVDataset

3,000+starsacrossallAlpamayoGitHubrepos

2026ChallengewInnERSLeaderboardSubmitAbout

Method

Team

MMS

MMS(selected)

MMS(heavyrai.

MMS(constructio...

TTVLM-BaseQwen3.52BMD+BV0.25

STLA-MINES

5.15

4.52

6.44

5.89

4.45

4.84

KinematicsV2

A1.5

Rangers

KE:SAI

4.31

4.31

4.18

4.21

4.39

4.50

Zero-shot3rdplaceonKITScenesLongTailChallenge!②nVIDIA(KE:SAITeam)

NVIDIAAlpamayoOpenPlatform

Chain-of-CausationReasoningLabels

100%manually-verifiedlabelstobootstrapreasoningmodelR&D,evaluation,andbenchmarks

·Explorereasoningimmediately,fromtrainingtoevaluation.

ienieiaotheefesie.snteointhe

center.Unlesstheegoischanginglane8,butthemissionis"KeepLane7;sovahice1mugntnothecritical

eieceiforhnaedsteceteoisaeceleratin,but

vehicle3;

Thisvehicleisontherightside.Unlesstheegowas

chngnglanesbutthemisionisTKepLane,somaybe

sehiaveticte4unhevonue4sthntnegi

vehicle4

egosfane,Lhascuudbeapoterislconfict.

byect

ReasoningAuto-LabelingPipeline

Toolstogeneratereasoninglabelsonyourowndata

·Generatemeta-actions,identifykeyframes,and/orproducechain-of-causationreasoninglabels.

·ExtensibleAPls;useyourfavoriteserverorlocalVLM/NVlabs/alpamavo-coc-autolabeler

Recipes

EachrecipefoldercontainsitsownREADMEwithinstallationandtraininginstructions.

Recipe

Description

recipes/alpamayo1sft/

Alpamayo1supervisedfine-tuning(HuggingFaceTrainer+DeepSpeed)

recipes/alpamayo1_5_sft/

Alpamayo1.5SFT(HuggingFaceTrainer+DeepSpeed)

recipes/alpamayo1xrl/

Alpamayo1and1.5RLpost-training(Cosmos-RL/GRPO)

UtilityScripts

Script

Purpose

scripts/curate_p

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论