开发者内部AI+模型风险报告_第1页
开发者内部AI+模型风险报告_第2页
开发者内部AI+模型风险报告_第3页
开发者内部AI+模型风险报告_第4页
开发者内部AI+模型风险报告_第5页
已阅读5页,还剩52页未读 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

April27th,2026

RiskReportingfor

Developers’InternalAIModelUse

AUTHORS

OscarDelaney–ResearchAssociate†

SambhavMaheshwari–ResearchAssociateJoeO’Brien–Researcher

TheoBearman–Researcher

OliverGuest–ResearchManager

†WorkdonewhileatIAPS

1

Abstract

FrontierAIcompaniesfirstdeploytheirmostadvancedmodelsinternally,forweeksormonthsofsafetytesting,evaluation,anditeration,beforeapossiblepublicrelease.Forexample,Anthropic

recentlydevelopedanewclassofmodelwithadvancedcyberoffense-relevantcapabilities,MythosPreview,whichwasavailableinternallyforatleastsixweeksbeforeitwaspubliclyannounced.Thisinternalusecreatesrisksthatexternaldeploymentframeworksmayfailtoaddress.

Legalframeworks,notablyCalifornia'sTransparencyinFrontierArtificialIntelligenceAct(SB53),NewYork'sResponsibleAISafetyAndEducation(RAISE)Act,andtheEU'sGeneral-PurposeAICodeofPractice,alldiscussrisksfrominternalAIuse.Theyrequirefrontierdeveloperstomakeandimplementplansforhowtomanagerisksfrominternaluse,andtoproduceinternaluseriskreportsdescribingtheirsafeguardsandanyresidualrisks.Thisguideprovidesaharmonized

standardforcompaniestoproduceinternaluseriskreportssuitableforallthreeregulatory

frameworks.ItisaddressedprimarilytoevaluationandsafetyteamsatfrontierAIdevelopers,andsecondarilytoregulatorsandauditorsseekingtounderstandwhatgoodreportinglookslike.

GiventhepaceofAIR&Dautomationandthelimitedexternalvisibilityintohowcompaniesuse

theirmostcapablemodelsinternally,regularanddetailedriskreportingmaybeoneofthefew

mechanismsavailabletoensurethattherisksfrominternalAIuseareidentifiedandmanaged

beforetheymaterialize.Wheneverasubstantiallymorecapableorriskiermodelisdeployed

internally,thedevelopershouldcreateariskreportandarguewhythemodelissafetodeploy.Westructurethereportingframeworkaroundtwothreatvectors—autonomousAImisbehaviorand

insiderthreats—andthreeriskfactorsforeach:means,motive,andopportunity.

2

ExecutiveSummary

FrontierAIcompaniesfirstdeploytheirmostadvancedmodelsinternally,forweeksormonthsofsafetytesting,evaluation,anditeration,beforeapossiblepublicrelease.Forexample,Anthropic

recentlydevelopedanewclassofmodelwithadvancedcyberoffense-relevantcapabilities,MythosPreview,whichwasavailableinternallyforatleastsixweeksbeforeitwaspubliclyannounced.

1

Thisinternaluse

2

createsrisksthatexternaldeploymentframeworksmayfailtoaddress.

Legalframeworks,notablyCalifornia'sTransparencyinFrontierArtificialIntelligenceAct(SB53),NewYork'sResponsibleAISafetyAndEducation(RAISE)Act,andtheEU'sGeneral-PurposeAICodeofPractice,alldiscussrisksfrominternalAIuse.Theyrequirefrontierdeveloperstomakeandimplementplansforhowtomanagerisksfrominternaluse,andtoproduceinternaluseriskreportsdescribingtheirsafeguardsandanyresidualrisks.

3

Thisguideprovidesaharmonized

standardforcompaniestoproduceinternaluseriskreportssuitableforallthreeregulatory

frameworks.ItisaddressedprimarilytoevaluationandsafetyteamsatfrontierAIdevelopers,andsecondarilytoregulatorsandauditorsseekingtounderstandwhatgoodreportinglookslike.

Figure1:RisktaxonomyforinternalAImodeluse.Theconjunctionofallthreeriskfactors(left)andoneorboththreatvectors(center)canleadtoanyoftheharmfuloutcomes(right).

1Anthropic,“

ProjectGlasswing:SecuringcriticalsoftwarefortheAIera

,”announcingMythosPreview,waspublishedonApril7,2026;Anthropic,"

SystemCard:ClaudeMythosPreview

,"13,confirmsthatanearlyversionofMythosPreviewwasavailableinternallyatAnthropicfromFebruary24,2026.

2Inthisdocument,"internaluse"referstobothunreleasedmodelsandinternaldeploymentsofpubliclyreleasedmodelsthatfeatureenhancedcapabilities,expandedaccess,orreducedsafetyfilters.

3NotethattheCodeofPracticeonlyappliestosignatories(althoughotherfrontierdevelopersstillneedtodemonstratecompliancewiththeAIAct).SB53andRAISErequirecompaniestodescribetheirriskmitigationpracticesandcomplywiththeplantheyoutline,butthereisnominimumstandardforthisplan.

3

Figure1summarizestherisklandscapeforinternalAImodels.

4

GiventhepaceofAIR&D

automationandthelimitedexternalvisibilityintohowcompaniesusetheirmostcapablemodels

internally,regularanddetailedriskreportingmaybeoneofthefewmechanismsavailabletoensurethattherisksfrominternalAIuseareidentifiedandmanagedbeforetheymaterialize.Wheneverasubstantiallymorecapableorriskiermodelisdeployedinternally,thedevelopershouldcreateariskreportandarguewhythemodelissafetodeploy.Table1summarizessomekeyindicatorsthe

reportshouldincludeforeachcombinationofthreatvectorandriskfactor.

KeyInternalUseRiskIndicators

ThreatVector

Means

Motive

Opportunity

AutonomousAImisbehavior

(Section4)

Advancedcapabilitybenchmarks;AIR&Dandsoftware

engineeringevals;

real-worldR&D

contributionmetrics;covertreasoning

assessments

Behavioral"honeypot"evaluations;reward

hackingincidents;

interpretabilityresults;observedmisbehaviorlogs

Internalusecasesandpermissions;monitoringandoversightsystems;securitytesting;controlevaluations

Insiderthreats(Section5)

Upliftassessmentsvs.publicmodels;

dangerouscapability

benchmarks(bio,cyber,chemical);jailbreak

resistance

Securityvetting

processes;security

violationrecords;

insiderrisk

managementstatistics

Accesscontrolsandcountsbytier;

monitoringandlogging;preventionmechanisms;socialengineering

defenses

Table1:Keyinternaluseriskindicatorstoreport.Thereportingfocuscangenerallybelimitedtothemostcapableorhighestriskinternalmodel.Insiderthreatindicators(bottomrow)areprimarily

intendedforconfidentialsubmissiontoregulatorsratherthanpublicsummary.

4Non-deliberatecapabilityfailures—suchasmodelsproducingincorrectcodebecauseataskexceedstheirabilities—areimportantbutoutsidethescopeofthisreport,whichfocusesonintentionalorgoal-directedrisks.

4

TableofContents

Abstract 1

ExecutiveSummary 2

TableofContents 4

1.WhyInternalModelsPoseDistinctiveRisks 5

ThreatVectorsandHarmfulOutcomes 6

2.LegalBasisforInternalUseRiskReports 8

3.RiskFactorsandReportingScope 10

3.1TheRiskFramework 10

3.2ReportingCoverageandCadence 11

3.3ReportingStandards 13

4.AutonomousAIMisbehavior 15

4.1Evidence:Means 15

4.2Evidence:Motive 17

4.3Evidence:Opportunity 19

5.InsiderThreats 21

5.1Evidence:Means 21

5.2Evidence:Motive 22

5.3Evidence:Opportunity 23

Conclusion 24

Acknowledgements 24

References 25

5

1.WhyInternalModelsPoseDistinctiveRisks

FrontierAIcompanies'internalmodelsposerisksdistinctfrom,andsometimesgreaterthan,risksfromexternallydeployedsystems.5

Internalmodelshaveprivilegedaccesstosensitivesystems.Internalmodels—andthe

employeeswhousethem—oftenhaveaccesstosensitivesystemssuchastraininginfrastructure,safetyevaluationpipelines,modelweights,securitycontrols,andproprietarycodebases.Asthesemodelsbecomeintegratedintocriticalworkflows,suchascodereview,securitymonitoring,andresearch,theattacksurfacetheycreateexpandscorrespondingly.

Internalmodelslackexternaloversight.Modelsusedexclusivelywithinacompanyare

developedandtestedwithlimitedexternalscrutiny.TheInternationalAISafetyReport(2025)notedthat"verylittleispubliclyknownaboutinternaldeployments."6Somemodelsmayneverbe

releasedpubliclyatall,meaningtheyneverreceivetheexternalredteaming,auditing,andpublicscrutinythataccompanyaproductlaunch.7

Internalmodelsaremorecapablethanpublicmodels.Themostcapablesystemsatany

givencompanyaretypicallydeployedinternallyforweeksormonthsbeforepublicrelease.8

OpenAI'sGPT-4,forinstance,wasusedinternallyforroughlysixmonthsbeforeitspubliclaunch.9Morerecently,however,thisinternal-onlyperiodhasshortened;GPT-5,forexample,underwent

justthreeweeksofsafetytestingbycontractors.10Anthropic'sClaudeMythosPreview—releasedtoasmallsetofcybersecuritypartnersinApril2026—wasdeployedinternallyfromlateFebruary

2026.However,Anthropicdoesnotplantomakethemodelgenerallycommerciallyavailable,on

thegroundsthatitsautonomouszero-dayvulnerabilitydiscoveryandexploitationcapabilities

would—ifbroadlyavailable—acceleratecyberoffensiveactivityagainstmajoroperatingsystemsandbrowsers.11Greatercapabilitiesentailgreaterpotentialforbothaccidentalandintentionalharm.

5AcharyaandDelaney,"

ManagingRisksfromInternalAISystems

.";Chan,"

AIModelsCanBeDangerousBeforePublic

Deployment

.";Stixetal.,"

AIBehindClosedDoors:aPrimeronTheGovernanceofInternalDeployment

.";Kwonand

Casper,"

InternalDeploymentGapsinAIRegulation

."

6Bengioetal.,"

InternationalAISafetyReport2025

,"35.

7SafeSuperintelligenceInc.,"

SafeSuperintelligenceInc

."

8DelaneyandAcharya,"

TheHiddenAIFrontier

."

9E.g.,GPT-4hadasixmonthperiodofinternaluseandsafetytesting.SeeOpenAI,"

GPT-4TechnicalReport

,"59.Morerecentmodelslikelyhavehadshorterinternaluseperiods,thoughtheexactlengthisoftenunknown.

10Notethattheremayhavebeenalongertestingperiodinternally—thereportonlysaysthatMETRwasgivenaccesstothemodelforthreeweeksbeforepublicdeployment.SeeOpenAI,"

OpenAIGPT-5SystemCard

,"41.

11Anthropic,“

ProjectGlasswing:SecuringcriticalsoftwarefortheAIera

,”announcingMythosPreview,waspublishedonApril7,2026;Anthropic,"

SystemCard:ClaudeMythosPreview

,"13,confirmsthatanearlyversionofMythosPreview

wasavailableinternallyatAnthropicfromFebruary24,2026.

6

Furthermore,insidersatfrontiercompaniesmayuseordevelophelpful-onlymodelvariants—modelversionswithoutthesafetyfilterstypicallyappliedtopublic-facingmodels.Thispractice

substantiallyincreasesthepotentialformisuse.12

ThesefeatureswillbeexacerbatedasAImodelsbecomemoredeeplyintegratedintoautomatedAIR&D:

●CompaniesmaybecomemorereluctanttopubliclydeploymodelsthatmateriallyspeedupAIR&D,lesttheyhelptheircompetitors.13Thisreluctancecouldwidenthegapbetween

internalandpublicAIcapabilities,makinginternalmodelsmoredangerous,lessvisible,andmoreattractivetargetsforexfiltrationbyadversaries.14

●ThesuperhumanpaceofautomatedinternalR&Dcouldstrainhumanoversightcapacity.

●AImodelsintendedtoautomateAIresearchwillneedextensiveaccesstocompany

codebases,traininginfrastructure,andcomputetorunexperiments.Thisprivilegedaccessincreasesthepotentialdamagefrommodelmisbehavior.15

Consequently,theinternaluseofAImodelsforautomatedR&Disacriticalfocusforriskreporting.DetailedreportingonthescopeofAIR&Dautomationcanprovidegovernmentsandthepublic

withaleadingindicatorofacceleratingAIcapabilities.

ThreatVectorsandHarmfulOutcomes

Therearetwomainthreatvectorsfrominternalmodels,bothrecognizedintheFrontierAISafety

CommitmentsagreedattheAISeoulSummit16andreflectedinthesafetyframeworksofleadingAIdevelopers,includingAnthropic,17GoogleDeepMind,18andOpenAI:19

12Forarecentexampleofhelpful-onlyvariantsusedinpractice,seeAnthropic,"

SystemCard:ClaudeMythosPreview

,"§1.1.4(notingthat"helpfulonly"snapshotswithoutsafeguardsexistalongsidethestandardmodelduringtraining)and§

(deployingahelpful-onlyMythosPreviewsnapshotinavirologyprotocoluplifttrialagainstOpus4.6-assistedandunassistedcontrols).Thesystemcardalsoreportsrunningsafeguardevasionandinfluence-operationevaluationsagainsthelpful-onlyvariantstoisolaterawcapabilities(§§4.4and8.3).

13Anthropic,"

Detectingandpreventingdistillationattacks

.";Robinson,"

AnthropicRevokesOpenAI'sAccesstoClaude

."

14AcharyaandDelaney,"

ManagingRisksfromInternalAISystems

."

15AcharyaandDelaney,"

ManagingRisksfromInternalAISystems

."

16TheFrontierAISafetyCommitments(signedby16majorAIcompaniesinSeoul)explicitlyrequireparticipantsto

"assesstherisksposedbytheirfrontiermodelsorsystemsacrosstheAIlifecycle,includingbeforedeployingthatmodelorsystem,and,asappropriate,beforeandduringtraining."SeeDepartmentforScience,Innovation,andTechnology,"

FrontierAISafetyCommitments,AISeoulSummit2024

."

17AnthropiclistsAIsystemsengaginginautonomousinternalsabotageasapotentialthreat.SeeAnthropic,

"

ResponsibleScalingPolicy

,"(v3.0,Feb.24,2026),§1(capabilitythresholdfor"High-stakessabotageopportunities");Anthropic,"

RiskReport:February2026

,"§2.1.

18Googledescribesillustrativemitigationsfor"high-stakesinternaldeploymentswherethereissignificantriskofthemodelundermininghumancontrol."SeeGoogleDeepMind,"

FrontierSafetyFrameworkVersion3.0

."

19OpenAInotesthatsafeguards"needtoberobusttobothmaliciousactors(eitherinternalorexternal)andmodelmisalignmentrisks."SeeOpenAI,"

PreparednessFramework

,"13.

7

●AutonomousAImisbehavior.Modelsmaytakeharmfulactionsontheirowninitiative.Thismayincludesubtleinterferencewithevaluations,suchassandbagging(deliberatelyunderperformingtohidecapabilities)oralignmentfaking(behavingwellduringevaluationswhilepursuingdifferentgoalsduringdeployment).Thiscouldalsoincludecovertactions,suchasmodelsattemptingtoexfiltratetheirownweightsorsabotagethetrainingof

next-generationmodels.

20

●Insiderthreats.Employeesorcontractorswithprivilegedaccessmaymisuse,exfiltrate,ormaliciouslymodifyinternalmodels,potentiallyincollaborationwithexternalactorssuchasforeigngovernmentsorcriminalorganizations.

21

InternalAImodelsamplifytraditionalinsiderthreatsintwoways:theyprovidepowerfultoolsthatcandramaticallyincreasethedamageamaliciousinsidercaninflict,andthemodelsthemselvesarehigh-valuetargetsfortheftorsabotage.

Theprincipalharmfuloutcomesthatboththreatvectorscanproducefallintothreebroadcategories:

●SabotageofinternalAIR&D.AroguehumanorAImodelcouldunderminesafetyandalignmentresearch,ormanipulatethetrainingofsubsequentgenerationsofAImodels(forexample,bypoisoningsyntheticpretrainingdata,reinforcementlearningenvironments,

evaluations,oralignmentdatatheyhelpproduce,orbycorruptingthetrainingcodeorAI-basedevaluatorsusedinthepipeline).Theseactionscouldcompoundovertimeandmaybedifficulttodetect.

22

●Directcatastrophicmisuse.AdvancedinternalmodelcapabilitiescouldbeusedbypeoplewithinAIcompaniesforbioweaponsdevelopment,large-scalecyberattacks,orotheractionscausingwidespreadharm.

23

●Modelweightexfiltration.AroguehumanorAIcouldmakeanddistributean

unauthorizedcopyofthemodelweights.Unlikesabotageandmisuse,whichinvolveharmfrommodelsbeingusedinternally,exfiltrationremovesthemodelfromthedeveloper's

controlentirely,enablingadversariestofine-tuneawaysafetymeasures,automatetheirownAIR&D,orconductlarge-scalemisusewithoutanyofthedeveloper'ssafeguards.

24

20Anthropic,"

RiskReport:February2026

.”;Anthropic,“

SabotageRiskReport:ClaudeOpus4.6

.”

21Outsiderthreats—suchasexternalhackersorstate-sponsoredactorsseekingdirectunauthorizedaccesstointernalmodels—arerelevanttotheoverallsecuritypostureoffrontierAIcompanies.Butstealinganinternally-deployedmodelislikelynoeasier(andpossiblyharder)thanstealingtheweightsofapubliclydeployedmodel.Sowedonotfocuson

outsiderthreats,exceptviacollusionwithinsiderthreats,whichisaddressedinSection5.

22Greenblattetal.,"

AIControl:ImprovingSafetyDespiteIntentionalSubversion

."

23Shevlaneetal.,"

Modelevaluationforextremerisks

.";Martin,"

OverviewofTransformativeAIMisuseRisks:WhatCould

GoWrongBeyondMisalignment

."

24Nevoetal.,"

SecuringAIModelWeights

.";Qietal.,"

Fine-tuningAlignedLanguageModelsCompromisesSafety,Even

WhenUsersDoNotIntendTo!

"

8

2.LegalBasisforInternalUseRiskReports

ThreemajorAIregulatoryframeworksinCalifornia,NewYork,andtheEUdiscussrisksfrominternaluse.AIdevelopersarerequiredtoproduceinternaluseriskreportsdescribingtheirsafeguardsandresidualrisksnotsufficientlyaddressedbysafeguards.

California’sSB53andNewYork’sRAISEAct,whichhavesubstantiallyidenticaltext,requiredevelopersoffrontierAImodelstowriteandpublish(withnecessaryredactions)afrontierAI

framework,includingprovisionsaddressingrisksfrominternaluse.Specifically,eachfrontierdevelopermustdescribeitsapproachto:

Assessingandmanagingcatastrophicriskresultingfromtheinternaluseofitsfrontiermodels,includingrisksresultingfromafrontiermodelcircumventingoversight

mechanisms.

25

Additionally,frontierdevelopersshouldreviewthe"adequacyofmitigationsaspartofthedecisiontodeployafrontiermodeloruseitextensivelyinternally."

26

Theymustalsoreportontheirinternaluserisks"everythreemonthsorpursuanttoanotherreasonableschedule."

27

TheEUframeworkhaslessclear-cutapplicability,butarguablyaddressesinternalusethrough

severalprovisions.TheEUAIActrequiresAIproviderstoassessandmitigaterisksfromthe

developmentanduseofgeneral-purposeAImodelswithsystemicrisk,

28

thoughitisnotspecifiedexplicitlywhetherinternaluseisalsoinscope.

29

TheCodeofPracticeforGeneral-PurposeAI

Models,writtenwithextensiveinputfromAIcompanies,clarifieshowAIproviderscanachieve

compliancewiththeAIAct.Itdefines“use”asincluding“useofthemodelbytheSignatoryor

otheractors”(Glossary),includes"capabilitiestoautomateAIresearchanddevelopment"amongthesourcesofsystemicriskthatsignatoriesmustassess(Appendix1.3.1),andrequiressignatoriesto"protectagainstinsiderthreats,includingintheformof(self-)exfiltrationorsabotagecarriedoutbymodels"(Appendix4.4).

30

Moreover,signatoriesoftheCodeofPracticemustprovidethe

EuropeanAIOfficewith"adescriptionofhowthemodelhasbeenusedandisexpectedtobe

25Section22757.12(a)(10)SB53andParagraph§1421.1(j)RAISEAct.SeeCaliforniaStateLegislature,"

SB-53

TransparencyinFrontierArtificialIntelligenceAct

."

26Section22757.12(a)(4)SB53andParagraph1421.1(d)RAISEAct.

27Section22757.12(d)SB53andParagraph1422.2RAISEAct.

28EUAIActArticle55.

29ForadetailedlegalanalysisseePistillo,"

InternalDeploymentintheEUAIAct

."

30Samwaldetal.,"

CodeofPracticeforGeneral-PurposeAIModels

."

9

used,includingitsuseinthedevelopment,oversight,and/orevaluationof[other]models"(Measure7.1)and"adescriptionofallsecuritymitigationsimplemented"(Measure7.3).

31

AIdevelopersareleftwithsignificantdiscretionaboutexactlywhattoincludeininternaluseriskreports,especiallyintheminimalSB53andRAISElaws.Thisguidefillsthatgapbyofferingaharmonizedreportingstandardcompatiblewithallthreelegalframeworks.

32

Itaimstosetoutgold-standardreportingpractices:recommendationsthatgobeyondthelegal

minimuminseveralrespects,drawingonandextendingexistingindustrypractice.Manyofthe

recommendationsinSections4and5reflectassessmentsthatleadingdeveloperssuchas

Anthropichavepublishedintheirriskreports;

33

formalizingthemintoareportingtemplatepromotesconsistencyandenablescross-companycomparison.Wherethisguiderecommendsreporting

thatgoesbeyondcurrentindustrypractice—particularlyaroundinsiderthreats—itdoesso

becausethesearesafetygaps,andjustifiestheserecommendationsonsafetygroundsintherelevantsections.

Thisguidefocusesonperiodicriskreportsratherthanincidentreporting,butdevelopersshouldalsoensuretheirincidentreportingprocessesfollowstandardsandbestpractices.

34

Finally,

developersandregulatorsshouldnotethatthetwothreatvectorsaddressedinthisguidehavedifferentdisclosuresensitivities.Informationaboutinsiderthreatmitigations(Section5)mayitselfprovidearoadmapformaliciousinsidersifmadepublic;thisguidethereforerecommendsthatSection5reportingbesubmittedtoregulatorsinaconfidentialannex,separatelyfromtheAI

misbehaviorreportinginSection4,whichismoresuitableforpublicsummary.

31Samwaldetal.,"

CodeofPracticeforGeneral-PurposeAIModels

."

32Notethatthisisnotlegalguidance,andAIdevelopersshouldseeklegalcounseloncompliancewithanyrelevantlegislation!

33Anthropic,"

RiskReport:February2026

."

34SB53(Section22757.13)requiresreportingcriticalsafetyincidentswithin15days,orwithin24hoursifthereisan

imminentriskofdeathorseriousinjury.Moreover,theCodeofPractice(Measure9.3)requiresseriousincidentreportingwithin2to15daysdependingonseverity.

10

3.RiskFactorsandReportingScope

Havingestablishedwhyinternaluseposesdistinctiverisks(Section1)andthelegalbasisfor

reportingonthem(Section2),thissectionturnstohowcompaniesshouldstructuretheirreports.ItintroducestherisktaxonomyusedinSections4and5,andthendiscussesthescopeand

standardsforinternaluseriskreports.

3.1TheRiskFramework

AIdevelopersshouldaimtobuilda"safetycase,"i.e.,astructuredargument,supportedby

evidence,thattheAIsystemisacceptablysafe.35Asafetycaseneednotbeaformalproofof

safety;rather,itisacomprehensiveassessmentoftheavailableevidence,organizedtomakethereasoningtransparentandcontestable.36ThisapproachalignswiththeEUCodeofPractice's

requirementfora"systemicriskacceptancedetermination"with"detailedjustificationforwhythesystemicrisksstemmingfromthemodelareacceptable"(Measure7.2).Forrisksfrominternal

models,thisrequiresseparateargumentsfortheautonomousAImisbehaviorandtherogueinsiderthreatpathways.Tomakeastrongsafetycase,developersshoulddiscusseachofthreerisk

factorsacrossbothautonomousAIandhumaninsiderthreatmodels(seeTable1):37

●Means:Doesthemodelorinsiderhavethetechnicalcapabilitiestocausetherelevant

harm,includingtheabilitytocarryouttheharmfulactioncovertly,evadingmonitoringandoversightsystems?38

●Motive:Isthemodelorinsiderinclinedtocauseharm?

●Opportunity:Couldthemodelorinsidergetthechancetocauseharmundetected,giventheaccesscontrols,permissions,monitoring,andothersafeguardsinplace?

Safetyargumentsthatrestontheabsenceofmeansarethemostlikelytoerodeovertime,asmodelsbecomemorecapable.Argumentsrelyingonopportunity—thataccesscontrolsandmonitoringwouldpreventordetectharmfulactions—arestrongestwhenregularlyvalidated

throughadversarialtesting,sincecontrolscanfailorbecircumventedinunforeseenways.Inthelongrun,asmodelsbecomemorecapableandmoredeeplyintegratedintocriticalsystems,

motive-basedsafetyargumentswilllikelybecomethemostimportant.Sufficientlycapableand

35Clymeretal.,"

SafetyCases:HowtoJustifytheSafetyofAdvancedAISystems

."

36Notethat“safetycase,”asusedinsomeindustries,mayrefertoahigherstandardofproof.

37Thisthree-factorriskmodelisdrawnfromMETR,forthcoming.

38Foramodel,thismayincludestealth-forwardcapabilitieslikesteganography(

Zolkowskietal.,“EarlySignsof

SteganographicCapabilitiesinFrontierLLMs

.”)orlatent-spacereasoning(Haoetal.,“

TrainingLargeLanguageModelsto

ReasoninaContinuousLatentSpace

.”).

11

broadlydeployedmodelsmayhaveboththemeansandtheopportunitytocauseharmshouldtheychooseto,andrelyingsolelyoncontrolsandsafeguardstoreininapowerfulmisalignedmodelwouldbeafragiledefense.

39

Toensuredefenseindepth,AIdevelopersshouldprovideevidenceandargumentsacrossmeans,motive,andopportunity,mitigatingtheriskifnew

evidenceunderminesanysinglecomponent.

Importantly,asafetycaseforinternaldeploymentdoesnotrequiredemonstratingthatamodel

poseszerorisk.Rather,itshouldarguethattheinternaldeploymentdoesnotposemeaningfullygreatermarginalriskrelativetopubliclyavailableAIsystemsandothertools.

40

Thisframing

acknowledgesthatsomebaselineriskalreadyexistsfrompubliclydeployedmodels,andfocusesthesafetycaseontheadditionalriskcreatedbytheinternaldeployment'senhancedcapabilities,expandedaccess,orreducedsafetyfilters—exactlythefeaturesthatdistinguishinternalusefrompublicdeployment.ManyoftheevidencecategoriesinSections4and5arethereforemost

informativewhenreportedascomparisonsagainstpubliclyavailablemodels.

3.2ReportingCoverageandCadence

AIdevelopersshouldfocusonreportingrisksfromtheirhighest-riskinternalsystems,withthe

comprehensivenessofreportingscalingwiththemagnitudeofmodelcapabilityadvancessincethelastriskreport.Quarterlyreportsserveasabackstopincasemanysmallchangescollectively

constitutealargeupdatetotheriskprofile.Figure2presentsthedecisionlogicbywhichAI

developersshoulddeterminewhatreportstomaketotherelevantregulatorabouttheirinternalAIuse.

39Greenblattetal.,"

AIControl:ImprovingSafetyDespiteInten

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论