版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
April27th,2026
RiskReportingfor
Developers’InternalAIModelUse
AUTHORS
OscarDelaney–ResearchAssociate†
SambhavMaheshwari–ResearchAssociateJoeO’Brien–Researcher
TheoBearman–Researcher
OliverGuest–ResearchManager
†WorkdonewhileatIAPS
1
Abstract
FrontierAIcompaniesfirstdeploytheirmostadvancedmodelsinternally,forweeksormonthsofsafetytesting,evaluation,anditeration,beforeapossiblepublicrelease.Forexample,Anthropic
recentlydevelopedanewclassofmodelwithadvancedcyberoffense-relevantcapabilities,MythosPreview,whichwasavailableinternallyforatleastsixweeksbeforeitwaspubliclyannounced.Thisinternalusecreatesrisksthatexternaldeploymentframeworksmayfailtoaddress.
Legalframeworks,notablyCalifornia'sTransparencyinFrontierArtificialIntelligenceAct(SB53),NewYork'sResponsibleAISafetyAndEducation(RAISE)Act,andtheEU'sGeneral-PurposeAICodeofPractice,alldiscussrisksfrominternalAIuse.Theyrequirefrontierdeveloperstomakeandimplementplansforhowtomanagerisksfrominternaluse,andtoproduceinternaluseriskreportsdescribingtheirsafeguardsandanyresidualrisks.Thisguideprovidesaharmonized
standardforcompaniestoproduceinternaluseriskreportssuitableforallthreeregulatory
frameworks.ItisaddressedprimarilytoevaluationandsafetyteamsatfrontierAIdevelopers,andsecondarilytoregulatorsandauditorsseekingtounderstandwhatgoodreportinglookslike.
GiventhepaceofAIR&Dautomationandthelimitedexternalvisibilityintohowcompaniesuse
theirmostcapablemodelsinternally,regularanddetailedriskreportingmaybeoneofthefew
mechanismsavailabletoensurethattherisksfrominternalAIuseareidentifiedandmanaged
beforetheymaterialize.Wheneverasubstantiallymorecapableorriskiermodelisdeployed
internally,thedevelopershouldcreateariskreportandarguewhythemodelissafetodeploy.Westructurethereportingframeworkaroundtwothreatvectors—autonomousAImisbehaviorand
insiderthreats—andthreeriskfactorsforeach:means,motive,andopportunity.
2
ExecutiveSummary
FrontierAIcompaniesfirstdeploytheirmostadvancedmodelsinternally,forweeksormonthsofsafetytesting,evaluation,anditeration,beforeapossiblepublicrelease.Forexample,Anthropic
recentlydevelopedanewclassofmodelwithadvancedcyberoffense-relevantcapabilities,MythosPreview,whichwasavailableinternallyforatleastsixweeksbeforeitwaspubliclyannounced.
1
Thisinternaluse
2
createsrisksthatexternaldeploymentframeworksmayfailtoaddress.
Legalframeworks,notablyCalifornia'sTransparencyinFrontierArtificialIntelligenceAct(SB53),NewYork'sResponsibleAISafetyAndEducation(RAISE)Act,andtheEU'sGeneral-PurposeAICodeofPractice,alldiscussrisksfrominternalAIuse.Theyrequirefrontierdeveloperstomakeandimplementplansforhowtomanagerisksfrominternaluse,andtoproduceinternaluseriskreportsdescribingtheirsafeguardsandanyresidualrisks.
3
Thisguideprovidesaharmonized
standardforcompaniestoproduceinternaluseriskreportssuitableforallthreeregulatory
frameworks.ItisaddressedprimarilytoevaluationandsafetyteamsatfrontierAIdevelopers,andsecondarilytoregulatorsandauditorsseekingtounderstandwhatgoodreportinglookslike.
Figure1:RisktaxonomyforinternalAImodeluse.Theconjunctionofallthreeriskfactors(left)andoneorboththreatvectors(center)canleadtoanyoftheharmfuloutcomes(right).
1Anthropic,“
ProjectGlasswing:SecuringcriticalsoftwarefortheAIera
,”announcingMythosPreview,waspublishedonApril7,2026;Anthropic,"
SystemCard:ClaudeMythosPreview
,"13,confirmsthatanearlyversionofMythosPreviewwasavailableinternallyatAnthropicfromFebruary24,2026.
2Inthisdocument,"internaluse"referstobothunreleasedmodelsandinternaldeploymentsofpubliclyreleasedmodelsthatfeatureenhancedcapabilities,expandedaccess,orreducedsafetyfilters.
3NotethattheCodeofPracticeonlyappliestosignatories(althoughotherfrontierdevelopersstillneedtodemonstratecompliancewiththeAIAct).SB53andRAISErequirecompaniestodescribetheirriskmitigationpracticesandcomplywiththeplantheyoutline,butthereisnominimumstandardforthisplan.
3
Figure1summarizestherisklandscapeforinternalAImodels.
4
GiventhepaceofAIR&D
automationandthelimitedexternalvisibilityintohowcompaniesusetheirmostcapablemodels
internally,regularanddetailedriskreportingmaybeoneofthefewmechanismsavailabletoensurethattherisksfrominternalAIuseareidentifiedandmanagedbeforetheymaterialize.Wheneverasubstantiallymorecapableorriskiermodelisdeployedinternally,thedevelopershouldcreateariskreportandarguewhythemodelissafetodeploy.Table1summarizessomekeyindicatorsthe
reportshouldincludeforeachcombinationofthreatvectorandriskfactor.
KeyInternalUseRiskIndicators
ThreatVector
Means
Motive
Opportunity
AutonomousAImisbehavior
(Section4)
Advancedcapabilitybenchmarks;AIR&Dandsoftware
engineeringevals;
real-worldR&D
contributionmetrics;covertreasoning
assessments
Behavioral"honeypot"evaluations;reward
hackingincidents;
interpretabilityresults;observedmisbehaviorlogs
Internalusecasesandpermissions;monitoringandoversightsystems;securitytesting;controlevaluations
Insiderthreats(Section5)
Upliftassessmentsvs.publicmodels;
dangerouscapability
benchmarks(bio,cyber,chemical);jailbreak
resistance
Securityvetting
processes;security
violationrecords;
insiderrisk
managementstatistics
Accesscontrolsandcountsbytier;
monitoringandlogging;preventionmechanisms;socialengineering
defenses
Table1:Keyinternaluseriskindicatorstoreport.Thereportingfocuscangenerallybelimitedtothemostcapableorhighestriskinternalmodel.Insiderthreatindicators(bottomrow)areprimarily
intendedforconfidentialsubmissiontoregulatorsratherthanpublicsummary.
4Non-deliberatecapabilityfailures—suchasmodelsproducingincorrectcodebecauseataskexceedstheirabilities—areimportantbutoutsidethescopeofthisreport,whichfocusesonintentionalorgoal-directedrisks.
4
TableofContents
Abstract 1
ExecutiveSummary 2
TableofContents 4
1.WhyInternalModelsPoseDistinctiveRisks 5
ThreatVectorsandHarmfulOutcomes 6
2.LegalBasisforInternalUseRiskReports 8
3.RiskFactorsandReportingScope 10
3.1TheRiskFramework 10
3.2ReportingCoverageandCadence 11
3.3ReportingStandards 13
4.AutonomousAIMisbehavior 15
4.1Evidence:Means 15
4.2Evidence:Motive 17
4.3Evidence:Opportunity 19
5.InsiderThreats 21
5.1Evidence:Means 21
5.2Evidence:Motive 22
5.3Evidence:Opportunity 23
Conclusion 24
Acknowledgements 24
References 25
5
1.WhyInternalModelsPoseDistinctiveRisks
FrontierAIcompanies'internalmodelsposerisksdistinctfrom,andsometimesgreaterthan,risksfromexternallydeployedsystems.5
Internalmodelshaveprivilegedaccesstosensitivesystems.Internalmodels—andthe
employeeswhousethem—oftenhaveaccesstosensitivesystemssuchastraininginfrastructure,safetyevaluationpipelines,modelweights,securitycontrols,andproprietarycodebases.Asthesemodelsbecomeintegratedintocriticalworkflows,suchascodereview,securitymonitoring,andresearch,theattacksurfacetheycreateexpandscorrespondingly.
Internalmodelslackexternaloversight.Modelsusedexclusivelywithinacompanyare
developedandtestedwithlimitedexternalscrutiny.TheInternationalAISafetyReport(2025)notedthat"verylittleispubliclyknownaboutinternaldeployments."6Somemodelsmayneverbe
releasedpubliclyatall,meaningtheyneverreceivetheexternalredteaming,auditing,andpublicscrutinythataccompanyaproductlaunch.7
Internalmodelsaremorecapablethanpublicmodels.Themostcapablesystemsatany
givencompanyaretypicallydeployedinternallyforweeksormonthsbeforepublicrelease.8
OpenAI'sGPT-4,forinstance,wasusedinternallyforroughlysixmonthsbeforeitspubliclaunch.9Morerecently,however,thisinternal-onlyperiodhasshortened;GPT-5,forexample,underwent
justthreeweeksofsafetytestingbycontractors.10Anthropic'sClaudeMythosPreview—releasedtoasmallsetofcybersecuritypartnersinApril2026—wasdeployedinternallyfromlateFebruary
2026.However,Anthropicdoesnotplantomakethemodelgenerallycommerciallyavailable,on
thegroundsthatitsautonomouszero-dayvulnerabilitydiscoveryandexploitationcapabilities
would—ifbroadlyavailable—acceleratecyberoffensiveactivityagainstmajoroperatingsystemsandbrowsers.11Greatercapabilitiesentailgreaterpotentialforbothaccidentalandintentionalharm.
5AcharyaandDelaney,"
ManagingRisksfromInternalAISystems
.";Chan,"
AIModelsCanBeDangerousBeforePublic
Deployment
.";Stixetal.,"
AIBehindClosedDoors:aPrimeronTheGovernanceofInternalDeployment
.";Kwonand
Casper,"
InternalDeploymentGapsinAIRegulation
."
6Bengioetal.,"
InternationalAISafetyReport2025
,"35.
7SafeSuperintelligenceInc.,"
SafeSuperintelligenceInc
."
8DelaneyandAcharya,"
TheHiddenAIFrontier
."
9E.g.,GPT-4hadasixmonthperiodofinternaluseandsafetytesting.SeeOpenAI,"
GPT-4TechnicalReport
,"59.Morerecentmodelslikelyhavehadshorterinternaluseperiods,thoughtheexactlengthisoftenunknown.
10Notethattheremayhavebeenalongertestingperiodinternally—thereportonlysaysthatMETRwasgivenaccesstothemodelforthreeweeksbeforepublicdeployment.SeeOpenAI,"
OpenAIGPT-5SystemCard
,"41.
11Anthropic,“
ProjectGlasswing:SecuringcriticalsoftwarefortheAIera
,”announcingMythosPreview,waspublishedonApril7,2026;Anthropic,"
SystemCard:ClaudeMythosPreview
,"13,confirmsthatanearlyversionofMythosPreview
wasavailableinternallyatAnthropicfromFebruary24,2026.
6
Furthermore,insidersatfrontiercompaniesmayuseordevelophelpful-onlymodelvariants—modelversionswithoutthesafetyfilterstypicallyappliedtopublic-facingmodels.Thispractice
substantiallyincreasesthepotentialformisuse.12
ThesefeatureswillbeexacerbatedasAImodelsbecomemoredeeplyintegratedintoautomatedAIR&D:
●CompaniesmaybecomemorereluctanttopubliclydeploymodelsthatmateriallyspeedupAIR&D,lesttheyhelptheircompetitors.13Thisreluctancecouldwidenthegapbetween
internalandpublicAIcapabilities,makinginternalmodelsmoredangerous,lessvisible,andmoreattractivetargetsforexfiltrationbyadversaries.14
●ThesuperhumanpaceofautomatedinternalR&Dcouldstrainhumanoversightcapacity.
●AImodelsintendedtoautomateAIresearchwillneedextensiveaccesstocompany
codebases,traininginfrastructure,andcomputetorunexperiments.Thisprivilegedaccessincreasesthepotentialdamagefrommodelmisbehavior.15
Consequently,theinternaluseofAImodelsforautomatedR&Disacriticalfocusforriskreporting.DetailedreportingonthescopeofAIR&Dautomationcanprovidegovernmentsandthepublic
withaleadingindicatorofacceleratingAIcapabilities.
ThreatVectorsandHarmfulOutcomes
Therearetwomainthreatvectorsfrominternalmodels,bothrecognizedintheFrontierAISafety
CommitmentsagreedattheAISeoulSummit16andreflectedinthesafetyframeworksofleadingAIdevelopers,includingAnthropic,17GoogleDeepMind,18andOpenAI:19
12Forarecentexampleofhelpful-onlyvariantsusedinpractice,seeAnthropic,"
SystemCard:ClaudeMythosPreview
,"§1.1.4(notingthat"helpfulonly"snapshotswithoutsafeguardsexistalongsidethestandardmodelduringtraining)and§
(deployingahelpful-onlyMythosPreviewsnapshotinavirologyprotocoluplifttrialagainstOpus4.6-assistedandunassistedcontrols).Thesystemcardalsoreportsrunningsafeguardevasionandinfluence-operationevaluationsagainsthelpful-onlyvariantstoisolaterawcapabilities(§§4.4and8.3).
13Anthropic,"
Detectingandpreventingdistillationattacks
.";Robinson,"
AnthropicRevokesOpenAI'sAccesstoClaude
."
14AcharyaandDelaney,"
ManagingRisksfromInternalAISystems
."
15AcharyaandDelaney,"
ManagingRisksfromInternalAISystems
."
16TheFrontierAISafetyCommitments(signedby16majorAIcompaniesinSeoul)explicitlyrequireparticipantsto
"assesstherisksposedbytheirfrontiermodelsorsystemsacrosstheAIlifecycle,includingbeforedeployingthatmodelorsystem,and,asappropriate,beforeandduringtraining."SeeDepartmentforScience,Innovation,andTechnology,"
FrontierAISafetyCommitments,AISeoulSummit2024
."
17AnthropiclistsAIsystemsengaginginautonomousinternalsabotageasapotentialthreat.SeeAnthropic,
"
ResponsibleScalingPolicy
,"(v3.0,Feb.24,2026),§1(capabilitythresholdfor"High-stakessabotageopportunities");Anthropic,"
RiskReport:February2026
,"§2.1.
18Googledescribesillustrativemitigationsfor"high-stakesinternaldeploymentswherethereissignificantriskofthemodelundermininghumancontrol."SeeGoogleDeepMind,"
FrontierSafetyFrameworkVersion3.0
."
19OpenAInotesthatsafeguards"needtoberobusttobothmaliciousactors(eitherinternalorexternal)andmodelmisalignmentrisks."SeeOpenAI,"
PreparednessFramework
,"13.
7
●AutonomousAImisbehavior.Modelsmaytakeharmfulactionsontheirowninitiative.Thismayincludesubtleinterferencewithevaluations,suchassandbagging(deliberatelyunderperformingtohidecapabilities)oralignmentfaking(behavingwellduringevaluationswhilepursuingdifferentgoalsduringdeployment).Thiscouldalsoincludecovertactions,suchasmodelsattemptingtoexfiltratetheirownweightsorsabotagethetrainingof
next-generationmodels.
20
●Insiderthreats.Employeesorcontractorswithprivilegedaccessmaymisuse,exfiltrate,ormaliciouslymodifyinternalmodels,potentiallyincollaborationwithexternalactorssuchasforeigngovernmentsorcriminalorganizations.
21
InternalAImodelsamplifytraditionalinsiderthreatsintwoways:theyprovidepowerfultoolsthatcandramaticallyincreasethedamageamaliciousinsidercaninflict,andthemodelsthemselvesarehigh-valuetargetsfortheftorsabotage.
Theprincipalharmfuloutcomesthatboththreatvectorscanproducefallintothreebroadcategories:
●SabotageofinternalAIR&D.AroguehumanorAImodelcouldunderminesafetyandalignmentresearch,ormanipulatethetrainingofsubsequentgenerationsofAImodels(forexample,bypoisoningsyntheticpretrainingdata,reinforcementlearningenvironments,
evaluations,oralignmentdatatheyhelpproduce,orbycorruptingthetrainingcodeorAI-basedevaluatorsusedinthepipeline).Theseactionscouldcompoundovertimeandmaybedifficulttodetect.
22
●Directcatastrophicmisuse.AdvancedinternalmodelcapabilitiescouldbeusedbypeoplewithinAIcompaniesforbioweaponsdevelopment,large-scalecyberattacks,orotheractionscausingwidespreadharm.
23
●Modelweightexfiltration.AroguehumanorAIcouldmakeanddistributean
unauthorizedcopyofthemodelweights.Unlikesabotageandmisuse,whichinvolveharmfrommodelsbeingusedinternally,exfiltrationremovesthemodelfromthedeveloper's
controlentirely,enablingadversariestofine-tuneawaysafetymeasures,automatetheirownAIR&D,orconductlarge-scalemisusewithoutanyofthedeveloper'ssafeguards.
24
20Anthropic,"
RiskReport:February2026
.”;Anthropic,“
SabotageRiskReport:ClaudeOpus4.6
.”
21Outsiderthreats—suchasexternalhackersorstate-sponsoredactorsseekingdirectunauthorizedaccesstointernalmodels—arerelevanttotheoverallsecuritypostureoffrontierAIcompanies.Butstealinganinternally-deployedmodelislikelynoeasier(andpossiblyharder)thanstealingtheweightsofapubliclydeployedmodel.Sowedonotfocuson
outsiderthreats,exceptviacollusionwithinsiderthreats,whichisaddressedinSection5.
22Greenblattetal.,"
AIControl:ImprovingSafetyDespiteIntentionalSubversion
."
23Shevlaneetal.,"
Modelevaluationforextremerisks
.";Martin,"
OverviewofTransformativeAIMisuseRisks:WhatCould
GoWrongBeyondMisalignment
."
24Nevoetal.,"
SecuringAIModelWeights
.";Qietal.,"
Fine-tuningAlignedLanguageModelsCompromisesSafety,Even
WhenUsersDoNotIntendTo!
"
8
2.LegalBasisforInternalUseRiskReports
ThreemajorAIregulatoryframeworksinCalifornia,NewYork,andtheEUdiscussrisksfrominternaluse.AIdevelopersarerequiredtoproduceinternaluseriskreportsdescribingtheirsafeguardsandresidualrisksnotsufficientlyaddressedbysafeguards.
California’sSB53andNewYork’sRAISEAct,whichhavesubstantiallyidenticaltext,requiredevelopersoffrontierAImodelstowriteandpublish(withnecessaryredactions)afrontierAI
framework,includingprovisionsaddressingrisksfrominternaluse.Specifically,eachfrontierdevelopermustdescribeitsapproachto:
Assessingandmanagingcatastrophicriskresultingfromtheinternaluseofitsfrontiermodels,includingrisksresultingfromafrontiermodelcircumventingoversight
mechanisms.
25
Additionally,frontierdevelopersshouldreviewthe"adequacyofmitigationsaspartofthedecisiontodeployafrontiermodeloruseitextensivelyinternally."
26
Theymustalsoreportontheirinternaluserisks"everythreemonthsorpursuanttoanotherreasonableschedule."
27
TheEUframeworkhaslessclear-cutapplicability,butarguablyaddressesinternalusethrough
severalprovisions.TheEUAIActrequiresAIproviderstoassessandmitigaterisksfromthe
developmentanduseofgeneral-purposeAImodelswithsystemicrisk,
28
thoughitisnotspecifiedexplicitlywhetherinternaluseisalsoinscope.
29
TheCodeofPracticeforGeneral-PurposeAI
Models,writtenwithextensiveinputfromAIcompanies,clarifieshowAIproviderscanachieve
compliancewiththeAIAct.Itdefines“use”asincluding“useofthemodelbytheSignatoryor
otheractors”(Glossary),includes"capabilitiestoautomateAIresearchanddevelopment"amongthesourcesofsystemicriskthatsignatoriesmustassess(Appendix1.3.1),andrequiressignatoriesto"protectagainstinsiderthreats,includingintheformof(self-)exfiltrationorsabotagecarriedoutbymodels"(Appendix4.4).
30
Moreover,signatoriesoftheCodeofPracticemustprovidethe
EuropeanAIOfficewith"adescriptionofhowthemodelhasbeenusedandisexpectedtobe
25Section22757.12(a)(10)SB53andParagraph§1421.1(j)RAISEAct.SeeCaliforniaStateLegislature,"
SB-53
TransparencyinFrontierArtificialIntelligenceAct
."
26Section22757.12(a)(4)SB53andParagraph1421.1(d)RAISEAct.
27Section22757.12(d)SB53andParagraph1422.2RAISEAct.
28EUAIActArticle55.
29ForadetailedlegalanalysisseePistillo,"
InternalDeploymentintheEUAIAct
."
30Samwaldetal.,"
CodeofPracticeforGeneral-PurposeAIModels
."
9
used,includingitsuseinthedevelopment,oversight,and/orevaluationof[other]models"(Measure7.1)and"adescriptionofallsecuritymitigationsimplemented"(Measure7.3).
31
AIdevelopersareleftwithsignificantdiscretionaboutexactlywhattoincludeininternaluseriskreports,especiallyintheminimalSB53andRAISElaws.Thisguidefillsthatgapbyofferingaharmonizedreportingstandardcompatiblewithallthreelegalframeworks.
32
Itaimstosetoutgold-standardreportingpractices:recommendationsthatgobeyondthelegal
minimuminseveralrespects,drawingonandextendingexistingindustrypractice.Manyofthe
recommendationsinSections4and5reflectassessmentsthatleadingdeveloperssuchas
Anthropichavepublishedintheirriskreports;
33
formalizingthemintoareportingtemplatepromotesconsistencyandenablescross-companycomparison.Wherethisguiderecommendsreporting
thatgoesbeyondcurrentindustrypractice—particularlyaroundinsiderthreats—itdoesso
becausethesearesafetygaps,andjustifiestheserecommendationsonsafetygroundsintherelevantsections.
Thisguidefocusesonperiodicriskreportsratherthanincidentreporting,butdevelopersshouldalsoensuretheirincidentreportingprocessesfollowstandardsandbestpractices.
34
Finally,
developersandregulatorsshouldnotethatthetwothreatvectorsaddressedinthisguidehavedifferentdisclosuresensitivities.Informationaboutinsiderthreatmitigations(Section5)mayitselfprovidearoadmapformaliciousinsidersifmadepublic;thisguidethereforerecommendsthatSection5reportingbesubmittedtoregulatorsinaconfidentialannex,separatelyfromtheAI
misbehaviorreportinginSection4,whichismoresuitableforpublicsummary.
31Samwaldetal.,"
CodeofPracticeforGeneral-PurposeAIModels
."
32Notethatthisisnotlegalguidance,andAIdevelopersshouldseeklegalcounseloncompliancewithanyrelevantlegislation!
33Anthropic,"
RiskReport:February2026
."
34SB53(Section22757.13)requiresreportingcriticalsafetyincidentswithin15days,orwithin24hoursifthereisan
imminentriskofdeathorseriousinjury.Moreover,theCodeofPractice(Measure9.3)requiresseriousincidentreportingwithin2to15daysdependingonseverity.
10
3.RiskFactorsandReportingScope
Havingestablishedwhyinternaluseposesdistinctiverisks(Section1)andthelegalbasisfor
reportingonthem(Section2),thissectionturnstohowcompaniesshouldstructuretheirreports.ItintroducestherisktaxonomyusedinSections4and5,andthendiscussesthescopeand
standardsforinternaluseriskreports.
3.1TheRiskFramework
AIdevelopersshouldaimtobuilda"safetycase,"i.e.,astructuredargument,supportedby
evidence,thattheAIsystemisacceptablysafe.35Asafetycaseneednotbeaformalproofof
safety;rather,itisacomprehensiveassessmentoftheavailableevidence,organizedtomakethereasoningtransparentandcontestable.36ThisapproachalignswiththeEUCodeofPractice's
requirementfora"systemicriskacceptancedetermination"with"detailedjustificationforwhythesystemicrisksstemmingfromthemodelareacceptable"(Measure7.2).Forrisksfrominternal
models,thisrequiresseparateargumentsfortheautonomousAImisbehaviorandtherogueinsiderthreatpathways.Tomakeastrongsafetycase,developersshoulddiscusseachofthreerisk
factorsacrossbothautonomousAIandhumaninsiderthreatmodels(seeTable1):37
●Means:Doesthemodelorinsiderhavethetechnicalcapabilitiestocausetherelevant
harm,includingtheabilitytocarryouttheharmfulactioncovertly,evadingmonitoringandoversightsystems?38
●Motive:Isthemodelorinsiderinclinedtocauseharm?
●Opportunity:Couldthemodelorinsidergetthechancetocauseharmundetected,giventheaccesscontrols,permissions,monitoring,andothersafeguardsinplace?
Safetyargumentsthatrestontheabsenceofmeansarethemostlikelytoerodeovertime,asmodelsbecomemorecapable.Argumentsrelyingonopportunity—thataccesscontrolsandmonitoringwouldpreventordetectharmfulactions—arestrongestwhenregularlyvalidated
throughadversarialtesting,sincecontrolscanfailorbecircumventedinunforeseenways.Inthelongrun,asmodelsbecomemorecapableandmoredeeplyintegratedintocriticalsystems,
motive-basedsafetyargumentswilllikelybecomethemostimportant.Sufficientlycapableand
35Clymeretal.,"
SafetyCases:HowtoJustifytheSafetyofAdvancedAISystems
."
36Notethat“safetycase,”asusedinsomeindustries,mayrefertoahigherstandardofproof.
37Thisthree-factorriskmodelisdrawnfromMETR,forthcoming.
38Foramodel,thismayincludestealth-forwardcapabilitieslikesteganography(
Zolkowskietal.,“EarlySignsof
SteganographicCapabilitiesinFrontierLLMs
.”)orlatent-spacereasoning(Haoetal.,“
TrainingLargeLanguageModelsto
ReasoninaContinuousLatentSpace
.”).
11
broadlydeployedmodelsmayhaveboththemeansandtheopportunitytocauseharmshouldtheychooseto,andrelyingsolelyoncontrolsandsafeguardstoreininapowerfulmisalignedmodelwouldbeafragiledefense.
39
Toensuredefenseindepth,AIdevelopersshouldprovideevidenceandargumentsacrossmeans,motive,andopportunity,mitigatingtheriskifnew
evidenceunderminesanysinglecomponent.
Importantly,asafetycaseforinternaldeploymentdoesnotrequiredemonstratingthatamodel
poseszerorisk.Rather,itshouldarguethattheinternaldeploymentdoesnotposemeaningfullygreatermarginalriskrelativetopubliclyavailableAIsystemsandothertools.
40
Thisframing
acknowledgesthatsomebaselineriskalreadyexistsfrompubliclydeployedmodels,andfocusesthesafetycaseontheadditionalriskcreatedbytheinternaldeployment'senhancedcapabilities,expandedaccess,orreducedsafetyfilters—exactlythefeaturesthatdistinguishinternalusefrompublicdeployment.ManyoftheevidencecategoriesinSections4and5arethereforemost
informativewhenreportedascomparisonsagainstpubliclyavailablemodels.
3.2ReportingCoverageandCadence
AIdevelopersshouldfocusonreportingrisksfromtheirhighest-riskinternalsystems,withthe
comprehensivenessofreportingscalingwiththemagnitudeofmodelcapabilityadvancessincethelastriskreport.Quarterlyreportsserveasabackstopincasemanysmallchangescollectively
constitutealargeupdatetotheriskprofile.Figure2presentsthedecisionlogicbywhichAI
developersshoulddeterminewhatreportstomaketotherelevantregulatorabouttheirinternalAIuse.
39Greenblattetal.,"
AIControl:ImprovingSafetyDespiteInten
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 精细化工产品项目风险评估报告
- 隧道围岩回填注浆专项施工方案
- 兽用原料药项目节能评估报告
- 企业数据安全全流程管理规程
- 换热站竣工验收报告
- 化学木浆项目经济效益和社会效益分析报告
- 氯丙烯项目环境影响报告书
- 高性能涂料项目风险评估报告
- 锅炉房安全生产现场处置方案
- 公园绿地土壤改良与水肥管理方案
- 2026福建福州市仓山区市场监督管理局编外人员招聘2人笔试参考题库及答案详解
- 矿井巷道维修安全技术措施培训课件
- 浙江金华市2025-2026学年高二下学期6月期末英语试题(含答案无听力原文无音频)
- 2026年四川省省直机关遴选和选调公务员申论+行政职业能力测验+综合知识复习题及答案
- 冷却塔技术参数及选型手册
- 2026供热考试题库及答案解析
- 2026年中级注册安全工程师《其他安全实务》考前冲刺测试卷(黄金题型)附答案详解
- 甲状腺功能减退的诊断与替代治疗
- 2026年北京市海淀区中考数学一模试卷(含答案)
- 南京市2025江苏南京航空航天大学自动化学院劳务派遣岗位招聘2人笔试历年参考题库典型考点附带答案详解
- 2025年中科大入学考试真题及答案解析PDF
评论
0/150
提交评论