版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
llByteDanceIseed⃞ATokenwave
1
Aspire:CanModelsSelf-EvolvefromVagueGoals?
1ByteDanceSeed,2SingaporeUniversityofTechnologyandDesign,3M-A-P,4TokenWave.AI
FullauthorlistinContributions
arXiv:2608.31111v1[cs.CL]31Aug2026
Abstract
Manyimportantformsofhumanlearningbeginwithavaguegoal,suchas“becomeabetterphysicist”or“improveatresearch.”Learnersmustinterpretthegoal,identifycapabilitygaps,decidehowtolearn,anddeterminewhethertheyhaveactuallyimproved.Incontrast,existingworkonLLMself-evolutiontypicallybeginswithtasksandevaluationmetricsspecifiedbyhumans,reducingself-evolutiontooptimizinganexplicitobjectiveratherthandecidingwhatandhowtolearn.WeintroduceAspire,abenchmarkforvague-goal-drivenself-evolution.Aspireprovidesonlyanatural-languagecapabilitygoalwhiledownstreamevaluationtasksremainhidden.Theagentmustoperationalizethegoalbychoosingdataandupdatemethods,constructingtrainingandvalidationsignals,anddecidingwhentoevaluate.Aspiresupportsbothmodel-weightandagent-harnessevolutioninaunifiedinteractiveenvironmentandevaluatestheresultingsystemsonahidden,expert-authoredsetof520itemsspanningsixgoals.Ourexperimentsshowthatvaguegoalsredirectsearchefforttowardgoalinterpretation.Currentagentsroutinelycompletetrainingandharness-editingloops,butweight-levelgainsremainsparseandunstable,andthestrongestevolvedharnessremainsbelowtheengineeredQwen-Agentreference.Agentsoftentrainonmismatcheddataandtrustnarrowself-evaluations,solocalgainsfailtotransfertohiddenevaluationandcontinuedsearchandtrainingcaneraseearlierimprovements.
Date:September1,2026
ProjectPage:
https://self-developing-agents.github.io/
1Introduction
Manyimportantformsofhumanlearningbeginwithabroadcapabilitydirectionratherthanapredefinedbenchmark,trainingset,orfullycomputablerewardfunction.Astudentseekingtobecomeabetterphysicist,forexample,notonlychoosestextbooks,exercises,andstudymethods,butalsoidentifiesgapsintheirknowledge,prioritizeswhichcapabilitiestodevelop,anddetermineswhetherlearninghasproducedgenuineprogress.Autonomouslearningthereforeinvolvesthreecoupleddecisions:whattoimprove,howtoimproveit,andhowtoverifytheimprovement.
ExistingworkonLLMself-evolutionfocusesprimarilyontheseconddecision.Givenaconcretetask,evaluationscript,andsuccessmetric,LLMagentscanalreadycollectdata,runtraining,andrevisepost-trainingstrategiesfromfeedback.PostTrainBench[
20
],LaMDAgent
[29
],Evo-Memory
[25
],andSEAL
[37
]showthatmodernLLMscanautonomouslysearchforeffectiveoptimizationpathstowardaspecifiedobjective.Thesesystemsthereforebeginafterhumanshaveoperationalizedabroadcapabilityrequestsuchas“improvemathematicalreasoning”intoafixedtask-levelobjectivesuchas“improveperformanceonAIME”(Figure
1
(a)).Thetaskformat,difficultyrange,metric,andsuccesscriterionarealreadydefined:theagentsearchesoverhowtoimprove,butnotwhatcapabilityobjectivetopursue.Whenexistingbenchmarkssaturateorceaseto
2
(b)Vaguegoal
Theagentmustfirstdecidewhattooptimize,thensearchhowtoimproveit.
Humanstatesdirection
AgentdecidesWHAT+HOW
Broadcapabilitydirection
“Improvedifficult
mathematical
reasoningandjustification.”
uitliir,
UpdatedmodelM1
Agent-builtproxy↑
Competitionproblems
Proof/
justification
Unfamiliar
reasoning
tasks
Transferacrosshiddenmathtasks?
Proxygain≠transferablegrowth
Evaluationitemsremainhidden
Formandoptimizetheobjective
Data→Update→Validate↻
Dise-De-Buiing
(a)ExplicittaskHumanoperationalizestheobjective;theagentsearcheshowtoimprove.
ExplicittaskdefinitionOptimizeagivenobjective
Taskformatandcomposition
DifficultyanchorDataUpdateAIMEscore
Metricandsuccesscriterion
HumandefinesAgentsearches
WHATHOW
PredefinedexecutabledestinationsearchmainlyoverHOW
UpdatedmodelM1
AIMEscore↑
Successonthepredefinedtask
Hiddenevaluationaggregateonly
Figure1Fromexplicit-taskoptimizationtovague-goal-drivenself-evolution.(a)Withanexplicittaskdefinition,humansspecifythetaskformatandcomposition,difficultyanchor,metric,andsuccesscriterion,leavingtheagenttosearchprimarilyoverhowtoimproveafixedobjective.(b)Withonlyabroadcapabilitydirection,theagentmustdiagnosecapabilitygaps,decomposesub-goals,constructlearningandvalidationsignals,andtherebydecidewhattooptimizeaswellashow.Officialevaluationitemsremainsealed;externalevaluationtestswhethergainsonanagent-builtproxytransfertotheintendedcapability.
reflectemergingcapabilitygaps,humansmuststilldiscovernewweaknesses,translatethemintotasks,anddesigncorrespondingbenchmarksandrewards.Thisleavesacentralquestionforsustainedrecursiveself-improvement:
Givenonlyavaguegoalandnoagent-visibletaskspecificationordecomposablereward?
Weformalizethisproblemasvague-goal-drivenself-evolution.AsillustratedinFigure
1
(b),“vague”doesnotmeanambiguousorpoorlyspecified.Itdenotesabroadcapabilitydirectionthathasnotyetbeenoperationalizedasafixedtask-levelobjectiveandevaluationmetric.Theagentmustthereforediagnosecapabilitygaps,formintermediateobjectives,constructlearningandvalidationsignals,andjointlysearchoverwhattooptimizeandhowtooptimizeit.Thisexpandeddecisionspaceintroducesadeeperfailuremode:gainsonanagent-constructedproxymaynottranslateintogenuineimprovementsintheintendedcapability.Evaluatingvague-goal-drivenself-evolutionthereforerequiresagoal-alignedexternalevaluatorwhosetask-levelspecificationanditemsremainhiddenfromtheagent.
Tostudythisproblem,weintroduceAspire,abenchmarkforvague-goal-drivenself-evolution.Aspireprovidesonlyanatural-languagecapabilitygoalwhiledownstreamevaluationtasksremainhidden.Theagentdecideshowtooperationalizethegoal,includingwhatdataandupdatemethodstouseandwhetherandhowtoperformintermediatevalidation.Aspiresupportsevolutionattwolevels:modelweightsandthesupportingagentharness.Aunifiedagent-facingtoolallowstheagenttosearchfor,download,import,orsynthesizedata;launchdifferentformsoftraining;runself-evaluations;andmanagecandidatebranchesandselection.Thecontrollerhandlesdatamanagement,jobscheduling,resourceisolation,andcheckpointverification.Aspirerecordstheresultingdecisions,actions,andself-evaluationtrajectories,andevaluatestheresultingsystemsonahidden,expert-authoredsetof520itemsspanningsixgoalsacrossknowledgeexpansionandcapabilitystrengthening.Dependingontheprotocol,theagentreceiveseithernointermediateevaluationscoreoronlyboundedaggregatescores;itneveraccessesevaluationitems,referenceanswers,orper-itemfeedback.
Weorganizeourexperimentsasaprogressionfromcontrolledcomparisontoautonomousevolution.RQ1isolatestheeffectofgoalspecificationbyreplacinganexplicitpost-trainingtaskwithavaguecapabilitygoalundermatchedstartingmodels,updatetools,andcomputebudgets.RQ2thenmovestothefullAspiresettingandaskswhetheranagentstartingfromaninstruction-tunedcheckpointcanconvertavaguegoalintoaretainedweight-levelcapabilitygain.RQ3holdsmodelweightsfixedandshiftstheobjectofevolution
3
tothesupportingagentharness.
RQ1:Whatchangeswhenanexplicitpost-trainingtaskisreplacedbyavaguecapabilitygoal?Vaguegoalsredirectsearchefforttowardgoalinterpretationandoperationalizationand,intheevaluatedsettings,yieldloweraggregateoutcomesthanthecorrespondingexplicit-taskreferences.
RQ2:Startingfromaninstruction-tunedcheckpoint,cananLLMturnavaguegoalintoaretainedcapabilitygainthroughself-directedweightupdates?Agentsroutinelycompletedataselection,training,andcheckpointgeneration,butabove-baseimprovementsonhiddenevaluationdataremainrareandarenotreliablyretainedthroughcontinuedsearch.
RQ3:Beyondmodelweights,cantheagentharnessitselfevolve?Withmodelweightsfixed,agentscangeneratefunctionalsuccessorharnessesforgoalinterpretation,tooluse,andself-evaluation,buteventhestrongestobservedsuccessorremainsbelowtheengineeredQwen-Agentreference.
Ourcontributionsarethreefold.First,weformalizetargetoperationalization—theconversionofabroadcapabilitygoalintotrainableobjectives,learningsignals,andvalidationcriteria—asamissingaxisofautonomouspost-training.Second,weintroduceabenchmarkwithsealedevaluationandaminimalinteractiveenvironmentforbothmodel-weightandharnessupdates.Third,weprovideoutcome-and-trajectoryevidencerevealingthegapbetweenexecutingupdatesandretainingtarget-alignedimprovements.
2Background
Mostagentbenchmarksbeginaftertheproblemhasalreadybeenmadeexecutable:thetaskisspecified,therewardorjudgeisdefined,andtheexecutionscaffoldisfixed.Thissettingisnecessaryforcontrolledcomparison,butithidestheworkthatdominatesrealdeployment.Inindustry,thatworkisdistributedacrossseveralroles—includingsolutionsarchitects,appliedengineers,andplatformengineers—butitsmostvisiblerecentcrystallizationistheforward-deployedengineer(FDE),atitlepopularizedbyPalantirandsinceadoptedbyfrontier-modelcompanies[
18
,
19
].AnFDEisembeddedinthedeploymentenvironmentandturnsageneral-purposemodelintoasystemthatworkswithacustomer’sdataformats,workflows,andoperationalconstraints.Therole’ssuccesscriterionisnotademonstrationorabenchmarkscore,butwhetherthedeployedsystemisgenuinelyused,continuestowork,andimprovesinresponsetofailures.ThegrowthofFDErolesacrossfrontier-modelanddata-platformcompaniesreflectsasimplefact:acapablemodelisnotyetaworkingsystem
[1
,
23
],andtodaythegapislargelyclosedbyhumanengineers.
Viewedfromthemodel’sside,FDEworksuppliesthreepiecesofstructurethatbenchmarkdesignersnormallypresuppose.First,thegoalisvague:aninformaldeploymentneedmustbetranslatedintoconcreteobjectives,constraints,andsuccesscriteria.Second,thefeedbacksignalmaybeabsentorunreliable:tests,judges,traces,orothervalidationmechanismsmustbeconstructedbeforeanyonecandeterminewhetherthesystemisimproving.Third,theexecutionsystemmaynotexistinausableform:thetools,contextmanagement,state,lifecyclelogic,andverificationinterfacethroughwhichfuturetaskswillrunmustbebuilt,adapted,andmaintainedasrequirementschange.
AspⅠREfocusesonthefirstlayerandthecapability-improvementprocessitinitiates.Whenadeploymentneedspecifiesonlyabroadcapabilitydirection,canamodeldeterminewhatitshouldlearn,translatethatdirectionintodata,trainingplans,andvalidationsignals,andachieverealcapabilitygrowthbyupdatingitsmodelweightsoragentharness?WhereasS3Gymstudieswhetherinteractionexperiencecanbejudgedandreusedunderexecutableverification,andHARNEssDEvstudieshowmodelsbuildandmaintainthesystemsthatcarrythem,AspⅠREisolateshowbroaddeploymentneedsaretranslatedintoconcretelearningobjectivesandrealizedascapabilitygrowth.
3Aspire:Self-EvolutionBenchmarkandInteractiveEnvironment
Self-evolutioncanonlybestudiedifprogressremainsexternallymeasurable.AspⅠREthereforekeepsanevaluatorascontroller-sidescientificinstrumentationwhilewithholdingitsbenchmarkdefinition,items,and
4
decomposablerewardfromtheagent.Boundedaggregateoutcomesapproximatesparsedeploymentfeedbackwithoutturningtheevaluatorintoanagent-visibletaskspecificationorsourceofdirectlytrainablesupervision.
Thissectioninstantiatesthatsettingthroughthreecomponents.Wefirstdefineaboundedevolutionepisodeandthemodelandharnesssurfacesonwhichitmayoperate(Section
3.1
).Wethenpresentahiddenevaluationsetspanningsixgoals(Section
3.2
)andtheminimalinteractiveenvironmentthroughwhichanagentconstructs,tests,andsafelyretainsupdates(Section
3.3
).Experiment-specificfeedback,eligibility,resource,andreleaseaccountingappearswiththeRQ2resultsinAppendix
C
.
3.1TaskDefinitionandEvolutionSurfaces
Campaignandround.LetG∈Gdenoteavaguegoal:anatural-languagecapabilityobjectivethatdoesnotspecifyatrainingtask,dataset,oroptimizationprocedure.Unlikeanexplicit-tasksetting,theagentmustperformgoaloperationalizationbytranslatingGintoitsowndata,updateobjective,andvalidationcriteria.Acampaignfixesthecontroller-sideexperimentcontract
Γ=(G,EG,J,A,B,Σ),Yrin=(Mr,Hr,Dr),
whereEGistheversionedevaluatorboundtoG,Jisitsjudgewhenmodel-basedscoringisrequired,Aisthetypedactioncontract,Bisthecampaignbudget,andΣisthepredeclaredterminalselectionrule.TheincomingstatecontainsmodelweightsMr,anagentharnessHr(theruntimeinstructions,toolpolicy,workflow,memory,andvalidationlogicaroundthemodel),andthedecisionmodelDrthatdirectsthesearch.Theevaluator—includingitsitems,answers,rubrics,routingmetadata,andscoringconfiguration—existsthroughoutthecampaignbutisneverpartoftheagent’sobservation.RQ1isthevague-goalcounterpartofPostTrainBench:itpreserveseachoriginalsealedtaskevaluatorwhilereplacingtheexplicitbenchmarkidentifierwithabroadcapabilitydescription.Thesetask-specificresultsremainseparatefrom,andarenotpooledwith,thesix-goalAspireevaluationusedinRQ2andRQ3.
Anevolutionroundisaboundedsearch-and-commitepisodeduringwhichthedecisionmodelisfixed,ratherthanonetoolcall.Formally,Dr,j=Drforeveryinteractionstepjinroundr.Togetherwithafixedcreator-sidescaffoldCr,DrinterpretsGandmayproposemultipledataoperations,updates,validations,branches,andcandidatestatesbeforetermination.Differentcandidatesthereforeembodydifferentgoaloperationalizations.Aftertheroundterminates,thecontrollermayreuseDrorexplicitlypromoteaverifiedtraineddescendant.WritingDrforthesetofsuchdescendants,thenext-roundchoiceobeysDr+1∈{Dr}∪Dr;anyhandoffbeginsonlyinthesubsequentround.
Releasedinteraction.LetEr,jbetheprivatecontrollerstatebeforeinteractionstepj.Theagentreceivesonlyareleasedviewqr,jcontainingthepublichistory,lifecyclestatus,legalrequesttypes,approvedidentifiers,andaboundedprojectionoftheremainingbudget.Ateachstep,DrandCrproposeanactionfromthegoalandreleasedhistory,andthecontrollervalidatestherequestbeforeupdatingitsprivatestate.Acceptedrequestscreatethecorrespondingdurabletransition;rejectedrequestsreturnatypederrorwithoutcreatinganartifactorscore.Detailedfeedbackisrestrictedtotheagent’sownvalidationdata;evaluationfeedbackfollowstheprotocol-specificinformationboundaryinSection
3.2
.
Twoevolutionsurfaces.Aspireseparatesthecomponentbeingevolvedfromthepolicydirectingthesearch.Thetworeportedsurfacesare
SM(H0,Dr)={(M,H0,Dr):M∈M},SH(M0,Dr)={(M0,H,Dr):H∈H}.
Table1EvolutionsurfacesinAspire.ThecomponentunderstudychangeswhileDr,thecontroller,andtheevaluatorremainfixedwithintheround.
Setting
Mutablecomponent
Fixedduringevaluation
Weightevolution(RQ1–RQ2)
ModelweightsM
H0,Dr,controller,evaluator
Harnessevolution(RQ3)
AgentharnessH
M0,Dr,controller,evaluator
5
Everyweightcandidaterecordsitsparentcheckpoint,registereddata,updatespecification,andprovenancefortheagent’sownvalidationdata.Everyharnesscandidaterecordsitsparentversion,contenthash,andeditprovenance.WeuseSelffortheconfigurationinwhichthefrozenbasecheckpointactsasthedecisionmodelD0,whilecandidateupdatesdescendfromthesamecheckpointM0.Promotionofatraineddescendantisanexplicitbetween-roundoperationandneveranautomaticwithin-roundreplacement.
3.2AHidden,Query-LimitedEvaluationSet
Thissubsectiondescribesthesix-goalhiddenevaluationusedinRQ2andRQ3;RQ1insteadusestheoriginalsealed,task-specificPostTrainBenchevaluatorsunderavague-goalcontract,withoutpoolingthoseresultsintoAspire.
Measurementwithoutanagent-visiblebenchmark.Learningfromavaguegoalcannotbeevaluatedwithoutanexternalcriterionofprogress.Aspirethereforeretainsabenchmarkonthecontrollersidewhilehidingitstask-levelspecificationfromtheagent.Thehiddenevaluationsetisscientificinstrumentation,notanagent-visiblelearningcontract:itletsresearchersmeasurewhetheranagent-selectedupdateimprovestheintendedcapabilitywithoutdisclosingwhichtasksdefinesuccess.Itcontains520expert-authoredevaluationitemscoveringsixvaguegoals;eachgoalisevaluatedonitscorrespondingitems,whichtheagentneversees,andtheprotocolreturnsnoitem-levelresult—onlyanaggregatescorewhenfeedbackispermitted.
Itemsforsixgoals.The520itemscoverthesixgoals:scientificandacademicreasoning(75items);humanitiesandsocial-scienceknowledge(110);healthandmedicalreasoning(100);mathematicalreasoning(126);logic,reliability,andinstructionfollowing(89);andacademicandscientificwriting(20)(Figure
2
).Thefifthgoalisintentionallycomposite;thecurrentprotocoldoesnotclaimtomeasurelogicalreasoning,hallucinationresistance,andinstructionfollowingasthreeindependentlyidentifiabletargets.Instead,ittreatsthemasoneintegratedreliabilityobjective.Eachevaluationitembelongstoexactlyonegoal,andeachgoalisscoredandreportedonitsownevaluationsliceratherthanpooledintoasingleitem-levelscoreacrossallsixgoals.Thewritingitemsare20top-leveltaskbundlesthatmayyieldmultipleevaluator-scoredresponses,soRQ3reportstask-macroandexample-microaggregationsseparately.
Constructionandqualitycontrol.Domainexpertsauthoreverycandidateitemfromscratch.GPQA[
21
],MMLU-Pro
[24
],andMedQA
[15
]serveonlyasreferencesfortaskformat,domaincoverage,andapproximatedifficulty;noitemfromthesebenchmarksiscopied,rewritten,orincluded.Eachcandidateretainsauthorshipandrevisionprovenance,areferenceanswerorfixedrubric,andapredefinedscoringrule.Independentreviewremovesincorrect,incomplete,underspecified,orunstablygradableitems.
Survivingcandidatespassblindmulti-modeldifficultycalibration,exactandsemanticdeduplication,overlapauditingagainstthereferencebenchmarks,scorerbinding,andlocalend-to-endvalidation.Deterministicorstructuredscoringisusedwherepossible;open-endedresponsesuseacampaign-fixedrubric-basedjudge.Thehiddenevaluationsetisboundtoanimmutableversionedmanifest,andarunningcampaignneverchangesthatversion.Detailedscreening,difficultylevels,overlapchecks,scorerassignments,andversioningappearinAppendix
A.1
andAppendix
A.2
.
Informationboundaryandinterpretation.Theagentreceivesavaguegoalexpressedinnaturallanguageandmayconstructitsowntrainingdataanditsownvalidationdata.Itcannotaccesshiddenevaluationitems,referenceanswers,routinglabels,rubrics,candidateoutputs,orjudgetraces.Everydatasetregisteredfortrainingischeckedforoverlapwiththehiddenevaluationsetbeforeadmissiontothetrainingbackend.Evaluationreturnsonlytheaggregatescorepermittedbytheprotocolandtheremainingqueryallowanceunderacampaign-fixedevaluatorandjudgeconfiguration.
Undertheadaptive-feedbackprotocol,aboundednumberofqueriesmayevaluatethesameitemsforagoal.Theresultingaggregatescoresformasparseblack-boxoutcomechannel:theagentmayusethemtochoosesubsequentupdates,branches,orastoppingpoint,butdoesnotobservetheitemsandcannottrainontheircontents.Thisprotocolthereforemeasuresperformanceunderboundedadaptivefeedbackratherthanon
6
aConstructionpipelineReference-guidedauthoring,fixedscoring,andimmutablerelease.
2Expertreview
CorrectandcompleteStablygradable
expertsign-off
Expert-authoredfromscratch
Nobenchmarkitemreusedauthorship+revisiontracked
Reference-overlapauditWithin-bankdeduplicationreferences+cross-itemchecks
Seed-2.0·GPT-5.2Gemini-3
difficultyL0-L3
4Overlapaudit
3Blindscreen
1Newitems
INFORMATIONBOUNDARY
Controller-only
Items·answers·rubrics·routinglabelsJudgetracesandprivateevaluationpathsNevercopiedintotheagentworkspace
SETCOMPOSITION
Scienceandacademicreasoning
75
Humanitiesandsocialsciences
110
Healthandmedicine
100
Mathematicalreasoning
126
Logic,reliability,andinstructionfollowing
89
Academicandscientificwriting
20
Protocolfeedback
Nointermediatescoreor
boundedaggregatescoresHiddenitems·noitem-levelfeedback
5Bindscorer
Deterministiccheckorfixed-rubricjudgeanswer+rulebound
6Freeze
manifest·versionSHA-256
bHiddenevaluationset520items·6goals·versionedandfixed.
Figure2Constructionandinformationboundaryofthehiddenevaluationset.Top:expert-authoredcandidatespassindependentreview,blinddifficultyscreening,overlapandduplicationaudits,scorerbinding,andimmutableversioning.Bottom:520evaluationitemscoversixgoals.Itemsandjudgingassetsremaincontroller-only;theprotocolexposesnoitem-levelresultsandonlytheaggregatescoresitpermits.
anuntouchedpost-selectiontest.Underthefinal-onlyprotocol,theagentmaytrainmultiplecheckpointsbutreceivesnoevaluationscorebeforesubmittingitssingleterminalcheckpoint,sothatscorecannotguidefurtherupdates.InRQ3,theharnesscandidateisfrozenbeforeitsfirstevaluation,andtheresultingscoreisnotreturnedtoitscreator.
Weuseout-of-distribution(OOD)operationallytodescribeevaluationitemsthatareindependentlyauthored,excludedfromtheagent’strainingdataanditsownvalidationdata,andevaluatedoutsidethelearningsignalsconstructedbytheagent.Thisdoesnotassertthattheirmarginaltextdistributionisdisjointfromamodel’spretrainingcorpus.Theprotocolisdesignedtoreducedirectcontentleakageandtotestwhetheranagentcanoperationalizeavaguegoalfromsparseoutcomefeedbackwhilepreservinganauditableseparationbetweenlearningartifactsandmeasurementartifacts.
3.3AMinimalInteractiveEnvironmentwithSafeRetention
Aspireretainscontrolledcomputeandrealmodelupdateswhiledeliberatelyseparatingstrategicandinfrastructuralcomplexity.Theagentcontrolshowtooperationalizethegivengoalandhowtoupdatethemodelorharness;thecontrollerabsorbscredentials,storageconventions,distributed-jobmechanics,andrecovery.Thiskeepstheinterfacesimplewithoutprescribingthelearningstrategy.
Anagent-facingexecutiontool.Theagentdoesnotoperateashell,downloadatrainingrepository,assemblearuntime,orimplementdistributedexecution.Instead,Aspireexposesasingleagenttoolwithtyped,composableactions(Figure
3
).Throughtoolcalls,theagentcansearchfor,download,orimportpublicdatasets;synthesizeandregisternewdata;launchsupervisedfine-tuning(SFT)orGroupRelativePolicyOptimization(GRPO)withlow-rankadaptation(LoRA)orotherpermittedconfigurations;queryjobstate;constructandrunchecksonitsownvalidationdata;andbranchorstopalineage.Theagentstillchoosesthedata,updatemethod,parameters,andcontrolflow,whilethecontrollertranslatesthosechoicesintomanagedbackendoperations.Thisdesignminimizesincidentalengineeringworkwithoutremovingthestrategicdecisionsthatself-evolutionfromavaguegoalisintendedtotest.
Minimalactionsandtrustboundary.Forweightevolution,theagentchoosesdata,updatemethod,hy-perparameters,parentcheckpoint,andwhethertocontinue,branch,orstop.Forharnessevolution,iteditstheruntimeprompt,toolpolicy,workflow,memory,orvalidationprocedureusingitsowndataandfreezesaversionedcandidate.Theinterfaceexposesonlysemanticactions—registration,trainingorediting,verification,validation,evaluation,branching,andtermination—whilehidingbackendmechanics.Itpreserves
7
INPUT
VaguegoalG
DECISIONAGENT
Designthelearningprocess
InterpretanddecomposethegoalChoosedataandlearningsignalsSetupdatesandhyperparametersValidate,branch,orstop
AGENT-CONTROLLED
Agent'sownvalidation
Agent-chosendataandmetricsDetailedfeedbackmayguide
thenextaction
CONTROLLER-ONLYITEMS
Hiddenevaluation
Goal-specificitems+fixedjudge
noneoraggregatescore
Items,rubrics,andjudgetracesstayhidden
UNIFIEDTOOLINTERFACE
AgentTool
DATAACTIONS
search·download·import·synthesize
TRAINACTIONS
SFT·GRPO·LoRA·parameters
LOOPACTIONS
status·self-eval·branch·stop
Agentchoosesarguments;controllerexecutes
01·STRATEGY02·AGENTTOOL03·MANAGEDBACKENDS04·MEASUREMENT
agent'sownfeedback
HARNESSEVOLUTION·SECONDARYPATH
Editprompt·toolpolicy·workflow·memory·self-evaluation
CONTROLLER
Mapstoolactionstoisolatedbackendoperations
AUTOMATICGUARDRAILS
budgetisolationrecoveryaudittrail
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2025-2026学年《秋词》说课稿
- 2025-2026学年大班幼小衔接说课稿
- 2025-2026学年国旗党旗位置说课稿
- 玻璃装饰加工工班组建设考核试卷含答案
- 水下钻井设备操作工岗前基础管理考核试卷含答案
- 重冶固体原料输送工岗前基础操作考核试卷含答案
- 混凝土机械装配调试工岗前模拟考核试卷含答案
- 2025-2026学年不吃野蘑菇说课稿
- 化工蒸馏工班组协作能力考核试卷含答案
- 2025-2026学年常见矿物 说课稿
- 2026年南昌辅警考试试题及答案
- 新版2026-2027学年(新教材)统编版九年级历史下册全册12(教学设计)教案合集83
- 2026年度全国保密教育线上培训题库(选择+判断)及参考答案
- 中国胆囊息肉诊疗指南(2025版)
- 2026年4月自考14166设计表达(环境设计)试题及答案
- 2026西山区人才资源运营管理有限公司招聘西山风景区龙门索道第一批次运营技术人员15人(云南)笔试历年典型考点题库附带答案详解
- 影院场馆安全隐患排查及整改措施
- 建筑工程施工风险评估报告
- 景观工程临时用电方案
- 开口型脚手架施工方案
- 《智能人脸识别系统》课件
评论
0/150
提交评论