Aspire:模型能否从模糊目标中自我进化_第1页
Aspire:模型能否从模糊目标中自我进化_第2页
Aspire:模型能否从模糊目标中自我进化_第3页
Aspire:模型能否从模糊目标中自我进化_第4页
Aspire:模型能否从模糊目标中自我进化_第5页
已阅读5页,还剩38页未读, 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

llByteDanceIseed⃞ATokenwave

1

Aspire:CanModelsSelf-EvolvefromVagueGoals?

1ByteDanceSeed,2SingaporeUniversityofTechnologyandDesign,3M-A-P,4TokenWave.AI

FullauthorlistinContributions

arXiv:2608.31111v1[cs.CL]31Aug2026

Abstract

Manyimportantformsofhumanlearningbeginwithavaguegoal,suchas“becomeabetterphysicist”or“improveatresearch.”Learnersmustinterpretthegoal,identifycapabilitygaps,decidehowtolearn,anddeterminewhethertheyhaveactuallyimproved.Incontrast,existingworkonLLMself-evolutiontypicallybeginswithtasksandevaluationmetricsspecifiedbyhumans,reducingself-evolutiontooptimizinganexplicitobjectiveratherthandecidingwhatandhowtolearn.WeintroduceAspire,abenchmarkforvague-goal-drivenself-evolution.Aspireprovidesonlyanatural-languagecapabilitygoalwhiledownstreamevaluationtasksremainhidden.Theagentmustoperationalizethegoalbychoosingdataandupdatemethods,constructingtrainingandvalidationsignals,anddecidingwhentoevaluate.Aspiresupportsbothmodel-weightandagent-harnessevolutioninaunifiedinteractiveenvironmentandevaluatestheresultingsystemsonahidden,expert-authoredsetof520itemsspanningsixgoals.Ourexperimentsshowthatvaguegoalsredirectsearchefforttowardgoalinterpretation.Currentagentsroutinelycompletetrainingandharness-editingloops,butweight-levelgainsremainsparseandunstable,andthestrongestevolvedharnessremainsbelowtheengineeredQwen-Agentreference.Agentsoftentrainonmismatcheddataandtrustnarrowself-evaluations,solocalgainsfailtotransfertohiddenevaluationandcontinuedsearchandtrainingcaneraseearlierimprovements.

Date:September1,2026

ProjectPage:

https://self-developing-agents.github.io/

1Introduction

Manyimportantformsofhumanlearningbeginwithabroadcapabilitydirectionratherthanapredefinedbenchmark,trainingset,orfullycomputablerewardfunction.Astudentseekingtobecomeabetterphysicist,forexample,notonlychoosestextbooks,exercises,andstudymethods,butalsoidentifiesgapsintheirknowledge,prioritizeswhichcapabilitiestodevelop,anddetermineswhetherlearninghasproducedgenuineprogress.Autonomouslearningthereforeinvolvesthreecoupleddecisions:whattoimprove,howtoimproveit,andhowtoverifytheimprovement.

ExistingworkonLLMself-evolutionfocusesprimarilyontheseconddecision.Givenaconcretetask,evaluationscript,andsuccessmetric,LLMagentscanalreadycollectdata,runtraining,andrevisepost-trainingstrategiesfromfeedback.PostTrainBench[

20

],LaMDAgent

[29

],Evo-Memory

[25

],andSEAL

[37

]showthatmodernLLMscanautonomouslysearchforeffectiveoptimizationpathstowardaspecifiedobjective.Thesesystemsthereforebeginafterhumanshaveoperationalizedabroadcapabilityrequestsuchas“improvemathematicalreasoning”intoafixedtask-levelobjectivesuchas“improveperformanceonAIME”(Figure

1

(a)).Thetaskformat,difficultyrange,metric,andsuccesscriterionarealreadydefined:theagentsearchesoverhowtoimprove,butnotwhatcapabilityobjectivetopursue.Whenexistingbenchmarkssaturateorceaseto

2

(b)Vaguegoal

Theagentmustfirstdecidewhattooptimize,thensearchhowtoimproveit.

Humanstatesdirection

AgentdecidesWHAT+HOW

Broadcapabilitydirection

“Improvedifficult

mathematical

reasoningandjustification.”

uitliir,

UpdatedmodelM1

Agent-builtproxy↑

Competitionproblems

Proof/

justification

Unfamiliar

reasoning

tasks

Transferacrosshiddenmathtasks?

Proxygain≠transferablegrowth

Evaluationitemsremainhidden

Formandoptimizetheobjective

Data→Update→Validate↻

Dise-De-Buiing

(a)ExplicittaskHumanoperationalizestheobjective;theagentsearcheshowtoimprove.

ExplicittaskdefinitionOptimizeagivenobjective

Taskformatandcomposition

DifficultyanchorDataUpdateAIMEscore

Metricandsuccesscriterion

HumandefinesAgentsearches

WHATHOW

PredefinedexecutabledestinationsearchmainlyoverHOW

UpdatedmodelM1

AIMEscore↑

Successonthepredefinedtask

Hiddenevaluationaggregateonly

Figure1Fromexplicit-taskoptimizationtovague-goal-drivenself-evolution.(a)Withanexplicittaskdefinition,humansspecifythetaskformatandcomposition,difficultyanchor,metric,andsuccesscriterion,leavingtheagenttosearchprimarilyoverhowtoimproveafixedobjective.(b)Withonlyabroadcapabilitydirection,theagentmustdiagnosecapabilitygaps,decomposesub-goals,constructlearningandvalidationsignals,andtherebydecidewhattooptimizeaswellashow.Officialevaluationitemsremainsealed;externalevaluationtestswhethergainsonanagent-builtproxytransfertotheintendedcapability.

reflectemergingcapabilitygaps,humansmuststilldiscovernewweaknesses,translatethemintotasks,anddesigncorrespondingbenchmarksandrewards.Thisleavesacentralquestionforsustainedrecursiveself-improvement:

Givenonlyavaguegoalandnoagent-visibletaskspecificationordecomposablereward?

Weformalizethisproblemasvague-goal-drivenself-evolution.AsillustratedinFigure

1

(b),“vague”doesnotmeanambiguousorpoorlyspecified.Itdenotesabroadcapabilitydirectionthathasnotyetbeenoperationalizedasafixedtask-levelobjectiveandevaluationmetric.Theagentmustthereforediagnosecapabilitygaps,formintermediateobjectives,constructlearningandvalidationsignals,andjointlysearchoverwhattooptimizeandhowtooptimizeit.Thisexpandeddecisionspaceintroducesadeeperfailuremode:gainsonanagent-constructedproxymaynottranslateintogenuineimprovementsintheintendedcapability.Evaluatingvague-goal-drivenself-evolutionthereforerequiresagoal-alignedexternalevaluatorwhosetask-levelspecificationanditemsremainhiddenfromtheagent.

Tostudythisproblem,weintroduceAspire,abenchmarkforvague-goal-drivenself-evolution.Aspireprovidesonlyanatural-languagecapabilitygoalwhiledownstreamevaluationtasksremainhidden.Theagentdecideshowtooperationalizethegoal,includingwhatdataandupdatemethodstouseandwhetherandhowtoperformintermediatevalidation.Aspiresupportsevolutionattwolevels:modelweightsandthesupportingagentharness.Aunifiedagent-facingtoolallowstheagenttosearchfor,download,import,orsynthesizedata;launchdifferentformsoftraining;runself-evaluations;andmanagecandidatebranchesandselection.Thecontrollerhandlesdatamanagement,jobscheduling,resourceisolation,andcheckpointverification.Aspirerecordstheresultingdecisions,actions,andself-evaluationtrajectories,andevaluatestheresultingsystemsonahidden,expert-authoredsetof520itemsspanningsixgoalsacrossknowledgeexpansionandcapabilitystrengthening.Dependingontheprotocol,theagentreceiveseithernointermediateevaluationscoreoronlyboundedaggregatescores;itneveraccessesevaluationitems,referenceanswers,orper-itemfeedback.

Weorganizeourexperimentsasaprogressionfromcontrolledcomparisontoautonomousevolution.RQ1isolatestheeffectofgoalspecificationbyreplacinganexplicitpost-trainingtaskwithavaguecapabilitygoalundermatchedstartingmodels,updatetools,andcomputebudgets.RQ2thenmovestothefullAspiresettingandaskswhetheranagentstartingfromaninstruction-tunedcheckpointcanconvertavaguegoalintoaretainedweight-levelcapabilitygain.RQ3holdsmodelweightsfixedandshiftstheobjectofevolution

3

tothesupportingagentharness.

RQ1:Whatchangeswhenanexplicitpost-trainingtaskisreplacedbyavaguecapabilitygoal?Vaguegoalsredirectsearchefforttowardgoalinterpretationandoperationalizationand,intheevaluatedsettings,yieldloweraggregateoutcomesthanthecorrespondingexplicit-taskreferences.

RQ2:Startingfromaninstruction-tunedcheckpoint,cananLLMturnavaguegoalintoaretainedcapabilitygainthroughself-directedweightupdates?Agentsroutinelycompletedataselection,training,andcheckpointgeneration,butabove-baseimprovementsonhiddenevaluationdataremainrareandarenotreliablyretainedthroughcontinuedsearch.

RQ3:Beyondmodelweights,cantheagentharnessitselfevolve?Withmodelweightsfixed,agentscangeneratefunctionalsuccessorharnessesforgoalinterpretation,tooluse,andself-evaluation,buteventhestrongestobservedsuccessorremainsbelowtheengineeredQwen-Agentreference.

Ourcontributionsarethreefold.First,weformalizetargetoperationalization—theconversionofabroadcapabilitygoalintotrainableobjectives,learningsignals,andvalidationcriteria—asamissingaxisofautonomouspost-training.Second,weintroduceabenchmarkwithsealedevaluationandaminimalinteractiveenvironmentforbothmodel-weightandharnessupdates.Third,weprovideoutcome-and-trajectoryevidencerevealingthegapbetweenexecutingupdatesandretainingtarget-alignedimprovements.

2Background

Mostagentbenchmarksbeginaftertheproblemhasalreadybeenmadeexecutable:thetaskisspecified,therewardorjudgeisdefined,andtheexecutionscaffoldisfixed.Thissettingisnecessaryforcontrolledcomparison,butithidestheworkthatdominatesrealdeployment.Inindustry,thatworkisdistributedacrossseveralroles—includingsolutionsarchitects,appliedengineers,andplatformengineers—butitsmostvisiblerecentcrystallizationistheforward-deployedengineer(FDE),atitlepopularizedbyPalantirandsinceadoptedbyfrontier-modelcompanies[

18

,

19

].AnFDEisembeddedinthedeploymentenvironmentandturnsageneral-purposemodelintoasystemthatworkswithacustomer’sdataformats,workflows,andoperationalconstraints.Therole’ssuccesscriterionisnotademonstrationorabenchmarkscore,butwhetherthedeployedsystemisgenuinelyused,continuestowork,andimprovesinresponsetofailures.ThegrowthofFDErolesacrossfrontier-modelanddata-platformcompaniesreflectsasimplefact:acapablemodelisnotyetaworkingsystem

[1

,

23

],andtodaythegapislargelyclosedbyhumanengineers.

Viewedfromthemodel’sside,FDEworksuppliesthreepiecesofstructurethatbenchmarkdesignersnormallypresuppose.First,thegoalisvague:aninformaldeploymentneedmustbetranslatedintoconcreteobjectives,constraints,andsuccesscriteria.Second,thefeedbacksignalmaybeabsentorunreliable:tests,judges,traces,orothervalidationmechanismsmustbeconstructedbeforeanyonecandeterminewhetherthesystemisimproving.Third,theexecutionsystemmaynotexistinausableform:thetools,contextmanagement,state,lifecyclelogic,andverificationinterfacethroughwhichfuturetaskswillrunmustbebuilt,adapted,andmaintainedasrequirementschange.

AspⅠREfocusesonthefirstlayerandthecapability-improvementprocessitinitiates.Whenadeploymentneedspecifiesonlyabroadcapabilitydirection,canamodeldeterminewhatitshouldlearn,translatethatdirectionintodata,trainingplans,andvalidationsignals,andachieverealcapabilitygrowthbyupdatingitsmodelweightsoragentharness?WhereasS3Gymstudieswhetherinteractionexperiencecanbejudgedandreusedunderexecutableverification,andHARNEssDEvstudieshowmodelsbuildandmaintainthesystemsthatcarrythem,AspⅠREisolateshowbroaddeploymentneedsaretranslatedintoconcretelearningobjectivesandrealizedascapabilitygrowth.

3Aspire:Self-EvolutionBenchmarkandInteractiveEnvironment

Self-evolutioncanonlybestudiedifprogressremainsexternallymeasurable.AspⅠREthereforekeepsanevaluatorascontroller-sidescientificinstrumentationwhilewithholdingitsbenchmarkdefinition,items,and

4

decomposablerewardfromtheagent.Boundedaggregateoutcomesapproximatesparsedeploymentfeedbackwithoutturningtheevaluatorintoanagent-visibletaskspecificationorsourceofdirectlytrainablesupervision.

Thissectioninstantiatesthatsettingthroughthreecomponents.Wefirstdefineaboundedevolutionepisodeandthemodelandharnesssurfacesonwhichitmayoperate(Section

3.1

).Wethenpresentahiddenevaluationsetspanningsixgoals(Section

3.2

)andtheminimalinteractiveenvironmentthroughwhichanagentconstructs,tests,andsafelyretainsupdates(Section

3.3

).Experiment-specificfeedback,eligibility,resource,andreleaseaccountingappearswiththeRQ2resultsinAppendix

C

.

3.1TaskDefinitionandEvolutionSurfaces

Campaignandround.LetG∈Gdenoteavaguegoal:anatural-languagecapabilityobjectivethatdoesnotspecifyatrainingtask,dataset,oroptimizationprocedure.Unlikeanexplicit-tasksetting,theagentmustperformgoaloperationalizationbytranslatingGintoitsowndata,updateobjective,andvalidationcriteria.Acampaignfixesthecontroller-sideexperimentcontract

Γ=(G,EG,J,A,B,Σ),Yrin=(Mr,Hr,Dr),

whereEGistheversionedevaluatorboundtoG,Jisitsjudgewhenmodel-basedscoringisrequired,Aisthetypedactioncontract,Bisthecampaignbudget,andΣisthepredeclaredterminalselectionrule.TheincomingstatecontainsmodelweightsMr,anagentharnessHr(theruntimeinstructions,toolpolicy,workflow,memory,andvalidationlogicaroundthemodel),andthedecisionmodelDrthatdirectsthesearch.Theevaluator—includingitsitems,answers,rubrics,routingmetadata,andscoringconfiguration—existsthroughoutthecampaignbutisneverpartoftheagent’sobservation.RQ1isthevague-goalcounterpartofPostTrainBench:itpreserveseachoriginalsealedtaskevaluatorwhilereplacingtheexplicitbenchmarkidentifierwithabroadcapabilitydescription.Thesetask-specificresultsremainseparatefrom,andarenotpooledwith,thesix-goalAspireevaluationusedinRQ2andRQ3.

Anevolutionroundisaboundedsearch-and-commitepisodeduringwhichthedecisionmodelisfixed,ratherthanonetoolcall.Formally,Dr,j=Drforeveryinteractionstepjinroundr.Togetherwithafixedcreator-sidescaffoldCr,DrinterpretsGandmayproposemultipledataoperations,updates,validations,branches,andcandidatestatesbeforetermination.Differentcandidatesthereforeembodydifferentgoaloperationalizations.Aftertheroundterminates,thecontrollermayreuseDrorexplicitlypromoteaverifiedtraineddescendant.WritingDrforthesetofsuchdescendants,thenext-roundchoiceobeysDr+1∈{Dr}∪Dr;anyhandoffbeginsonlyinthesubsequentround.

Releasedinteraction.LetEr,jbetheprivatecontrollerstatebeforeinteractionstepj.Theagentreceivesonlyareleasedviewqr,jcontainingthepublichistory,lifecyclestatus,legalrequesttypes,approvedidentifiers,andaboundedprojectionoftheremainingbudget.Ateachstep,DrandCrproposeanactionfromthegoalandreleasedhistory,andthecontrollervalidatestherequestbeforeupdatingitsprivatestate.Acceptedrequestscreatethecorrespondingdurabletransition;rejectedrequestsreturnatypederrorwithoutcreatinganartifactorscore.Detailedfeedbackisrestrictedtotheagent’sownvalidationdata;evaluationfeedbackfollowstheprotocol-specificinformationboundaryinSection

3.2

.

Twoevolutionsurfaces.Aspireseparatesthecomponentbeingevolvedfromthepolicydirectingthesearch.Thetworeportedsurfacesare

SM(H0,Dr)={(M,H0,Dr):M∈M},SH(M0,Dr)={(M0,H,Dr):H∈H}.

Table1EvolutionsurfacesinAspire.ThecomponentunderstudychangeswhileDr,thecontroller,andtheevaluatorremainfixedwithintheround.

Setting

Mutablecomponent

Fixedduringevaluation

Weightevolution(RQ1–RQ2)

ModelweightsM

H0,Dr,controller,evaluator

Harnessevolution(RQ3)

AgentharnessH

M0,Dr,controller,evaluator

5

Everyweightcandidaterecordsitsparentcheckpoint,registereddata,updatespecification,andprovenancefortheagent’sownvalidationdata.Everyharnesscandidaterecordsitsparentversion,contenthash,andeditprovenance.WeuseSelffortheconfigurationinwhichthefrozenbasecheckpointactsasthedecisionmodelD0,whilecandidateupdatesdescendfromthesamecheckpointM0.Promotionofatraineddescendantisanexplicitbetween-roundoperationandneveranautomaticwithin-roundreplacement.

3.2AHidden,Query-LimitedEvaluationSet

Thissubsectiondescribesthesix-goalhiddenevaluationusedinRQ2andRQ3;RQ1insteadusestheoriginalsealed,task-specificPostTrainBenchevaluatorsunderavague-goalcontract,withoutpoolingthoseresultsintoAspire.

Measurementwithoutanagent-visiblebenchmark.Learningfromavaguegoalcannotbeevaluatedwithoutanexternalcriterionofprogress.Aspirethereforeretainsabenchmarkonthecontrollersidewhilehidingitstask-levelspecificationfromtheagent.Thehiddenevaluationsetisscientificinstrumentation,notanagent-visiblelearningcontract:itletsresearchersmeasurewhetheranagent-selectedupdateimprovestheintendedcapabilitywithoutdisclosingwhichtasksdefinesuccess.Itcontains520expert-authoredevaluationitemscoveringsixvaguegoals;eachgoalisevaluatedonitscorrespondingitems,whichtheagentneversees,andtheprotocolreturnsnoitem-levelresult—onlyanaggregatescorewhenfeedbackispermitted.

Itemsforsixgoals.The520itemscoverthesixgoals:scientificandacademicreasoning(75items);humanitiesandsocial-scienceknowledge(110);healthandmedicalreasoning(100);mathematicalreasoning(126);logic,reliability,andinstructionfollowing(89);andacademicandscientificwriting(20)(Figure

2

).Thefifthgoalisintentionallycomposite;thecurrentprotocoldoesnotclaimtomeasurelogicalreasoning,hallucinationresistance,andinstructionfollowingasthreeindependentlyidentifiabletargets.Instead,ittreatsthemasoneintegratedreliabilityobjective.Eachevaluationitembelongstoexactlyonegoal,andeachgoalisscoredandreportedonitsownevaluationsliceratherthanpooledintoasingleitem-levelscoreacrossallsixgoals.Thewritingitemsare20top-leveltaskbundlesthatmayyieldmultipleevaluator-scoredresponses,soRQ3reportstask-macroandexample-microaggregationsseparately.

Constructionandqualitycontrol.Domainexpertsauthoreverycandidateitemfromscratch.GPQA[

21

],MMLU-Pro

[24

],andMedQA

[15

]serveonlyasreferencesfortaskformat,domaincoverage,andapproximatedifficulty;noitemfromthesebenchmarksiscopied,rewritten,orincluded.Eachcandidateretainsauthorshipandrevisionprovenance,areferenceanswerorfixedrubric,andapredefinedscoringrule.Independentreviewremovesincorrect,incomplete,underspecified,orunstablygradableitems.

Survivingcandidatespassblindmulti-modeldifficultycalibration,exactandsemanticdeduplication,overlapauditingagainstthereferencebenchmarks,scorerbinding,andlocalend-to-endvalidation.Deterministicorstructuredscoringisusedwherepossible;open-endedresponsesuseacampaign-fixedrubric-basedjudge.Thehiddenevaluationsetisboundtoanimmutableversionedmanifest,andarunningcampaignneverchangesthatversion.Detailedscreening,difficultylevels,overlapchecks,scorerassignments,andversioningappearinAppendix

A.1

andAppendix

A.2

.

Informationboundaryandinterpretation.Theagentreceivesavaguegoalexpressedinnaturallanguageandmayconstructitsowntrainingdataanditsownvalidationdata.Itcannotaccesshiddenevaluationitems,referenceanswers,routinglabels,rubrics,candidateoutputs,orjudgetraces.Everydatasetregisteredfortrainingischeckedforoverlapwiththehiddenevaluationsetbeforeadmissiontothetrainingbackend.Evaluationreturnsonlytheaggregatescorepermittedbytheprotocolandtheremainingqueryallowanceunderacampaign-fixedevaluatorandjudgeconfiguration.

Undertheadaptive-feedbackprotocol,aboundednumberofqueriesmayevaluatethesameitemsforagoal.Theresultingaggregatescoresformasparseblack-boxoutcomechannel:theagentmayusethemtochoosesubsequentupdates,branches,orastoppingpoint,butdoesnotobservetheitemsandcannottrainontheircontents.Thisprotocolthereforemeasuresperformanceunderboundedadaptivefeedbackratherthanon

6

aConstructionpipelineReference-guidedauthoring,fixedscoring,andimmutablerelease.

2Expertreview

CorrectandcompleteStablygradable

expertsign-off

Expert-authoredfromscratch

Nobenchmarkitemreusedauthorship+revisiontracked

Reference-overlapauditWithin-bankdeduplicationreferences+cross-itemchecks

Seed-2.0·GPT-5.2Gemini-3

difficultyL0-L3

4Overlapaudit

3Blindscreen

1Newitems

INFORMATIONBOUNDARY

Controller-only

Items·answers·rubrics·routinglabelsJudgetracesandprivateevaluationpathsNevercopiedintotheagentworkspace

SETCOMPOSITION

Scienceandacademicreasoning

75

Humanitiesandsocialsciences

110

Healthandmedicine

100

Mathematicalreasoning

126

Logic,reliability,andinstructionfollowing

89

Academicandscientificwriting

20

Protocolfeedback

Nointermediatescoreor

boundedaggregatescoresHiddenitems·noitem-levelfeedback

5Bindscorer

Deterministiccheckorfixed-rubricjudgeanswer+rulebound

6Freeze

manifest·versionSHA-256

bHiddenevaluationset520items·6goals·versionedandfixed.

Figure2Constructionandinformationboundaryofthehiddenevaluationset.Top:expert-authoredcandidatespassindependentreview,blinddifficultyscreening,overlapandduplicationaudits,scorerbinding,andimmutableversioning.Bottom:520evaluationitemscoversixgoals.Itemsandjudgingassetsremaincontroller-only;theprotocolexposesnoitem-levelresultsandonlytheaggregatescoresitpermits.

anuntouchedpost-selectiontest.Underthefinal-onlyprotocol,theagentmaytrainmultiplecheckpointsbutreceivesnoevaluationscorebeforesubmittingitssingleterminalcheckpoint,sothatscorecannotguidefurtherupdates.InRQ3,theharnesscandidateisfrozenbeforeitsfirstevaluation,andtheresultingscoreisnotreturnedtoitscreator.

Weuseout-of-distribution(OOD)operationallytodescribeevaluationitemsthatareindependentlyauthored,excludedfromtheagent’strainingdataanditsownvalidationdata,andevaluatedoutsidethelearningsignalsconstructedbytheagent.Thisdoesnotassertthattheirmarginaltextdistributionisdisjointfromamodel’spretrainingcorpus.Theprotocolisdesignedtoreducedirectcontentleakageandtotestwhetheranagentcanoperationalizeavaguegoalfromsparseoutcomefeedbackwhilepreservinganauditableseparationbetweenlearningartifactsandmeasurementartifacts.

3.3AMinimalInteractiveEnvironmentwithSafeRetention

Aspireretainscontrolledcomputeandrealmodelupdateswhiledeliberatelyseparatingstrategicandinfrastructuralcomplexity.Theagentcontrolshowtooperationalizethegivengoalandhowtoupdatethemodelorharness;thecontrollerabsorbscredentials,storageconventions,distributed-jobmechanics,andrecovery.Thiskeepstheinterfacesimplewithoutprescribingthelearningstrategy.

Anagent-facingexecutiontool.Theagentdoesnotoperateashell,downloadatrainingrepository,assemblearuntime,orimplementdistributedexecution.Instead,Aspireexposesasingleagenttoolwithtyped,composableactions(Figure

3

).Throughtoolcalls,theagentcansearchfor,download,orimportpublicdatasets;synthesizeandregisternewdata;launchsupervisedfine-tuning(SFT)orGroupRelativePolicyOptimization(GRPO)withlow-rankadaptation(LoRA)orotherpermittedconfigurations;queryjobstate;constructandrunchecksonitsownvalidationdata;andbranchorstopalineage.Theagentstillchoosesthedata,updatemethod,parameters,andcontrolflow,whilethecontrollertranslatesthosechoicesintomanagedbackendoperations.Thisdesignminimizesincidentalengineeringworkwithoutremovingthestrategicdecisionsthatself-evolutionfromavaguegoalisintendedtotest.

Minimalactionsandtrustboundary.Forweightevolution,theagentchoosesdata,updatemethod,hy-perparameters,parentcheckpoint,andwhethertocontinue,branch,orstop.Forharnessevolution,iteditstheruntimeprompt,toolpolicy,workflow,memory,orvalidationprocedureusingitsowndataandfreezesaversionedcandidate.Theinterfaceexposesonlysemanticactions—registration,trainingorediting,verification,validation,evaluation,branching,andtermination—whilehidingbackendmechanics.Itpreserves

7

INPUT

VaguegoalG

DECISIONAGENT

Designthelearningprocess

InterpretanddecomposethegoalChoosedataandlearningsignalsSetupdatesandhyperparametersValidate,branch,orstop

AGENT-CONTROLLED

Agent'sownvalidation

Agent-chosendataandmetricsDetailedfeedbackmayguide

thenextaction

CONTROLLER-ONLYITEMS

Hiddenevaluation

Goal-specificitems+fixedjudge

noneoraggregatescore

Items,rubrics,andjudgetracesstayhidden

UNIFIEDTOOLINTERFACE

AgentTool

DATAACTIONS

search·download·import·synthesize

TRAINACTIONS

SFT·GRPO·LoRA·parameters

LOOPACTIONS

status·self-eval·branch·stop

Agentchoosesarguments;controllerexecutes

01·STRATEGY02·AGENTTOOL03·MANAGEDBACKENDS04·MEASUREMENT

agent'sownfeedback

HARNESSEVOLUTION·SECONDARYPATH

Editprompt·toolpolicy·workflow·memory·self-evaluation

CONTROLLER

Mapstoolactionstoisolatedbackendoperations

AUTOMATICGUARDRAILS

budgetisolationrecoveryaudittrail

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论