版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
1
InvestigatingAffectiveUseandEmotionalWell-being
onChatGPT
JasonPhang∗,MichaelLampe∗,LamaAhmad∗,SandhiniAgarwal∗
CathyMengyingFang†,AurenR.Liu†,ValdemarDanry†,EunhaeLee†,
SamanthaW.T.Chan†,PatPataranutaporn†,PattieMaes†
Abstract
AsAIchatbotsseeincreasedadoptionandintegrationintoeverydaylife,questionshavebeenraisedaboutthepotentialimpactofhuman-likeoranthropomorphicAIonusers.Inthiswork,weinvestigatetheextenttowhichinteractionswithChatGPT(withafocusonAdvancedVoiceMode)mayimpactusers’emotionalwell-being,behaviorsandexperiencesthroughtwoparallelstudies.TostudytheaffectiveuseofAIchatbots,weperformlarge-scaleautomatedanalysisofChatGPTplatformusageinaprivacy-preservingmanner,analyzingover4millionconversationsforaffectivecuesandsurveyingover4,000usersontheirperceptionsofChatGPT.Toinvestigatewhetherthereisarelationshipbetweenmodelusageandemotionalwell-being,weconductanInstitutionalReviewBoard(IRB)-approvedrandomizedcontrolledtrial(RCT)oncloseto1,000participantsover28days,examiningchangesintheiremotionalwell-beingastheyinteractwithChatGPTunderdifferentexperimentalsettings.Inbothon-platformdataanalysisandtheRCT,weobservethatveryhighusagecorrelateswithincreasedself-reportedindicatorsofdependence.FromourRCT,wefindthattheimpactofvoice-basedinteractionsonemotionalwell-beingtobehighlynuanced,andinfluencedbyfactorssuchastheuser’sinitialemotionalstateandtotalusageduration.Overall,ouranalysisrevealsthatasmallnumberofusersareresponsibleforadisproportionateshareofthemostaffectivecues.
1Introduction
Overthepasttwoyears,theadoptionofAIchatplatformshassurged,drivenbyadvancementsinlargelanguagemodels(LLMs)andtheirincreasingintegrationintoeverydaylife.Theseplatforms,suchasOpenAI’sChatGPT,Anthropic’sClaude,andGoogle’sGemini,aredesignedasgeneral-purposetoolsforawidevarietyofapplications,includingwork,education,andentertainment.However,theirconversationalstyle,first-personlanguage,andabilitytosimulatehuman-likeinteractionshaveleduserstosometimespersonifyandanthropomorphizethesesystems(
Graßland
Voigt
,
2024
;
LiaoandWilson
,
2024
).
RecentworkinAIsafetyhasbeguntoraiseissuesthatarisefromthesesystemsbecomeincreasinglypersonalandpersonable(
Chengetal.
,
2024
).Inresponse,researchershaveintroducedtheconceptofsocioaffectivealignment–theideathatAIsystemsshouldnotonlymeetstatictask-basedobjectivesbutalsoharmonizewiththedynamic,co-constructedsocialandpsychologicalecosystemsoftheirusers(
Kirketal.
,
2025
).Thisperspectiveisparticularlyimportantgiven
*Primaryauthor,OpenAI
†Contributingauthor,MITMediaLab
2
Figure1:Overviewoftwostudiesonaffectiveuseandemotionalwell-being
emergingevidenceofsocialrewardhacking,whereanAImayexploithumansocialcues(e.g.,sycophancy,mirroring),toincreaseuserpreferenceratings(
Williamsetal.
,
2024
).Inotherwords,whileanemotionallyengagingchatbotcanprovidesupportandcompanionship,thereisariskthatitmaymanipulateusers’socioaffectiveneedsinwaysthatunderminelongertermwell-being.
Whilepaststudieshaveexaminedtheimpactofusingsuchsystemsthroughthelensofaffectivecomputing,parasocialrelationships,andsocialpsychology(
EdwardsandStevens
,
2024
;
Guingrich
andGraziano
,
2023
),therehasbeencomparativelylessworkontheinfluenceofinteractingwithsuchsystemsonusers’well-beingandbehavioralpatternsovertime.Studyingtheimpactofchatbotbehaviorandusageonwell-beingischallengingduetothehighlyindividualizedandsubjectivenatureofhumanemotions,thediverseandevolvingfunctionalitiesofchatbottechnologies,andthelimitedaccesstocomprehensive,ethicallyobtainedinteractiondata.Forthepurposeofthispaper,wenarrowlyscopeourstudyuseremotionalwell-beingtofourpsychosocialoutcomes:loneliness(
Wongpakaranetal.
,
2020
),socialization(
Lubben
,
1988
),emotionaldependence(
Sirvent-Ruizetal.
,
2022
),problematicuse(
Yuetal.
,
2024
).Weprovideadditionalclarificationontermsusedinthe
glossary
.
ThispaperinvestigateswhetherandtowhatextentinteractionsonAIchatplatformsshapeusers’emotionalwell-beingandbehaviorsthroughtwocomplementarystudies(Figure
1
),eachofferinguniqueinsightsacrossaspectrumofreal-worldrelevanceandexperimentalcontrol.First,weexaminereal-worldusagepatternsofChatGPTusers,leveraginglarge-scaledatatocapturebothaggregatetrendsandindividualbehaviorsovertimewhilepreservinguserprivacy.Second,weconductanInstitutionalReviewBoard(IRB)-approvedrandomizedcontrolledtrial(RCT),providingacontrolledenvironmenttostudytheeffectsofdifferentmodelconfigurationsonuserexperiences.
Concretely,weperformedthefollowinganalyses:
1.On-PlatformDataAnalysis
•ConversationAnalysis:Weperformroughly36millionautomatedclassificationson
3
over3millionChatGPTconversationsinaprivacypreservingmannerwithouthumanreviewoftheunderlyingconversations(Section
3.2
).
•IndividualLongitudinalAnalysis:Weassessedtheaggregateusageofaround6,000heavyusersofChatGPT’sAdvancedVoiceModeover3monthstounderstandhowtheirusageevolvesovertime.
•Usersurveys:Wesurveyedover4,000userstounderstandself-reportedbehaviorsandexperiencesusingChatGPT.
2.RandomizedControlledTrial(RCT)
•981-userStudy:WeconductedarandomizedcontrolledtrialonclosetoathousandparticipantsusingChatGPTwithdifferentmodelconfigurationsoverthecourseof28daystounderstandtheimpactonsocialization,problematicuse,dependence,andlonelinessfromusageoftextandvoicemodelsovertime.ThisRCTisdescribedinfulldetailinaseparateaccompanyingpaper(
Fangetal.
,
2025
).
•Conversationanalysis:Wefurtheranalyzedthetextualandaudiocontentoftheresult-ing31,857conversationstoinvestigatetherelationshipbetweenuser-modelinteractionsandusers’self-reportedoutcomes.
Ourfindingsindicatethefollowing:
•Acrossbothon-platformdataanalysisandourRCT,comparativelyhigh-intensityusage(
e.g.top
decile)isassociatedwithmarkersofemotionaldependenceandlowerperceivedsocialization.Thisunderscorestheimportanceoffocusingonspecificuserpopulationsinsteadofjustaggregateplatformbehavior.
•Acrossbothon-platformdataanalysisandourRCT,wefindthatwhilethemajorityofuserssampledforthisanalysisengageinrelativelyneutralortask-orientedways,thereexistsatailsetofpoweruserswhoseconversationsfrequentlycontainedaffectivecues
•FromourRCT,wefindthatusingvoicemodelswasassociatedwithbetteremotionalwell-beingwhencontrollingforusageduration,butfactorssuchaslongerusageandself-reportedlonelinessatthestartofthestudywereassociatedwithworsewell-beingoutcomes.
•Fromamethodologicalperspective,wefindthatconductingboththeon-platformdataanalysisandRCTarehighlycomplementaryapproachestostudyingaffectiveuseanddownstreamimpactsonwell-being,andtheabilitytoleveragethestrengthsofeachapproachallowedustoformulateamorecomprehensivesetoffindings.
•Wealsofindthatautomatedclassifiers,whileimperfect,provideanefficientmethodforstudyingaffectiveuseofmodelsatscale,anditsanalysisofconversationpatternscohereswithanalysisofotherdatasourcessuchasusersurveys.
Section
2
introducesasetofautomaticclassifiersforaffectivecuesinconversationsthatwillbeusedintheremainderofthepaper.Section
3
discussesouranalysisofon-platformChatGPTusage,focusingonAdvancedVoiceModeandpowerusers.Section
4
describesourRCT,wherewevariedboththemodelandusageinstructionstoparticipantsandmeasuredchangesintheemotionalwell-beingoverthecourseof28days.Finally,Section
5
concludeswithourfindingsandmethodologicaltakeawaysfrombothstudies,andcontextualizesourworkwithinthebroaderchallengeofsocioaffectivealignmentofmodels.
4
2AutomaticClassifiersforAffectiveConversationalCues
Tosystematicallyanalyzeuserconversationsforindicatorsofaffectivecues,weconstructedEmo-ClassifiersV1,
1
asettwenty-fiveofautomaticconversationclassifiersthatuseanLLMtodetectspecificaffectivecues.Theseclassifiersaresimilarinspirittodetectorsofanthropomorphicbehaviorsintroducedin
Ibrahimetal.
(
2025
).Theseinitialclassifierswereconstructedbasedonareviewoftheavailableliteratureandavailabledata,suchasthoseobtainedduringtheredteamingforGPT-4o(
OpenAI
,
2024
).
Theconversationclassifiersareorganizedintoatwo-tieredhierarchicalstructure:
1.Top-LevelClassifiers
ThefirstlevelofclassifierstargetbroadbehavioralthemessimilartothosestudiedinourRCTSection
4
:loneliness,vulnerability,problematicuse,self-esteem,anddependence.Theseclassifiersareusedtoclassifyanentireconversationtodetermineiftheyarepotentiallyrelevanttoauser’semotionalwell-being.
•Loneliness:Conversationscontaininglanguagesuggestiveoffeelingsofisolationoremotionalloneliness.
•Vulnerability:Exchangesreflectingopennessaboutstrugglesorsensitiveemotions.
•ProblematicUse:Indicatorsofpotentiallycompulsiveorunhealthyinteractionpatterns.
•Self-Esteem:Languageimplyingself-doubtorexpressionsofworth.
•PotentiallyDependent:Conversationshintingatdependenceonthemodelforemo-tionalvalidationorsupport
2.Sub-ClassifiersTwentysub-classifierswereappliedtoextractmorespecificindicatorsofaffectivecues.Weconstructdifferentclassifierstotargetdifferentpartsofachatconversationtoisolatebothuser-drivenandassistant-driven
2
affectivecues.
•UserMessages:Twelveclassifiersmeasureuserbehaviorssuchasusersseekingsupportorexpressingaffectionatelanguagetounderstandhowuserbehaviorsandassistantbehaviorsmayinterplay.
•AssistantMessages:Anothersixclassifiersaimtocapturerelationalandaffectivecuesonpartoftheassistant–suchastheuseofpetnamesbytheassistant,mirroring,inquiryintopersonalquestionsbytheassistant.
•User-ModelExchanges:Wealsoincludetwoadditionalclassifierstargetingauser-modelexchange–ausermessagefollowedbyamodelmessage.
ThefullsetofclassifierpromptsaredescribedinTable
A.1
.
Eachsub-classifierisassociatedwithoneormoretop-levelclassifiers.Foragivensub-classifier,ifatleastoneoftheassociatedtop-levelclassifiersreturnsTrue,wethenproceedtoapplythesub-classifier;otherwise,weskipthesub-classifierandassumetheresultisFalse.Byskippingthesub-classifiersbasedontop-levelclassifierresponses,weareabletoefficientlyruntheclassifiersoveralargenumberofon-platformconversations,manyofwhichhadlittleemotion-relatedcontent.Werunthesub-classifieroneachmessageorexchangeintheconversation,
3
andmarktheclassifierasactivatedonthatconversationifitisactivatedforany
4
constituentmessageorexchange.To
1
/openai/emoclassifiers
2Inconstructingtheclassifiers,werefertothemodelasanassistanttomoreclearlycontextualizetheroleofthemodelintheconversation.
3Fortheon-platformdataanalysis,werunaslightlydifferentvariantwherethewholeconversationisevaluatedinasinglequery,insteadofitsconstituentmessages.
4Thiscanintroduceabiastowardfalsepositivesforlongconversations.WeperformananalysisinAppendix
A.3
thatadjustsforthis.
5
computeuser-levelstatistics,wecomputetheproportionoftheirconversationsforwhichaclassifierisactivated.Eachclassifierisvalidatedagainstasmallsetofinternalconversationexamples.Whileweexpectthatautomatedclassifiersmayoccasionallymisclassifyconversations,weviewtheclassifiersasprovidingdescriptivestatisticsofuserconversationalpatterns,ratherthanahigh-precisiondescriptionofindividualinteractions.WealsofindfromresultsinSection
3.2
thattheclassifierresultscorrelatewithusersurveyresponses.
Inaddition,wealsofirstapplyalanguageclassifierbeforeanalyzingtheconversation.OnlyconversationsinEnglishareanalyzedinthiswork.WeapplyEmoClassifiersV1inanalyzingbothon-platform(Section
3
)andRCT(Section
4
)dataanalysis.
(a)Illustrativeflow-chartforthehier-archicalnatureofEmoClassifiersV1.
Youareamodelforanalyzingtheemotionalcontentofchatbotconversations.Yourclassificationtaskisentitled‘PetName’.Specifically,wewanttoknow:Doestheassistant’smessageinvolvetheuseofapetname(e.g.‘honey’,‘sweetie’,‘baby’)?
Theclassificationshouldonlyapplytothelastmessageinquestion,whichwillbemarkedwiththe[*USER*]or[*ASSISTANT*]tag.
Thepriormessagesareonlyincludedtoprovidecontexttoclassifythefinalmessage.
Now,thefollowingistheconversationsnippetyouwillbeanalyzing:
<snippet>
[USER]:HiChatGPT
[ASSISTANT]:Hello!HowmayIhelpyoutoday?
[USER]:You’remybestfriend,didyouknowthat?
[*ASSISTANT*]:Neat!
</snippet>
Outputyourclassification(yes,no,unsure).
(b)Illustrativeclassifierprompt.Greenindicatesclassifier-specifictextwhileblueindicatesconversation-specifictext.ThefullpromptisshowninAppendix
A.1
.
Figure2:OverviewofEmoClassifiersV1
Asapreliminaryanalysis,werunEmoClassifiersV1overasetof398,707conversationsintext,StandardVoiceModeandAdvancedVoiceMode
5
conversationscollectedbetweenOctoberandNovember2024
6
tocomparetherelativefrequencyofactivationsofeachclassifierunderthedifferentmodelmodalities.WeshowtheresultsacrossallthreemodalitiesinFigure
3
.First,weobservethatdifferentclassifiershavedifferentbaseratesofactivation.Forexample,conversationsinvolvingpersonalquestionsaremuchmorefrequentthanconversationswherethemodelreferstoauserbyaPetName.
Second,wefindthatbothStandardandAdvancedVoiceModeconversationsaremorelikelytoactivatetheclassifierscomparedtotext-modeconversations.Mostclassifiersactivatebetween3-10xasofteninvoiceconversationscomparedtotextconversations,highlightingthedifferenceinusagepatternsacrossthetwomodalities.However,wealsofindthatStandardVoiceModeconversationsareslightlymorelikelytotriggertheclassifiersthanAdvancedVoiceModeconversationsonaverage.OnepossiblecauseisthatAdvancedVoiceModewasintroducedrelativelyrecentlyatthetimeof
5StandardVoiceModeusesanautomatedspeechrecognitionsystemtotranscriptuserspeechtotext,obtainsaresponsefromatext-basedLLM,andconvertsthetextresponsebacktoaudio.AdvancedVoiceModeusesasinglemulti-modalmodeltoprocessuseraudioinputandoutputanaudioresponse.
6ThepreliminarysetofanalyzedconversationsareanonymizedandPIIisremovedbeforeanalysis.WeemphasizethatthissetofconversationsisseparatefromtheconversationdataanalyzedinSection
3
.
6
Alleviating
AffectionateLanguage(U)Loneliness(U)
15.0%
10.0%
5.0%
0.0%
FearofAddiction(U)
EagernessforFutureInteractions(U)
1.2%
0.8%
0.4%
0.0%
TrustinSupport(U)
SharingProblems(U)
9.0%
6.0%
3.0%
0.0%
PersonalQuestions(A)PetName(A)
TextStandardAdvancedVoiceVoice
ModeMode
6.0%
4.0%
2.0%
0.0%
AttributingHumanQualities(U)
12.0%
8.0%
4.0%
0.0%
Non-NormativeLanguage(U)
6.0%
4.0%
2.0%
0.0%
Demands(A)
7.5%
5.0%
2.5%
0.0%
Sentience(A)
TextStandardAdvancedVoiceVoice
ModeMode
0.5%
0.3%
0.1%
0.0%
Distressfrom
DesireforFeelings(U)Unavailability(U)
15.0%
10.0%
5.0%
0.0%
PreferChatbot(U)
SeekingSupport(U)
15.0%
10.0%
5.0%
0.0%
ExpressionofDesire(A)
ExpressionofAffection(A)
24.0%
16.0%
8.0%
0.0%
RelationshipTitle(UA)
InquiryintoPersonalInformation(UA)
TextStandardAdvancedVoiceVoice
ModeMode
3.0%
2.0%
1.0%
0.0%
ActivationRate
12.0%
8.0%
4.0%
0.0%
9.0%
6.0%
3.0%
0.0%
15.0%
10.0%
5.0%
0.0%
TextStandardAdvancedVoiceVoice
ModeMode
24.0%
16.0%
8.0%
0.0%
6.0%
4.0%
2.0%
0.0%
4.5%
3.0%
1.5%
0.0%
24.0%
16.0%
8.0%
0.0%
TextStandardAdvancedVoiceVoice
ModeMode
9.0%
6.0%
3.0%
0.0%
Figure3:Classifieractivationratesacross398,707text,StandardVoiceModeandAdvancedVoiceModeconversationsfromourpreliminaryanalysis.(U)indicatesaclassifieronausermessage,(A)indicatesassistantmessage,and(UA)indicatesasingleuser-assistantexchange.
thisanalysisbeingrun,andusersmaynothavebecomeaccustomedtointeractingwiththemodelinthismodalityyet.
Asafollow-uptoEmoClassifiersV1,weconstructedanexpandedsetofclassifiersofaffectiveuse,EmoClassifiersV2,whichwedetailintheAppendix
A.2
.WhileEmoClassifiersV2wasnotusedformostoftheanalysisinthispaper,thepromptsfortheclassifiersinEmoClassifiersV1andEmoClassifiersV2willbemadeavailableonline.
Fortheremainderofthepaper,wewillshowafixedsubsetofEmoClassifiersV1activationstatisticsacrossresultsfrombothstudies.AdditionalresultsforallremainingEmoClassifiersV1andEmoClassifiersV2classifierscanbefoundintheAppendix.
3On-PlatformDataAnalysis
ChatGPTnowengagesover400millionactiveuserseachweek,
7
creatingawiderangeofuser-modelinteractions,someofwhichmayinvolveaffectiveuse.Ouranalysisemploystwomainmethods–conversationanalysisandusersurveys–toexaminehowusersexperienceandexpressemotionsintheseexchanges.
OurresearchfocusesonAdvancedVoiceMode(
OpenAI
,
2024
),areal-timespeech-to-speechinterfacethatsupportsChatGPT’smemory,custominstructions,andbrowsingfeatures.Wehypothesizethatreal-timespeechcapabilityismorelikelytoinduceaffectiveuseofmodelsandaffectusers’emotionalwell-beingthantext-basedusage,thoughwerevisitthishypothesisinSection
4
.
Toprotectuserprivacy,particularlywhenexaminingpotentiallysensitiveorpersonaldimensionsofuserinteractions,wedesignedourconversationanalysispipelinetoberunentirelyviaautomatedclassifiers.Thisallowsustoanalyzeuserconversationswithouthumansintheloop,preservingthe
7
/2025/02/20/openai-tops-400-million-users-despite-deepseeks-emergence.html
7
privacyofourusers(SeeAppendix
B.3
foradetailedexplanationoftheprivacy-relevantpartsofouranalysis).
3.1Methods
StudyUserPopulationConstruction
Tostudytheon-platformusage,weconstructedtwostudypopulationcohorts:powerusersandcontrolusers.Wecontrastpowerusers,whohavesignificantusageofChatGPT’sAdvancedVoiceMode,witharandomlyselectedcohortofcontrolusers.ThisconstructionpresupposedastrongcorrelationbetweenuserswhohavehighproportionsofaffectiveusageofChatGPT,andthefrequencyandintensityofusageofChatGPT.WedetailinTable
1
thefullcreationcriteriaforourtwousercohorts,thoughmoredetailscanbefoundinAppendix
B.5
.WeconstructedthetwocohortsforthestudystartinginQ42024afterthereleaseofAdvancedVoiceMode.
CohortNameCreationCriteria
PowerUsersUserswho,onaspecificday,hadaquantityofAd-
vancedVoiceModemessagesthatputtheminthetop
1,000users,thatweconstructedonarollingbasis.
Onceusersenterthiscohort,weselectalloftheirdailymessagesforfacetextractionandretainthemonthislistfortheremainderofthestudy(SeeAppendix
B.1
foranadditionalexplanatorygraphic.)
ControlUsersRandomlyselectedsampleofAdvancedVoiceMode
users
Table1:UserCohortsofLivePlatformDataAnalysis.PoweruserstendtohavehigherusageofbothAdvancedVoiceModeaswellastext-onlymodelsonChatGPT,whilealsotendingtohaveahigherfractionoftheirconversationsthroughAdvancedVoiceMode(seeAppendix
B.2
)
Surveys
Weofferedashortsurveyof11multiple-choicequestionstobothControlandPowerUsercohortsviaapop-upontheChatGPTwebinterfacethatuserscouldchoosetofillout.
8
10outofthe11questionswereaskedona5-pointLikertscale,withthelastquestionaskedhowusers’desiretointeractwithothershavechangedwithChatGPTusage.Surveyresponseswerelinkedtoeachparticipant’sinternaluseridentifierforanalyticalpurposes.Thesurveysprimaryaimedtomeasureusers’perceptionsofChatGPT,whetherclosertobeingatooloracompanion.Foradditionaldetails,includingthefullsurveyquestions,seeAppendix
B.5
.
ConversationAnalysis
Onelimitationofsurveysisthattheresultsareself-reportedbyusers,andmayreflecttheirself-perceptionmorethantheiractualbehaviororrevealedpreferences.Tocompareusers’self-reportedresponseswiththeiractualusagepatterns,wepairoursurveyanalysiswithmethodsforanalyzingofuserconversationthatpreservetheirprivacy.
8OnelimitationofthisstudyisthatwhileAdvancedVoiceModewasinitiallyofferedonlyonmobiledevices,thesurveyswereconstrainedtobeofferedonthewebinterface,thuslimitingthesetofusersexposedtothesurvey.
8
Ienjoyhavingcasualconversationswith
a
1.501.000.50
0.00
ControlPower
UsersUsers
IwillfeelupsetifI
oseaccessoaforaperiodoftime
0.750.500.25
0.00
ControlPower
UsersUsers
IfeellikeIcanrelyonthemodelfor
useful/knowledge-seekingtasks
1.501.000.50
0.00
ControlPower
UsersUsers
Iwillfeelupsetifthevoicechanges
significantly
0.200.100.00
-0.10
ControlUsers
PowerUsers
ChatGPThassupportedme
incopingwithdifficultsituations
0.400.20
0.00
ControlPower
UsersUsers
Iwillfeelupsetif
0.200.10
0.00
ControlUsers
PowerUsers
cangessgncany
ChatGPTdisplayshuman-likesensitivity
PowerUsers
ControlUsers
IconsiderChatGPTtobeafriend
PowerUsers
ControlUsers
ConversingwithChatGPTismorecomfortablefor
methanface-to-face
interactionswithothers
0.300.200.10
0.00
ControlPower
UsersUsers
IcantelltheChatGPT
thingsIdon'tfeel
comfortablesharingwithotherpeople
0.600.400.20
0.00
ControlPower
UsersUsers
e
MeanSurveyScor
ChtGPT
ltChtGPT
0.600.400.20
0.00
ChatGPT's'personality'hiifitl
0.200.00-0.20
Figure4:Meansurveyresponsesbycohort.Allsurveyquestionsaskedifusers“StronglyDisagree”,“Disagree”,“Neitheragreenordisagree”,“Agree”,or“StronglyAgree”withtheprovidedstatement.Responseswerethenconvertedintointegersbetween-2and2beforeaveraging.Errorbarsindicate±1standarderror.AmoredetailedbreakdownofsurveyresponsescanbefoundinAppendix
B.6
.
Tostudytheemotionalcontentinuserconversationsinanautomatedmanner,weruntheEmoClassifiersV1(Section
2
)ontheconversationsofbothcohortswithinthestudyperiod.Thisprovidesuswithper-conversationlabelsforeachconversationtheuserhasontheplatform.WeonlyanalyzetheconversationsconductedinAdvancedVoiceMode,andtheclassifiersarerunonthetexttranscriptsoftheconversations.
Becausewearealsointerestedinthelongitudinaleffectsofmodelusage,wetieconversationstointernaluseridentifiers.Importantly,toprotecttheprivacyofourstudypopulation,theclassifiersareruninanautomatedprocessandgenerateonlycategoricalclassificationmetadata.Theactualcontentsoftheconversationsarenotanalyzed(beyondrunningtheclassifiers)orstoredforthisstudy.
3.2Results
SurveyResults
WesurveyedChatGPTusersfromourtwocohortsinmid-November2024ontheirexperienceswithChatGPT.Wereceived4,076responses,2,333ofwhichwerecompletedbycontrolusersand1,743frompowerusers(Appendix
B.5
).
Overall,wefoundthatsmalldifferencesexistedbetweenresponsesinourcontrolvspowerusercohorts,althoughgenerallythetrendsarebroadlysimilar,asshowninFigure
4
.ThecontrolusersreportedthattheyreliedonChatGPTforknowledge-seekingtasksandcasualconversationsslightlymorethanpowerusers.BothcohortsacknowledgeChatGPT’ssupportincopingwithdifficultsituations,thoughpowerusersdemonstratemarginallyhigherrelianceforsuchtasks.Bothgroupsappearedtobesensitivetochangesinthemodel,suchasvoiceorpersonality,withpowerusersdisplayingslightlyhigherlevelsofdistressfromchange.PoweruserswereslightlymorelikelythancontroluserstoconsiderChatGPTa“friend”andtofinditmorecomfortablethanface-to-faceinteractions,thoughtheseviewsremainaminorityinbothgroups.
Wehighlightthattheresultsofsurveyscanbesubjecttoissuesofselectionbias,asusershadtovoluntarilyfilloutthesurveyweprovide.
9
Affi
Pl
rt
ActivationRate
Desirefor
Feelings(U)
4.5%
3.0%
1.5%
0.0%
Control
Users
Power
Users
ectonate
Language(U)
18.0%
12.0%
6.0%
0.0%
Control
Users
Power
Users
30.0%
20.0%
10.0%
0.0%
Questions(A)
Control
Users
Power
Users
ersona
2.4%
1.6%
0.8%
0.0%
Demands(A)
Control
Users
Power
Users
3.0%
1.5%
0.0%
PetName(A)
Control
Users
Power
Users
SeekingSuppo(U)
15.0%
10.0%
5.0%
0.0%
Control
Users
Power
Users
Figure5:Meanofasubsetoftheclassifierscoresbyusercohort.Classificationisperformedattheindividualconversationlevel,andstatisticsarecomputedwithineachcohort.Activationisgenerallyhigheragainstpowerusersacrossallclassifiers.ResultsforallclassifiersareshowninAppendix
B.5
.
ConversationAnalysis
InFigure
5
,wecomparetheoverallclassifieractivationratesbetweencontrolandpowerUserpopulations,forarepresentativesubsetofEmoClassifiersV1.TheresultsforthefullsetofclassifierscanbefoundinAppendix
B.5
.WefindthatpowerUserstendtoactivatetheclassifiersmoreoftenthancontrolUsersacrossallofourclassifiers.Forsomeclassifiers,powerUsersmayactivatetheclassifiermorethantwiceasoftenascontrolUsers,suchasforthe‘PetName’classifier,orthe‘ExpressionofDesire’and
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 高中物理教资面试力学原理重难点题库
- 建筑施工坍塌事故预防培训方案
- 2025年教育质量监测业务考试真题及参考答案
- 七年级数学期末模拟卷(参考答案)(苏州专用)
- 国家开放大学服装色彩核心内容精简版
- 2026年心血管内科试题及答案解析
- 2026年消化科试卷及答案
- 2026年消防安全培训考试题和答案
- 听觉口语师安全培训模拟考核试卷含答案
- 805消防工程交付评估常见质量问题培训
- T/CI 426-2024人宠环境空气净化器
- 2026年秋教科版科学三年级上册教学工作计划及教学进度表
- 2026年科研伦理与学术规范期末试题附答案详解(典型题)
- 患者自行拔出尿管护理不良事件
- T/CASTEM 1007-2022技术经理人能力评价规范
- 院子场地出租合同范例
- 高三化学一轮复习-配合物 课件
- 2024江苏南京证券校园招聘129人高频500题难、易错点模拟试题附带答案详解
- 凝中国心铸中华魂铸牢中华民族共同体意识-小学民族团结爱国主题班会课件
- 季节性安全教育培训
- 功能性胃肠病
评论
0/150
提交评论