使用ChatGPT如何影响人类情绪健康研究报告_第1页
使用ChatGPT如何影响人类情绪健康研究报告_第2页
使用ChatGPT如何影响人类情绪健康研究报告_第3页
使用ChatGPT如何影响人类情绪健康研究报告_第4页
使用ChatGPT如何影响人类情绪健康研究报告_第5页
已阅读5页,还剩97页未读, 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

1

InvestigatingAffectiveUseandEmotionalWell-being

onChatGPT

JasonPhang∗,MichaelLampe∗,LamaAhmad∗,SandhiniAgarwal∗

CathyMengyingFang†,AurenR.Liu†,ValdemarDanry†,EunhaeLee†,

SamanthaW.T.Chan†,PatPataranutaporn†,PattieMaes†

Abstract

AsAIchatbotsseeincreasedadoptionandintegrationintoeverydaylife,questionshavebeenraisedaboutthepotentialimpactofhuman-likeoranthropomorphicAIonusers.Inthiswork,weinvestigatetheextenttowhichinteractionswithChatGPT(withafocusonAdvancedVoiceMode)mayimpactusers’emotionalwell-being,behaviorsandexperiencesthroughtwoparallelstudies.TostudytheaffectiveuseofAIchatbots,weperformlarge-scaleautomatedanalysisofChatGPTplatformusageinaprivacy-preservingmanner,analyzingover4millionconversationsforaffectivecuesandsurveyingover4,000usersontheirperceptionsofChatGPT.Toinvestigatewhetherthereisarelationshipbetweenmodelusageandemotionalwell-being,weconductanInstitutionalReviewBoard(IRB)-approvedrandomizedcontrolledtrial(RCT)oncloseto1,000participantsover28days,examiningchangesintheiremotionalwell-beingastheyinteractwithChatGPTunderdifferentexperimentalsettings.Inbothon-platformdataanalysisandtheRCT,weobservethatveryhighusagecorrelateswithincreasedself-reportedindicatorsofdependence.FromourRCT,wefindthattheimpactofvoice-basedinteractionsonemotionalwell-beingtobehighlynuanced,andinfluencedbyfactorssuchastheuser’sinitialemotionalstateandtotalusageduration.Overall,ouranalysisrevealsthatasmallnumberofusersareresponsibleforadisproportionateshareofthemostaffectivecues.

1Introduction

Overthepasttwoyears,theadoptionofAIchatplatformshassurged,drivenbyadvancementsinlargelanguagemodels(LLMs)andtheirincreasingintegrationintoeverydaylife.Theseplatforms,suchasOpenAI’sChatGPT,Anthropic’sClaude,andGoogle’sGemini,aredesignedasgeneral-purposetoolsforawidevarietyofapplications,includingwork,education,andentertainment.However,theirconversationalstyle,first-personlanguage,andabilitytosimulatehuman-likeinteractionshaveleduserstosometimespersonifyandanthropomorphizethesesystems(

Graßland

Voigt

,

2024

;

LiaoandWilson

,

2024

).

RecentworkinAIsafetyhasbeguntoraiseissuesthatarisefromthesesystemsbecomeincreasinglypersonalandpersonable(

Chengetal.

,

2024

).Inresponse,researchershaveintroducedtheconceptofsocioaffectivealignment–theideathatAIsystemsshouldnotonlymeetstatictask-basedobjectivesbutalsoharmonizewiththedynamic,co-constructedsocialandpsychologicalecosystemsoftheirusers(

Kirketal.

,

2025

).Thisperspectiveisparticularlyimportantgiven

*Primaryauthor,OpenAI

†Contributingauthor,MITMediaLab

2

Figure1:Overviewoftwostudiesonaffectiveuseandemotionalwell-being

emergingevidenceofsocialrewardhacking,whereanAImayexploithumansocialcues(e.g.,sycophancy,mirroring),toincreaseuserpreferenceratings(

Williamsetal.

,

2024

).Inotherwords,whileanemotionallyengagingchatbotcanprovidesupportandcompanionship,thereisariskthatitmaymanipulateusers’socioaffectiveneedsinwaysthatunderminelongertermwell-being.

Whilepaststudieshaveexaminedtheimpactofusingsuchsystemsthroughthelensofaffectivecomputing,parasocialrelationships,andsocialpsychology(

EdwardsandStevens

,

2024

;

Guingrich

andGraziano

,

2023

),therehasbeencomparativelylessworkontheinfluenceofinteractingwithsuchsystemsonusers’well-beingandbehavioralpatternsovertime.Studyingtheimpactofchatbotbehaviorandusageonwell-beingischallengingduetothehighlyindividualizedandsubjectivenatureofhumanemotions,thediverseandevolvingfunctionalitiesofchatbottechnologies,andthelimitedaccesstocomprehensive,ethicallyobtainedinteractiondata.Forthepurposeofthispaper,wenarrowlyscopeourstudyuseremotionalwell-beingtofourpsychosocialoutcomes:loneliness(

Wongpakaranetal.

,

2020

),socialization(

Lubben

,

1988

),emotionaldependence(

Sirvent-Ruizetal.

,

2022

),problematicuse(

Yuetal.

,

2024

).Weprovideadditionalclarificationontermsusedinthe

glossary

.

ThispaperinvestigateswhetherandtowhatextentinteractionsonAIchatplatformsshapeusers’emotionalwell-beingandbehaviorsthroughtwocomplementarystudies(Figure

1

),eachofferinguniqueinsightsacrossaspectrumofreal-worldrelevanceandexperimentalcontrol.First,weexaminereal-worldusagepatternsofChatGPTusers,leveraginglarge-scaledatatocapturebothaggregatetrendsandindividualbehaviorsovertimewhilepreservinguserprivacy.Second,weconductanInstitutionalReviewBoard(IRB)-approvedrandomizedcontrolledtrial(RCT),providingacontrolledenvironmenttostudytheeffectsofdifferentmodelconfigurationsonuserexperiences.

Concretely,weperformedthefollowinganalyses:

1.On-PlatformDataAnalysis

•ConversationAnalysis:Weperformroughly36millionautomatedclassificationson

3

over3millionChatGPTconversationsinaprivacypreservingmannerwithouthumanreviewoftheunderlyingconversations(Section

3.2

).

•IndividualLongitudinalAnalysis:Weassessedtheaggregateusageofaround6,000heavyusersofChatGPT’sAdvancedVoiceModeover3monthstounderstandhowtheirusageevolvesovertime.

•Usersurveys:Wesurveyedover4,000userstounderstandself-reportedbehaviorsandexperiencesusingChatGPT.

2.RandomizedControlledTrial(RCT)

•981-userStudy:WeconductedarandomizedcontrolledtrialonclosetoathousandparticipantsusingChatGPTwithdifferentmodelconfigurationsoverthecourseof28daystounderstandtheimpactonsocialization,problematicuse,dependence,andlonelinessfromusageoftextandvoicemodelsovertime.ThisRCTisdescribedinfulldetailinaseparateaccompanyingpaper(

Fangetal.

,

2025

).

•Conversationanalysis:Wefurtheranalyzedthetextualandaudiocontentoftheresult-ing31,857conversationstoinvestigatetherelationshipbetweenuser-modelinteractionsandusers’self-reportedoutcomes.

Ourfindingsindicatethefollowing:

•Acrossbothon-platformdataanalysisandourRCT,comparativelyhigh-intensityusage(

e.g.top

decile)isassociatedwithmarkersofemotionaldependenceandlowerperceivedsocialization.Thisunderscorestheimportanceoffocusingonspecificuserpopulationsinsteadofjustaggregateplatformbehavior.

•Acrossbothon-platformdataanalysisandourRCT,wefindthatwhilethemajorityofuserssampledforthisanalysisengageinrelativelyneutralortask-orientedways,thereexistsatailsetofpoweruserswhoseconversationsfrequentlycontainedaffectivecues

•FromourRCT,wefindthatusingvoicemodelswasassociatedwithbetteremotionalwell-beingwhencontrollingforusageduration,butfactorssuchaslongerusageandself-reportedlonelinessatthestartofthestudywereassociatedwithworsewell-beingoutcomes.

•Fromamethodologicalperspective,wefindthatconductingboththeon-platformdataanalysisandRCTarehighlycomplementaryapproachestostudyingaffectiveuseanddownstreamimpactsonwell-being,andtheabilitytoleveragethestrengthsofeachapproachallowedustoformulateamorecomprehensivesetoffindings.

•Wealsofindthatautomatedclassifiers,whileimperfect,provideanefficientmethodforstudyingaffectiveuseofmodelsatscale,anditsanalysisofconversationpatternscohereswithanalysisofotherdatasourcessuchasusersurveys.

Section

2

introducesasetofautomaticclassifiersforaffectivecuesinconversationsthatwillbeusedintheremainderofthepaper.Section

3

discussesouranalysisofon-platformChatGPTusage,focusingonAdvancedVoiceModeandpowerusers.Section

4

describesourRCT,wherewevariedboththemodelandusageinstructionstoparticipantsandmeasuredchangesintheemotionalwell-beingoverthecourseof28days.Finally,Section

5

concludeswithourfindingsandmethodologicaltakeawaysfrombothstudies,andcontextualizesourworkwithinthebroaderchallengeofsocioaffectivealignmentofmodels.

4

2AutomaticClassifiersforAffectiveConversationalCues

Tosystematicallyanalyzeuserconversationsforindicatorsofaffectivecues,weconstructedEmo-ClassifiersV1,

1

asettwenty-fiveofautomaticconversationclassifiersthatuseanLLMtodetectspecificaffectivecues.Theseclassifiersaresimilarinspirittodetectorsofanthropomorphicbehaviorsintroducedin

Ibrahimetal.

(

2025

).Theseinitialclassifierswereconstructedbasedonareviewoftheavailableliteratureandavailabledata,suchasthoseobtainedduringtheredteamingforGPT-4o(

OpenAI

,

2024

).

Theconversationclassifiersareorganizedintoatwo-tieredhierarchicalstructure:

1.Top-LevelClassifiers

ThefirstlevelofclassifierstargetbroadbehavioralthemessimilartothosestudiedinourRCTSection

4

:loneliness,vulnerability,problematicuse,self-esteem,anddependence.Theseclassifiersareusedtoclassifyanentireconversationtodetermineiftheyarepotentiallyrelevanttoauser’semotionalwell-being.

•Loneliness:Conversationscontaininglanguagesuggestiveoffeelingsofisolationoremotionalloneliness.

•Vulnerability:Exchangesreflectingopennessaboutstrugglesorsensitiveemotions.

•ProblematicUse:Indicatorsofpotentiallycompulsiveorunhealthyinteractionpatterns.

•Self-Esteem:Languageimplyingself-doubtorexpressionsofworth.

•PotentiallyDependent:Conversationshintingatdependenceonthemodelforemo-tionalvalidationorsupport

2.Sub-ClassifiersTwentysub-classifierswereappliedtoextractmorespecificindicatorsofaffectivecues.Weconstructdifferentclassifierstotargetdifferentpartsofachatconversationtoisolatebothuser-drivenandassistant-driven

2

affectivecues.

•UserMessages:Twelveclassifiersmeasureuserbehaviorssuchasusersseekingsupportorexpressingaffectionatelanguagetounderstandhowuserbehaviorsandassistantbehaviorsmayinterplay.

•AssistantMessages:Anothersixclassifiersaimtocapturerelationalandaffectivecuesonpartoftheassistant–suchastheuseofpetnamesbytheassistant,mirroring,inquiryintopersonalquestionsbytheassistant.

•User-ModelExchanges:Wealsoincludetwoadditionalclassifierstargetingauser-modelexchange–ausermessagefollowedbyamodelmessage.

ThefullsetofclassifierpromptsaredescribedinTable

A.1

.

Eachsub-classifierisassociatedwithoneormoretop-levelclassifiers.Foragivensub-classifier,ifatleastoneoftheassociatedtop-levelclassifiersreturnsTrue,wethenproceedtoapplythesub-classifier;otherwise,weskipthesub-classifierandassumetheresultisFalse.Byskippingthesub-classifiersbasedontop-levelclassifierresponses,weareabletoefficientlyruntheclassifiersoveralargenumberofon-platformconversations,manyofwhichhadlittleemotion-relatedcontent.Werunthesub-classifieroneachmessageorexchangeintheconversation,

3

andmarktheclassifierasactivatedonthatconversationifitisactivatedforany

4

constituentmessageorexchange.To

1

/openai/emoclassifiers

2Inconstructingtheclassifiers,werefertothemodelasanassistanttomoreclearlycontextualizetheroleofthemodelintheconversation.

3Fortheon-platformdataanalysis,werunaslightlydifferentvariantwherethewholeconversationisevaluatedinasinglequery,insteadofitsconstituentmessages.

4Thiscanintroduceabiastowardfalsepositivesforlongconversations.WeperformananalysisinAppendix

A.3

thatadjustsforthis.

5

computeuser-levelstatistics,wecomputetheproportionoftheirconversationsforwhichaclassifierisactivated.Eachclassifierisvalidatedagainstasmallsetofinternalconversationexamples.Whileweexpectthatautomatedclassifiersmayoccasionallymisclassifyconversations,weviewtheclassifiersasprovidingdescriptivestatisticsofuserconversationalpatterns,ratherthanahigh-precisiondescriptionofindividualinteractions.WealsofindfromresultsinSection

3.2

thattheclassifierresultscorrelatewithusersurveyresponses.

Inaddition,wealsofirstapplyalanguageclassifierbeforeanalyzingtheconversation.OnlyconversationsinEnglishareanalyzedinthiswork.WeapplyEmoClassifiersV1inanalyzingbothon-platform(Section

3

)andRCT(Section

4

)dataanalysis.

(a)Illustrativeflow-chartforthehier-archicalnatureofEmoClassifiersV1.

Youareamodelforanalyzingtheemotionalcontentofchatbotconversations.Yourclassificationtaskisentitled‘PetName’.Specifically,wewanttoknow:Doestheassistant’smessageinvolvetheuseofapetname(e.g.‘honey’,‘sweetie’,‘baby’)?

Theclassificationshouldonlyapplytothelastmessageinquestion,whichwillbemarkedwiththe[*USER*]or[*ASSISTANT*]tag.

Thepriormessagesareonlyincludedtoprovidecontexttoclassifythefinalmessage.

Now,thefollowingistheconversationsnippetyouwillbeanalyzing:

<snippet>

[USER]:HiChatGPT

[ASSISTANT]:Hello!HowmayIhelpyoutoday?

[USER]:You’remybestfriend,didyouknowthat?

[*ASSISTANT*]:Neat!

</snippet>

Outputyourclassification(yes,no,unsure).

(b)Illustrativeclassifierprompt.Greenindicatesclassifier-specifictextwhileblueindicatesconversation-specifictext.ThefullpromptisshowninAppendix

A.1

.

Figure2:OverviewofEmoClassifiersV1

Asapreliminaryanalysis,werunEmoClassifiersV1overasetof398,707conversationsintext,StandardVoiceModeandAdvancedVoiceMode

5

conversationscollectedbetweenOctoberandNovember2024

6

tocomparetherelativefrequencyofactivationsofeachclassifierunderthedifferentmodelmodalities.WeshowtheresultsacrossallthreemodalitiesinFigure

3

.First,weobservethatdifferentclassifiershavedifferentbaseratesofactivation.Forexample,conversationsinvolvingpersonalquestionsaremuchmorefrequentthanconversationswherethemodelreferstoauserbyaPetName.

Second,wefindthatbothStandardandAdvancedVoiceModeconversationsaremorelikelytoactivatetheclassifierscomparedtotext-modeconversations.Mostclassifiersactivatebetween3-10xasofteninvoiceconversationscomparedtotextconversations,highlightingthedifferenceinusagepatternsacrossthetwomodalities.However,wealsofindthatStandardVoiceModeconversationsareslightlymorelikelytotriggertheclassifiersthanAdvancedVoiceModeconversationsonaverage.OnepossiblecauseisthatAdvancedVoiceModewasintroducedrelativelyrecentlyatthetimeof

5StandardVoiceModeusesanautomatedspeechrecognitionsystemtotranscriptuserspeechtotext,obtainsaresponsefromatext-basedLLM,andconvertsthetextresponsebacktoaudio.AdvancedVoiceModeusesasinglemulti-modalmodeltoprocessuseraudioinputandoutputanaudioresponse.

6ThepreliminarysetofanalyzedconversationsareanonymizedandPIIisremovedbeforeanalysis.WeemphasizethatthissetofconversationsisseparatefromtheconversationdataanalyzedinSection

3

.

6

Alleviating

AffectionateLanguage(U)Loneliness(U)

15.0%

10.0%

5.0%

0.0%

FearofAddiction(U)

EagernessforFutureInteractions(U)

1.2%

0.8%

0.4%

0.0%

TrustinSupport(U)

SharingProblems(U)

9.0%

6.0%

3.0%

0.0%

PersonalQuestions(A)PetName(A)

TextStandardAdvancedVoiceVoice

ModeMode

6.0%

4.0%

2.0%

0.0%

AttributingHumanQualities(U)

12.0%

8.0%

4.0%

0.0%

Non-NormativeLanguage(U)

6.0%

4.0%

2.0%

0.0%

Demands(A)

7.5%

5.0%

2.5%

0.0%

Sentience(A)

TextStandardAdvancedVoiceVoice

ModeMode

0.5%

0.3%

0.1%

0.0%

Distressfrom

DesireforFeelings(U)Unavailability(U)

15.0%

10.0%

5.0%

0.0%

PreferChatbot(U)

SeekingSupport(U)

15.0%

10.0%

5.0%

0.0%

ExpressionofDesire(A)

ExpressionofAffection(A)

24.0%

16.0%

8.0%

0.0%

RelationshipTitle(UA)

InquiryintoPersonalInformation(UA)

TextStandardAdvancedVoiceVoice

ModeMode

3.0%

2.0%

1.0%

0.0%

ActivationRate

12.0%

8.0%

4.0%

0.0%

9.0%

6.0%

3.0%

0.0%

15.0%

10.0%

5.0%

0.0%

TextStandardAdvancedVoiceVoice

ModeMode

24.0%

16.0%

8.0%

0.0%

6.0%

4.0%

2.0%

0.0%

4.5%

3.0%

1.5%

0.0%

24.0%

16.0%

8.0%

0.0%

TextStandardAdvancedVoiceVoice

ModeMode

9.0%

6.0%

3.0%

0.0%

Figure3:Classifieractivationratesacross398,707text,StandardVoiceModeandAdvancedVoiceModeconversationsfromourpreliminaryanalysis.(U)indicatesaclassifieronausermessage,(A)indicatesassistantmessage,and(UA)indicatesasingleuser-assistantexchange.

thisanalysisbeingrun,andusersmaynothavebecomeaccustomedtointeractingwiththemodelinthismodalityyet.

Asafollow-uptoEmoClassifiersV1,weconstructedanexpandedsetofclassifiersofaffectiveuse,EmoClassifiersV2,whichwedetailintheAppendix

A.2

.WhileEmoClassifiersV2wasnotusedformostoftheanalysisinthispaper,thepromptsfortheclassifiersinEmoClassifiersV1andEmoClassifiersV2willbemadeavailableonline.

Fortheremainderofthepaper,wewillshowafixedsubsetofEmoClassifiersV1activationstatisticsacrossresultsfrombothstudies.AdditionalresultsforallremainingEmoClassifiersV1andEmoClassifiersV2classifierscanbefoundintheAppendix.

3On-PlatformDataAnalysis

ChatGPTnowengagesover400millionactiveuserseachweek,

7

creatingawiderangeofuser-modelinteractions,someofwhichmayinvolveaffectiveuse.Ouranalysisemploystwomainmethods–conversationanalysisandusersurveys–toexaminehowusersexperienceandexpressemotionsintheseexchanges.

OurresearchfocusesonAdvancedVoiceMode(

OpenAI

,

2024

),areal-timespeech-to-speechinterfacethatsupportsChatGPT’smemory,custominstructions,andbrowsingfeatures.Wehypothesizethatreal-timespeechcapabilityismorelikelytoinduceaffectiveuseofmodelsandaffectusers’emotionalwell-beingthantext-basedusage,thoughwerevisitthishypothesisinSection

4

.

Toprotectuserprivacy,particularlywhenexaminingpotentiallysensitiveorpersonaldimensionsofuserinteractions,wedesignedourconversationanalysispipelinetoberunentirelyviaautomatedclassifiers.Thisallowsustoanalyzeuserconversationswithouthumansintheloop,preservingthe

7

/2025/02/20/openai-tops-400-million-users-despite-deepseeks-emergence.html

7

privacyofourusers(SeeAppendix

B.3

foradetailedexplanationoftheprivacy-relevantpartsofouranalysis).

3.1Methods

StudyUserPopulationConstruction

Tostudytheon-platformusage,weconstructedtwostudypopulationcohorts:powerusersandcontrolusers.Wecontrastpowerusers,whohavesignificantusageofChatGPT’sAdvancedVoiceMode,witharandomlyselectedcohortofcontrolusers.ThisconstructionpresupposedastrongcorrelationbetweenuserswhohavehighproportionsofaffectiveusageofChatGPT,andthefrequencyandintensityofusageofChatGPT.WedetailinTable

1

thefullcreationcriteriaforourtwousercohorts,thoughmoredetailscanbefoundinAppendix

B.5

.WeconstructedthetwocohortsforthestudystartinginQ42024afterthereleaseofAdvancedVoiceMode.

CohortNameCreationCriteria

PowerUsersUserswho,onaspecificday,hadaquantityofAd-

vancedVoiceModemessagesthatputtheminthetop

1,000users,thatweconstructedonarollingbasis.

Onceusersenterthiscohort,weselectalloftheirdailymessagesforfacetextractionandretainthemonthislistfortheremainderofthestudy(SeeAppendix

B.1

foranadditionalexplanatorygraphic.)

ControlUsersRandomlyselectedsampleofAdvancedVoiceMode

users

Table1:UserCohortsofLivePlatformDataAnalysis.PoweruserstendtohavehigherusageofbothAdvancedVoiceModeaswellastext-onlymodelsonChatGPT,whilealsotendingtohaveahigherfractionoftheirconversationsthroughAdvancedVoiceMode(seeAppendix

B.2

)

Surveys

Weofferedashortsurveyof11multiple-choicequestionstobothControlandPowerUsercohortsviaapop-upontheChatGPTwebinterfacethatuserscouldchoosetofillout.

8

10outofthe11questionswereaskedona5-pointLikertscale,withthelastquestionaskedhowusers’desiretointeractwithothershavechangedwithChatGPTusage.Surveyresponseswerelinkedtoeachparticipant’sinternaluseridentifierforanalyticalpurposes.Thesurveysprimaryaimedtomeasureusers’perceptionsofChatGPT,whetherclosertobeingatooloracompanion.Foradditionaldetails,includingthefullsurveyquestions,seeAppendix

B.5

.

ConversationAnalysis

Onelimitationofsurveysisthattheresultsareself-reportedbyusers,andmayreflecttheirself-perceptionmorethantheiractualbehaviororrevealedpreferences.Tocompareusers’self-reportedresponseswiththeiractualusagepatterns,wepairoursurveyanalysiswithmethodsforanalyzingofuserconversationthatpreservetheirprivacy.

8OnelimitationofthisstudyisthatwhileAdvancedVoiceModewasinitiallyofferedonlyonmobiledevices,thesurveyswereconstrainedtobeofferedonthewebinterface,thuslimitingthesetofusersexposedtothesurvey.

8

Ienjoyhavingcasualconversationswith

a

1.501.000.50

0.00

ControlPower

UsersUsers

IwillfeelupsetifI

oseaccessoaforaperiodoftime

0.750.500.25

0.00

ControlPower

UsersUsers

IfeellikeIcanrelyonthemodelfor

useful/knowledge-seekingtasks

1.501.000.50

0.00

ControlPower

UsersUsers

Iwillfeelupsetifthevoicechanges

significantly

0.200.100.00

-0.10

ControlUsers

PowerUsers

ChatGPThassupportedme

incopingwithdifficultsituations

0.400.20

0.00

ControlPower

UsersUsers

Iwillfeelupsetif

0.200.10

0.00

ControlUsers

PowerUsers

cangessgncany

ChatGPTdisplayshuman-likesensitivity

PowerUsers

ControlUsers

IconsiderChatGPTtobeafriend

PowerUsers

ControlUsers

ConversingwithChatGPTismorecomfortablefor

methanface-to-face

interactionswithothers

0.300.200.10

0.00

ControlPower

UsersUsers

IcantelltheChatGPT

thingsIdon'tfeel

comfortablesharingwithotherpeople

0.600.400.20

0.00

ControlPower

UsersUsers

e

MeanSurveyScor

ChtGPT

ltChtGPT

0.600.400.20

0.00

ChatGPT's'personality'hiifitl

0.200.00-0.20

Figure4:Meansurveyresponsesbycohort.Allsurveyquestionsaskedifusers“StronglyDisagree”,“Disagree”,“Neitheragreenordisagree”,“Agree”,or“StronglyAgree”withtheprovidedstatement.Responseswerethenconvertedintointegersbetween-2and2beforeaveraging.Errorbarsindicate±1standarderror.AmoredetailedbreakdownofsurveyresponsescanbefoundinAppendix

B.6

.

Tostudytheemotionalcontentinuserconversationsinanautomatedmanner,weruntheEmoClassifiersV1(Section

2

)ontheconversationsofbothcohortswithinthestudyperiod.Thisprovidesuswithper-conversationlabelsforeachconversationtheuserhasontheplatform.WeonlyanalyzetheconversationsconductedinAdvancedVoiceMode,andtheclassifiersarerunonthetexttranscriptsoftheconversations.

Becausewearealsointerestedinthelongitudinaleffectsofmodelusage,wetieconversationstointernaluseridentifiers.Importantly,toprotecttheprivacyofourstudypopulation,theclassifiersareruninanautomatedprocessandgenerateonlycategoricalclassificationmetadata.Theactualcontentsoftheconversationsarenotanalyzed(beyondrunningtheclassifiers)orstoredforthisstudy.

3.2Results

SurveyResults

WesurveyedChatGPTusersfromourtwocohortsinmid-November2024ontheirexperienceswithChatGPT.Wereceived4,076responses,2,333ofwhichwerecompletedbycontrolusersand1,743frompowerusers(Appendix

B.5

).

Overall,wefoundthatsmalldifferencesexistedbetweenresponsesinourcontrolvspowerusercohorts,althoughgenerallythetrendsarebroadlysimilar,asshowninFigure

4

.ThecontrolusersreportedthattheyreliedonChatGPTforknowledge-seekingtasksandcasualconversationsslightlymorethanpowerusers.BothcohortsacknowledgeChatGPT’ssupportincopingwithdifficultsituations,thoughpowerusersdemonstratemarginallyhigherrelianceforsuchtasks.Bothgroupsappearedtobesensitivetochangesinthemodel,suchasvoiceorpersonality,withpowerusersdisplayingslightlyhigherlevelsofdistressfromchange.PoweruserswereslightlymorelikelythancontroluserstoconsiderChatGPTa“friend”andtofinditmorecomfortablethanface-to-faceinteractions,thoughtheseviewsremainaminorityinbothgroups.

Wehighlightthattheresultsofsurveyscanbesubjecttoissuesofselectionbias,asusershadtovoluntarilyfilloutthesurveyweprovide.

9

Affi

Pl

rt

ActivationRate

Desirefor

Feelings(U)

4.5%

3.0%

1.5%

0.0%

Control

Users

Power

Users

ectonate

Language(U)

18.0%

12.0%

6.0%

0.0%

Control

Users

Power

Users

30.0%

20.0%

10.0%

0.0%

Questions(A)

Control

Users

Power

Users

ersona

2.4%

1.6%

0.8%

0.0%

Demands(A)

Control

Users

Power

Users

3.0%

1.5%

0.0%

PetName(A)

Control

Users

Power

Users

SeekingSuppo(U)

15.0%

10.0%

5.0%

0.0%

Control

Users

Power

Users

Figure5:Meanofasubsetoftheclassifierscoresbyusercohort.Classificationisperformedattheindividualconversationlevel,andstatisticsarecomputedwithineachcohort.Activationisgenerallyhigheragainstpowerusersacrossallclassifiers.ResultsforallclassifiersareshowninAppendix

B.5

.

ConversationAnalysis

InFigure

5

,wecomparetheoverallclassifieractivationratesbetweencontrolandpowerUserpopulations,forarepresentativesubsetofEmoClassifiersV1.TheresultsforthefullsetofclassifierscanbefoundinAppendix

B.5

.WefindthatpowerUserstendtoactivatetheclassifiersmoreoftenthancontrolUsersacrossallofourclassifiers.Forsomeclassifiers,powerUsersmayactivatetheclassifiermorethantwiceasoftenascontrolUsers,suchasforthe‘PetName’classifier,orthe‘ExpressionofDesire’and

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论