版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
1
arXiv:2308.06088v1[cs.AI]11Aug2023
DepartmentofEducationalSciences
TUMSchoolofSocialSciencesandTechnology
TechnicalUniversityofMunich
AssessingStudentErrorsinExperimentationUsingArtificialIntelligenceandLargeLanguageModels:AComparativeStudywithHumanRaters
ArneBewersdorff1,KathrinSeßler1,ArminBaur2,EnkelejdaKasneci1,andClaudiaNerdel1
1TUMSchoolofSocialSciencesandTechnology,TechnicalUniversityofMunich
2UniversityofEducationHeidelbergAugust11,2023
Abstract—Identifyinglogicalerrorsincomplex,incompleteorevencontradictoryandoverallhetero-geneousdatalikestudents’experimentationprotocolsischallenging.Recognizingthelimitationsofcur-rentevaluationmethods,weinvestigatethepotentialofLargeLanguageModels(LLMs)forautomaticallyidentifyingstudenterrorsandstreamliningteacheras-sessments.Ouraimistoprovideafoundationforproductive,personalizedfeedback.Usingadatasetof65studentprotocols,anArtificialIntelligence(AI)
systembasedontheGPT-3.5andGPT-4serieswasde-velopedandtestedagainsthumanraters.OurresultsindicatevaryinglevelsofaccuracyinerrordetectionbetweentheAIsystemandhumanraters.TheAIsys-temcanaccuratelyidentifymanyfundamentalstudenterrors,forinstance,theAIsystemidentifieswhenastudentisfocusingthehypothesisnotonthedepen-dentvariablebutsolelyonanexpectedobservation(acc.=0.90),whenastudentmodifiesthetrialsinanongoinginvestigation(acc.=1),andwhetherastudentisconductingvalidtesttrials(acc.=0.82)re-liably.Theidentificationofother,usuallymorecom-plexerrors,likewhetherastudentconductsavalidcontroltrial(acc.=.60),posesagreaterchallenge.ThisresearchexploresnotonlytheutilityofAIineducationalsettings,butalsocontributestotheunder-standingofthecapabilitiesofLLMsinerrordetectionininquiry-basedlearninglikeexperimentation.
Keywords:ArtificialIntelligence·LargeLanguageModels·ScienceEducation·ScientificInquiry·Experimentation·FormativeAssessment·StudentErrors
1Introduction
CompetenciesforplanningandconductingscientificinquirylikeexperimentsserveasvitalcomponentsofsciencecurriculainGermany(KMK,2004)andaroundtheglobe(e.g.US:NationalResearchCouncil,
2013;UK:DepartmentforEducation,2014;Finnland:FinnishNationalBoardofEducation,2014)fosteringthedevelopmentofscientificthinkingandproblem-solvingskillsofstudents(Bewersdorffetal.,2020).Despiteitsimportance,studentsfrequentlyencounterchallengesduringtheplanningandimplementationofexperiments.Thesechallenges–orerrors–havebeenwell-documentedandempiricallyvalidatedthroughnumerousstudiesoverthepasttwodecades(Kranzetal.,2022).
Currenttoolsavailablefortheidentificationoftheseer-
rorsprimarilyconsistofratingschemes,paper-penciltests,oron-the-flyobservationswhichareemployedbyteachersorstudentsthemselvestoevaluatestudentperformance(Hildetal.,2019;Lehtinenetal.,2022).
However,thesemethodshaveseverallimitations.For
one,theyputasignificantburdenonteacherssincetheyrequirethemtometiculouslyreviewandassesseachstudents’workindividually.Additionally,theseratingschemesareoftenfoundtobecomplexandtime-consuming,makingthemlessaccessibleanduser-friendlyforteachers(Baur,2015).Anothercriticalconstraintofthesetoolsistheirrelianceontherelia-bilityandobjectivityoftheusers,primarilyteachersandresearchers.Asaresult,thefeedbackprovidedtostudentsmaynotalwaysbeconsistentoraccurate,limitingitseffectivenessinaddressingcommonstu-denterrorsandhencetheimprovementofexperimen-talcompetencies.Inlightoftheselimitations,thereisagrowingneedforthedevelopmentofmoreefficientandintuitivetoolstoanalyzeandaddresscommonerrorsmadebystudentsduringtheplanningandcon-ductingofexperiments.Suchtoolswouldnotonlyhelptostreamlinetheevaluationprocessforteachersbutcouldeventuallyhelptoprovidevaluablefeedbacktostudents,ultimatelyenhancingtheirunderstandingofexperimentaldesignandfosteringtheirdevelop-mentofcriticalscientificcompetencies.
AImodels,especiallyLargeLanguageModels,haveemergedasatransformativeforceinthefieldofeduca-
2
Toidentifystudenterrorsinexperimentation,thestu-
tion(Abdelghanietal.,2023;Bhatetal.,2022;Dijk-straetal.,2023;Jietal.,2023;MacNeiletal.,2023).TheseadvancedgenerativeAIsystemslikeGPT-3(Brown,2020),ChatGPT(OpenAI,2022),GPT-4(OpenAI,2023)ortherecentlyreleasedLaMDAmodel(Thoppilanetal.,2022),aredeeplearningar-chitecturesthatusemassivetextdatasetsandreinforce-mentlearningwithhumanfeedbacktolearnhowtogeneratehuman-liketexts.Thistrainingprocedureen-ablesthemtounderstandandrespondtoawiderangeofnaturallanguagequeries.Ineducationalsettings,LLM-basedAIsystemsareincreasinglybeinglever-agedforpersonalizedlearning(Murtazaetal.,2022),intelligenttutoringsystems(Marmo,2022),andsup-
portingcontentgeneration(Khosravietal.,2023).Bybeingabletoadapttoindividuallearners’needsand
providinginstantfeedback,LLM-basedAIsystemsholdgreatpotentialtoenhancetheoveralleducationalexperienceandbridgethegapbetweenstudentsand
expertknowledge(Kasnecietal.,2023).OftenLLM-
basedAIsystemsareprimarilyusedtogiverathergeneralfeedbacktostudents.Whilethisisagreatwaytosupportstudents’learning,theserepliesbytheLLM-basedAIsystemsarenotcomparableamongeachotherandthefocusofthefeedbackmightshiftleadingtoreducedreliabilityandvalidity.Tocounterthisissue,wedecidedtoaimforwell-describedstu-
denterrorsandtestanLLM-basedAIsystemforeacherroragainsthumans.
2Framework
2.1Identificationofstudenterrorsduringexperimentation
Thefirststepofanystudentassessmentistheidentifi-cationofthecurrentstateofthestudentperformance,e.g.theirerrorsinagiventask.Thedescriptionofstudenterrorsduringscientificinquiry,especiallyex-perimentation,hasalongtraditionineducationalsci-ences,leadingtoacomprehensivecollectionofstudent
errors(forareviewseeKranzetal.,2022).Whileotherpapersemploytermslike‘problems’,‘difficulties’,or‘challenges’tocharacterizeaspectsoftheexperimen- talprocess(e.g.Jong&vanJoolingen,1998;Kranzetal.,2022),throughoutthisworkandinlinewithSchwi-chowetal.(2022),butinabroaderunderstanding,wewillusetheterm‘error’.Theterm‘error’inthispaperreferstoactionstakenduringtheinquiryprocessthatpotentiallycomplicateorevenhinderstudentsfromreachingaconclusiveresult(cf.Baur,2023).
dentprotocolsoftheexperimentscanserveasavaliddatasource.Theprotocolsprovidedbystudentsshedlightonthefinaloutcomeoftheexperiment,butofferminimaltonoinsightsintotheactualprocesswhileconductingtheexperimentitself.Giventhatstudentprotocolsonlyillustratethefinalstateandoutcomeoftheexperiment,wewerecompelledtoexcludealler-rorsassociatedwiththeexperimentalprocedureitself.Thislimitationarisesfromthefactthattheseprotocolsfailtoidentifyerrorsthatmayhaveoccurredduringtheexperimentationprocessbutwerenotevidentinthestudents’protocols.Thisledtoalistof16commonstudenterrors,whicharepresentedinthefollowing
Table
1
.
2.2CurrentstateandvisionofAIineducation
Consideringthetoolscurrentlyavailableforerroridentification–primarilyratingschemes,paper-penciltests,andon-the-flyobservations(Hildetal.,2019;Lehtinenetal.,2022)–it’sevidentthatthereisaneedtodevelopmoreefficientandintuitivemethodstode-tectcommonstudenterrorsduringtheexperimentationprocess.AIsystemshavethepotentialtoaidinthiseffort.
Ingeneral,therearestrongargumentsfortheuseofAIsystemsintheeducationalcontextastheycouldim-proveaccesstoeducation(Osetskyietal.,2020),fos-
terpersonalizedlearning(Holmesetal.,2016),unlock
teachertime(Sadikuetal.,2021),reduceinequality(fordebatesee:Holstein&Doroudi,2021)andthere-
foreimprovelearningingeneral(Chenetal.,2020).Besidesthesepromises,theuseofhighlyintegratedAIsystemsalsocomeswithsomepotentialrisks.AI
systemsmayleadtoreducedstudentprivacy(X.Zhaietal.,2021)andchallengethepedagogicalrelationsaswellastheteachers’autonomy.TeachersandlearnershaveamanifoldofmisconceptionsandfearsaboutAI(Bewersdorffetal.,2023)whichmightleadtogen-eralskepticismtowardsAIsystemsintheclassroom
amongstakeholders(Doualietal.,2022),ultimatelyhinderingitseffectiveimplementation.AIsystemsmightevenincreaseinequalityamongstudents(Noy&Zhang,2023).
AsuccessfulintegrationofAIsystemsineducationshouldcomplementteachersontheirmissiontofosterstudents’learning.Therefore,itiscrucialtounder-standthatAIsystemsshouldnotreplacetheteacher,or,worse,beseenascompetitors,butthattheAIsys-
temalterstheirroleinthelearningprocess(Burbules
3
Phase
DefinitionLabel
Description/ExampleReference(s)fromthesample
(see4.2)
StateahypothesisHypothesisisnothyp_var_obs“IthinkbecauseofBaur,2018
focusedonthede-thewaterthelid
pendentvariable,
popsopen.”
butonanexpected
observation
Hypothesishyp_var_comb“IsuspectthattheBaur,2018;
consistsofaconescontractdueValanidesetal.,
combinationoftothecoldandthe2014
independentvari-moisture.”
ables
Hypothesishasnohyp_no_dep“Itneedswater.”-
dependentvariable
Nohypothesisishyp_existsproposed
Studentworkswith-J.Zhaietal.,2014outposingahypoth-
esis
Designandcon-Materialismissingmaterial_missductanexperiment
Thestudentdoesnotitemizethema-terialheisusing
Garcia-Mila&Andersen,2007
Missingtesttrialis_testNotrialwithoutBaur,2021anindependent
variable
Missingcontrolis_controlNotrialwhereDasguptaetal.,
trialallvariablesare2016;Germannet
presental.,1996
Studentplansandmissing_componentsThestudentcon-Baur,2021
preparesexperi-ductsanexperiment
mentaltrialsand
todeterminewhat
forgetstheneces-yeastneedstopro-
sarycomponentduceCO2,butwith-
outusingyeast
Trialswiththeno_variationStudentconductsH
K.Wu&C.-L.
samecontent(notrialswiththesameWu,2011
variation)
contentandthe
sameinstruments
Experimentaltri-alter_expThestudent(repeat-Baur,2021
alsarealterededly)altersrunning
experimentaltrials
-theyaddmore
ingredients,remove
astopper,stirthe
mixture,etc.
Onlyonetrialisone_trialThestudentcon-Hammannetal.,
conductedductsonlyonetrial2008
Documentationofno_implStudentdoesnotde-Garcia-Mila&
theimplementa-scribehisimple-Andersen,2007
tionismissingmentation
4
Observeandana-
lyzedata
Observationonlyinoneorafewtri-als
few_obs
ThestudentonlyBaur,2018
observessometri-
als,notall,focusing
primarilyononeor
afew
Result&Conclu-
sion
Resultfocuseson
whichisthebesttrial,nostatementaboutthevari-able(s)
best_result
“ItclosesthemostBaur,2018
inwater.”
Thestudents’ob-
servationorhy-pothesesaregivenastheresult
result_obs_hyp_same
Studentjustrepeatshishypothesesorobservationasare-
sult,like:“Blisters
haveformed.”
Boaventuraetal.,2013;García-Carmonaetal.,2017
Noresult
if_no_result
“Ihavenoresult.
Ithinkmyassump-
tioniswrong.”
Table1Definitionanddescriptionofstudents’errorseligibleforidentificationfromtheirexperimentationprotocols.
etal.,2020;Schiff,2020).Theteachers’rolechangesdependingonthedegreeofautomation.DifferentmodesanddegreesofimplementationofAIsystemsareimaginable(e.g.:SixlevelsofautomationofAI
ineducation:I.Teacheronly,II.Teacherassistance,III.Partialautomation,IV.Conditionalautomation,V.
HighautomationandVI.Fullautomation;Molenar,2022).ThiscanrangefromusingtheAIsystemtoprovidesupportiveinformation(II.)toautomaticallycontrollingtheentirelearningprocess(VI.).InlinewithMolenaaretal.(2017),ourgoalisahybridintel-ligencewithcombinedresponsibilitybetweentheAIsystemandtheteacher.Forasuccessfulintegration,wearguethatitiscrucialtorespecttheteachers’au-tonomyandthereforegivethemfullfreedomtodecidetowhichdegreetheywanttouseAIsystems.Theshifttowardsahybridintelligencewouldgiveteachersmoretimetoconcentrateonclarifyingcon-cepts,fosteringstudents’criticalthinkingskills,en-couragingcreativity,andfosteringanengaginganddynamiclearningenvironmentwhilestillbeinginfull
controlofthelearningprocessandassociatedpeda-gogicalconsiderations.
2.3AIbasedassessmentinscienceeducation
ApromisingapplicationofAIsystemsineducationisassessingstudentoutcomesinscienceeducation.
Therearetwogeneralapproachesofassessment:for-mativeassessmentwhichisongoingduringthelearn-ingprocess,andsummativeassessmentattheendofalearningunit(Harlen&James,1997).Formativeassessment,asapracticeinherentinteachingthat
focusesonthelearningprocess,isintendedtohelpcontinuouslyadapttheteachingtotheneedsofthestudents(Filsecker&Kerres,2012).Ithasbeen,de-spitesomecritics(Bennett,2010),identifiedasoneofthemostsignificantinfluencingfactorsforeffectivelearning(Hattie,2009),especiallyformsof(computer-based)‘rapidformativeassessment’havebeenshowntobehighlyeffective(Yeh,2010).AIdrivensys-temscouldhelpteacherswithformativeassessment(Swieckietal.,2022).
Someeducationalresearchersinthefieldofscience
educationraiseconcernsabouttheuseofAIsystemsforformativeassessment(Lietal.,2023).Centralpointsofcritiquearetheconfinementofthepeda-gogicalfacetofassessmentandthesideliningofpro-fessionalexpertiseaswellasthatAIbasedassess-mentmightonlyevaluatelimitedformsoflearningandleadtoasurveillancepedagogy(Swieckietal.,2022).OthervoicesarguethatAIsystemsarealreadybeingwidelyemployedinformativeassessmentacrossvariouseducationalcontextsandcallforashiftinper-spective,fromviewingAIasaproblemtobesolvedtorecognizingitspotentialforassessmentineducation(X.Zhai&Nehm,2023).Anexamplefortheimple-
5
mentationofanAIsysteminscienceassessmentsistheautomatedtextanalysiswhichisusedforscoring(Zhaietal.2020).TheseAIsystemsarevalidatedbycomparingthecomputer-assignedscorestohuman-assignedscores(Williamsonetal.,2012)
LLM-basedAIsystemsinthefieldofscienceeduca-tionarestillatanearlystage.Mooreetal.(2022)usedGPT-3basedmodelstoevaluatethequalityofstudent-generatedquestionsinacollegechemistrycourse.Theyreportdifficultiesfortheautomaticevaluation,withaccuraciesbetween.32and.4thusdemonstrat-ingonepotentialwaytohelpscalestudentassessmentbyusinglargelanguagemodels.X.Wuetal.(2023)designedanAIsystemforautomaticscoringinscienceeducation.TheyreportCohensKapparangingfrom.3to.57anddemonstrate–whilestillsomeroomforimprovement–thepreliminarypotentialoftheLLM-basedAIsystem.
3Objectives
ThemajorityofcontemporaryAIsystemsinscienceclassroomsconcentrateoncategorizingstudents’gen-eratedresponsesinscientificpractices,particularlyinareaslikeexplanationandargumentation(X.Zhaietal.,2020).X.Zhaietal.(2020)concludethatstud-iesareneededwhichexamineproceduresincomplexdecision-makingprocesses.ThisstudyinvestigatesthepotentialofLLM-basedAIsystemsinsupportingteachersbyanalyzingstudenterrors:WeinvestigatewhetherautomaticerroridentificationbyanLLM-basedAIsystemisasvalidandreliableasthatbyscienceeducators.
4Designandmethods
4.1Datacollection
Thedatawasgatheredfromasampleof37sixthtoeighth-gradestudentsattendingsecondaryschoolsinSouthernGermany.Toensureadiversesample,theacademicperformanceofthestudentswasestimatedbysumminguptheirschoolgradesinmathematics,
German,andscience.Teachersinvitedstudentswith
good,average,andpooracademicperformancetopar-ticipateinthestudy.Allparticipatingstudentsvolun-teeredandhadparentalconsent.
Datawascollectedthroughcompletingexperimenta-tionprotocolswithsectionsfor‘Hypothesis’,‘Mate-rial’,‘Sketchoftheexperimentalsetup’,‘Descriptionoftheimplementation’,‘Observation’,and‘Result’.
Theyweregiventwotasks(Figure
1
)wheretheyhadtoplan,execute,andevaluateexperiments.Thefirsttaskinvolvedayeastexperiment,wherestudentshadtodeterminetheconditionsnecessaryforyeasttopro-ducecarbondioxide(task:“Findoutwhatyeastneedstoproducecarbondioxide”).Thesecondtaskrequiredthemtoexplorethefactorscausingpineconescalestoclose(task:“Findoutwhattriggersconescalestoclose”).Variousmaterialswereprovidedforeachtask,andstudentswerefreetochoosewhichmateri-alstouse.Bothtaskswerecompletedeitheronthesamedayorontwoconsecutivedays,with60minutesallottedforeachexperiment.Thestudentsworkedin-dependently,supervisedbyatraineduniversitystudentasassistant.Theuniversitystudentassistantswerere-sponsibleforensuringtaskcomprehension,explainingtheexperimentationprotocolandremindingstudentstocontinuedocumentingtheirwork.Theydidnotassistinconductingtheexperiments.
4.2Sample
Thefinaldatasetitselfconsistsof65structuredstudent
protocolsinGermanlanguagefromlaboratorycondi-
tions,focusingonexperimentsrelatedtoconesandyeast.25protocolswereratedbyhumanratersandthenexclusivelyusedastrainingdataforadaptingandrevisingtheAIsystem.Theremaining40protocolswereexclusivelyusedforcalculatinginter-rateragree-mentbetweenhumansandtheAI(all40protocols)andbetweenthreehumanraters(15protocolsasasubsetofthe40protocols).The40protocolsforcalculatingtheinter-rateragreementbetweenhumansandtheAIsystemaswellasthe15protocolsforcalculatingtheinter-rateragreementbetweenthreehumansrespec-tivelywereselectedwithrespecttoadiversesampleregardingthetopicoftheexperiment(conesoryeast),thestudents’gender(femaleormale),theirgrade(5th,6th,7thor8thgrade)aswellastheiracademicper-formance(poor,average,good).ThecompositionisdisplayedinTable
2
.
4.3DevelopmentoftheAIsystem
Inourproject,weuseapre-trainedLargeLanguageModeltoanalyzetheexperimentalprotocolsforcom-monstudenterrors.Duetotheinherentstrengthofpre-trainedLLMstofollowtextualdescriptions(Brown,2020),wecanleveragetheircapabilitiesandoperateonourlimitedtrainingdatasetofmerely25studentprotocols.Mitigatingtheneedforalarger,customdataset,wecanexploittheirknowledgeandmakeac-
6
Figure1Thetwotasksgiventothestudentstoanalyzetheirexperimentalprocedure.Theleftimageshowsthechangeofapineconethatthestudentswereaskedtoreproduce.Therightpictureshowsthematerialavailable(salt,yeast,waterandflour)tostimulateyeasttoproducecarbondioxide.
Dataset
Topic
Gender
Grade
Academic
performance
Total
trainingdataset
cones:10
yeast:15
female:13
male:12
6th:7
7th:11
8th:7
poor:7
average:9
good:9
25
Inter-humanrating(subsetofHumanvs.AIdataset)
cones:7
yeast:8
female:7
male:8
6th:5
7th:5
8th:5
poor:4
average:6
good:5
15
Humanvs.AI
cones:20
yeast:20
female:20
male:20
6th:15
7th:13
8th:12
poor:13
average:13
good:14
40
Table2Compositionofthesamplesforcalculatingtheinter-rateragreement
curatepredictionsandassessments.
Avalidandpublishedratingscheme,encompass-ingcommonstudenterrorsbutoriginallyfocusingon
videotapedanalysis(Baur,2021),servedasthefoun-dationtobuildanAIsystemfordetectingtheseerrors.ThisAIsystemisbasedonmodelsoftheGPT-3.5series(Ouyangetal.,2022)aswellastheGPT-4se-ries(OpenAI,2023),specificallyusingthe"*-0613"snapshotscorrespondingtotheversionsoftheGPTmodelsfromJune2023.
Weuseddifferentpromptingtechniques.A‘prompt’istypicallyashortstringoftextthatincludesinstructionsforthetask(zero-shotlearning)orafewsamplesofthetask(fewshotslearning)(Liuetal.,2022;Mayeretal.,2023).TocustomizetheLLMsforourusecase,weusedChain-of-Thoughtprompting(Weietal.,2022)androleprompting(definingGPTsroleas“Youareascienceteacherlookingatstudent’sprotocolsofex-periments”orsimilar)amongothers.
Besidesthemyriadchallengesassociatedwithdeploy-ingAIsystemsbasedonLLMsforgeneralfeedback,aparticularlyprominentissueisaccuratelyidentifying
logicalerrorsincomplex,incompleteorevencontra-dictorydatalikestudents’experimentationprotocols.Forthistask,wegenerallyfollowedatwo-prongedap-proach.Firstly,weidentifiedcriticalelementsoftheexperiment,suchasthedependentandindependentvariablesbydissectingandunderstandingthestudents’hypothesis.Thisinitialstepformsthebasisforunder-standingthestructureanddesignofthewholeexper-iment.Secondly,weperformedasystematicandal-gorithmicamalgamationoftheseidentifiedelements.Thisinvolvesexamininge.g.whetherthenumberofindependentvariablesalignswiththenumberoftesttrialsconducted.Anydiscrepancymaypointtowardserrorsmadeintheexperimentalprocess.Therefore,thiscomplexprocessoferroridentificationrequiresaninterplayofmethods,basedonbothLLMsandpurelyalgorithmicprocedures.TheyellowboxesinFigure
2
representtheLLM-basedmethodsforidentifyingthestudenterrors.Thearrowsindicatetheflowofinfor-
mation.Someerrorscanbedirectlyidentifiedthroughpromptingandtextfromthestudents’protocol,whileothersundergopreprocessingbeforeactuallycheck-
7
Figure2SimplifiedflowchartofourdevelopedLLM-basedAIsystem
ingiftheerrorispresent.Inthispreprocessingstage(grayboxes),keyfeaturesoftheexperimentareex-tractedfromtheprotocolusingbothLLMs(prompt-ing)andalgorithmictechniques.Theidentificationoftheseerrorsisthenperformedbasedonthisextractedinformation.
4.4Methodsofdataanalysis
Theaimofthedataanalysisistwofold.First,ourobjectiveistocomparehowdifferenthumanraters,guidedbythesharedratingscheme,ratethestudentprotocols.Second,itaimstocomparetheseresultswiththeoutcomesproducedwhentheAIsystemisappliedonstudentprotocols.Therefore,firstwecalcu-latedinter-rateragreementamongthreehumanraters.Weemployedthreehumanraterstorigorouslyassesstheinstruments’reliability,goingbeyondthetypicalinter-rateragreementderivedfromtworaters.Theagreementbetweenthreehumanratersprovidesusanessentialbenchmark,demonstratingtheconsistency,validityandreplicabilityofourratingschemetode-tectstudenterrorsacrossdifferentindividuals.Next,wecalculatedtheinter-rateragreementbetweenhu-
manratersandtheAIsystem.TheagreementbetweenhumanratersandtheAIsystemallowsustoascertainthattheAIsystemiscapableofreplicatingthesameprocessaccuratelyanddemonstratingitsefficacy.Asaby-producttheinter-rateragreementbetweenhumans
andtheAIsystemisaddinganotherlayerofvaliditytotheusedratingscheme.
FortheevaluationoftheeffectivenessoftheAIsystem,differentmethodsestablishedinthefieldofcomputerscienceaswellasmethodscommoninthefieldofsocialsciencesareapplied.TheratingsarecomparedbetweenhumanratersandtheAI-generatedanalysesbymetricscommoninthefieldofAI(Accuracy)andinthefieldofsocialsciences(CohensKappa,Co-hen,1960;FleissKappa,Fleiss,1971andGwet’sAC1,Gwet,2014).Forclassificationtasksincom-puterscience,accuracyisacommonlyusedperfor-mancemetrictoevaluatetheeffectivenessofamodel.Accuracymeasurestheproportionofcorrectlyclas-sifiedinstancesoutofthetotalnumberofinstances(Goodfellowetal.,2016).Tocalculateinter-raterreliabilityamongtworatersweuseCohensKappa,forthreeraters,weuseFleissKappa(Fleiss,1971).Weenhanceouroverviewofinter-rateragreementbyincorporatingGwet’sAC1.Gwet’sAC1metricpro-videsamorestableinter-raterreliabilitycoefficientthanCohensKappaandislessaffectedbyprevalenceandmarginalprobabilitythanCohensKappa(Wong-pakaranetal.,2013).AswereportonmanyerrorswhichareverycommonorveryrareweuseGwet’sAC1forourinter-raterreliabilityanalysistogainmoreinterpretableresults.
ForthemetricsCohensKappa,Fleiss’Kappaand
8
Gwet’sAC1thescaleproposedbyLandis&Koch(1977)isapplicableandusedthroughoutthispaper:
.00-.20:slightagreement;.21-.40:fairagreement;.41-.60:moderateagreement;.61-.80:substantialagreementand.81-1:almostperfectagreement.
5Results
5.1Inter-rateragreementbetweenhuman
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 金融学概论课程标准大纲
- 新部编人教版五年级下册语文课内外阅读理解专项练习题及答案
- 小学英语教师业务素质考试试题及答案
- 夏季夜市饮品成分安全辨别知识
- 2026年秋季开学第一课:中学生消防安全技能教育
- 2026 年夏季高中水域安全隐患排查与学生防溺教育
- 2026 年中秋食品安全科普知识学习课
- 2026 年九月户外务工人员秋季安全宣讲
- 消费者心理学试题及答案
- 2025年海底光缆AI音乐诠释通信基础设施
- OEE培训教学课件
- 零内耗培训课件
- 广西医疗机构病历书写规范与治理规定(第三版)
- 慢性阻塞性肺疾病护理策略
- ECMO联合连续性肾脏替代治疗(CRRT)方案
- 2025重庆璧山区西算大数据有限公司招聘工作人员5人笔试历年备考题库附带答案详解2套试卷
- 冬季建设工程赶工技术方案
- 初中物理新课程标准解读
- 基层网格员安全监管培训课件
- 卡西欧手表OCW-T400(5054)说明书
- 蔬菜大棚现场管理制度
评论
0/150
提交评论