Reliability of Supervised Machine Learning Using 使用监督机器学习的可靠性

上传人：媚*** IP属地：境外上传时间：2024-04-06 格式：DOC 页数：28 大小：1.44MB 积分：20 举报 版权申诉

Reliability of Supervised Machine Learning Using 使用监督机器学习的可靠性_第2页

Reliability of Supervised Machine Learning Using 使用监督机器学习的可靠性_第3页

Reliability of Supervised Machine Learning Using 使用监督机器学习的可靠性_第4页

Reliability of Supervised Machine Learning Using 使用监督机器学习的可靠性_第5页

已阅读5页，还剩23页未读，继续免费阅读

版权说明：本文档由用户提供并上传，收益归属内容提供方，若内容存在侵权，请进行举报或认领

文档简介

OriginalPaper

DebbieRankin1PhD,Correspondingauthor,

d.rankin1@ulster.ac.uk

,+442871675841

MichaelaBlack1PhD,

mm.black@ulster.ac.uk

RaymondBond2PhD,

rb.bond@ulster.ac.uk

JonathanWallace2MSc,

jg.wallace@ulster.ac.uk

MauriceMulvenna2PhD,

md.mulvenna@ulster.ac.uk

GorkaEpelde3,4PhD,

gepelde@

1SchoolofComputing,EngineeringandIntelligentSystems,UlsterUniversity,Derry~Londonderry,NorthernIreland,UnitedKingdom

2SchoolofComputing,UlsterUniversity,Jordanstown,NorthernIreland,UnitedKingdom

3VicomtechFoundation,BasqueResearchandTechnologyAlliance(BRTA),Donostia-SanSebastián,Spain

4BiodonostiaHealthResearchInstitute,eHealthGroup,Donostia-SanSebastián,Spain

ReliabilityofSupervisedMachineLearningUsingSyntheticDatainHealthcare:AModeltoPreservePrivacyforDataSharing

Abstract

Background:

Theexploitationofsyntheticdatainhealthcareisatanearlystage.Syntheticdatagenerationcouldunlockthevastpotentialwithinhealthcaredatasetsthataretoosensitiveforreleaseduetoprivacyconcerns.Severalsyntheticdatageneratorshavebeendevelopedtodate,howeverstudiesevaluatingtheirefficacyandgeneralisabilityarescarce.

Objective:

Thisworksetsouttounderstandthedifferenceinperformanceofsupervisedmachinelearningmodelstrainedonsyntheticdatacomparedwiththosetrainedonrealdata.

Methods:

Atotalof19openhealthcaredatasetscontainingbothcategoricalandnumericaldatahavebeenselectedforexperimentalwork.SyntheticdataisgeneratedusingthreepopularsyntheticdatageneratorsthatapplyClassificationandRegressionTrees,parametricandBayesiannetworkapproaches.Realandsyntheticdataareused(separately)totrainfivesupervisedmachinelearningmodels:stochasticgradientdescent,decisiontree,k-nearestneighbors,randomforestandsupportvectormachine.Modelsaretestedonlyonrealdatatodeterminewhetheramodeldevelopedbytrainingonsyntheticdatacanbeputintousebyhealthcaredepartmentsandusedtoaccuratelyclassifynew,realexamples.Evaluationmetricsarecomputedanddifferentialsinthesescoresarecompared.Theimpactofstatisticaldisclosurecontrolonmodelperformanceisalsoassessed.

Results:

TheaccuracyofMLmodelstrainedonsyntheticdataislowerthanmodelstrainedonrealdatain92%ofcases.Tree-basedmodelstrainedonsyntheticdatahavedeviationsinaccuracyfrommodelstrainedonrealdataof17.7-19.3%,whilstothermodelshavelowerdeviationsof5.8-7.2%.Thewinningclassifierwhentrainedandtestedonrealdataversusmodelstrainedonsyntheticdataandtestedonrealdataisthesamein26.3%ofcasesforCARTandparametricsyntheticdata,andin21.1%ofcasesforBayesiannetworkgeneratedsyntheticdata.Tree-basedmodelsperformbestwithrealdataandarethewinningclassifierin94.7%ofcases.Thisisnotthecaseformodelstrainedonsyntheticdata.Whentree-basedmodelsarenotconsidered,thewinningclassifierforrealandsyntheticdataismatchedin73.7%,52.6%and68.4%ofcasesforCART,parametricandBayesiannetworksyntheticdata,respectively.Statisticaldisclosurecontrolmethodsdidnothaveanotableimpactondatautility.

Conclusions:

Theresultsofthisstudyarepromisingwithsmalldecreasesinaccuracyobservedinmodelstrainedwithsyntheticdatacomparedtomodelstrainedwithrealdata,wherebotharetestedonrealdata.Suchdeviationsareexpectedandmanageable.Tree-basedclassifiershavesomesensitivitytosyntheticdataandtheunderlyingcauserequiresfurtherinvestigation.Thisstudyhighlightsthepotentialofsyntheticdataandtheneedforfurtherevaluationitsrobustness.Syntheticdatamustensureindividualprivacyanddatautilityispreservedinordertoinstilconfidenceinhealthcaredepartmentswhenutilisingsuchdatatoinformpolicydecision-making.

Keywords:SyntheticData;SupervisedMachineLearning;DataUtility;Healthcare;DecisionSupport;StatisticalDisclosureControl

Introduction

Background

NationalHealthcareDepartmentsholdvastvolumesofdataonpatientsandthepopulationthatisnotbeingusedtoitsfullpotentialduetovalidprivacyconcerns.Machinelearning(ML)hasthepotentialtovastlyimprovedecisionsandoutcomesinhealthcareandyettheseimprovementshavenotyetbeenfullyrealised.Thereasonmaybeinpartrelatedtoanissuethatfacesmanydatascientistsandresearchersinthearea:thelimitedavailabilityoforaccesstodata,orthereadinessforhealthcareinstitutionstosharedata.Privacyconcernsoverpersonaldata,andinparticularhealthcaredata,meansthatalthoughthedataexists,itisdeemedtoosensitiveforpublicrelease[1],eveninthecaseofseriousresearch.

Onewaytoovercometheissueofdataavailabilityistousefullysyntheticdataasanalternativetorealdata.Theexploitationofsyntheticdatainhealthcareisatanearlystageandisgainingincreasingattention.Syntheticdataisdatathatissimulatedfromrealdatabyusingtheunderlyingstatisticalpropertiesoftherealdatatoproducesyntheticdatasetsthatexhibitthesesamestatisticalproperties.Syntheticdatacanrepresentthepopulationintheoriginaldatawhilstavoidinganydivulgenceofreal,potentiallypersonal,confidentialandsensitivedata.Inthecaseofhealth-relateddata,thiswouldensurethatactualpatientrecordsarenotdisclosedthusavoidinggovernanceandconfidentialityissues.Therearethreetypesofsyntheticdata:fullysynthetic,partiallysynthetic,andhybridsynthetic.Thisworkconsidersfullysyntheticdatawhichdoesnotcontainoriginaldata.

Syntheticdatacanbeusedintwoways:toaugmentanexistingdatasetthusincreasingitssize,fortimeswhenadatasetisunbalancedduetothelimitedoccurrenceofaneventorwhenmoreexamplesarerequired[2,3];andtogenerateafullysyntheticdatasetthatisrepresentativeoftheoriginaldataset,fortimeswhendataisnotavailableduetoitssensitivenature[4].Thelatterisconsideredinthisworkasakeyrequirementforhealthcaredatasharing.

Traditionally,dataperturbationtechniquessuchasdataswapping,datamasking,cellsuppressionandaddingnoise,havebeenappliedtorealdatatomodifyandthusprotectthedatafromdisclosurepriortoreleasingit.However,suchmethodsdonoteliminatedisclosureriskandcanimpacttheutilityofthedata,particularlyifmultivariaterelationshipsarenotconsidered[5].SyntheticdatawasfirstproposedbyRubin[6]andLittle[7].Raghunathan,ReiterandRubin[8]implementedandextendeduponthis,pioneeringthemultipleimputationapproachtosyntheticdatageneration,exemplifiedinarangeofstudies[9-14].Reiter[15]thenintroducedanalternativemethodofsynthesisingdatathroughanon-parametrictree-basedtechniquethatutilisesClassificationandRegressionTrees(CART).AmorerecenttechniqueproposesaBayesiannetworkapproachforsyntheticdatageneration[16].Syntheticdataisconsideredasecureapproachforenablingpublicreleaseofsensitivedataasitgoesbeyondtraditionalde-identificationmethodsbygeneratingafakedatasetthatdoesnotcontainanyoftheoriginal,identifiableinformationfromwhichitwasgenerated,whilstretainingthevalidstatisticalpropertiesoftherealdata.Therefore,theriskofdisclosureofarealpersonorreverseengineeringisconsideredtobeunlikely[17].

Whilstanumberofsyntheticdatageneratorshavebeendeveloped,empiricalevidenceoftheirefficacyhasnotbeenfullyexplored.Thisworkextendsapreliminarystudy[18]andinvestigateswhetherfullysyntheticdatacanpreservethehiddencomplexpatternsthatsupervisedMLcanuncoverfromrealdata,andthereforewhetheritcanbeusedasavalidalternativetorealdatawhendevelopingeHealthapplicationsandhealthcarepolicymakingsolutions.Thiswillbeachievedbyexperimentingwitharangeofopenhealthcaredatasets.Syntheticdatawillbegeneratedusingthreewellknownsyntheticdatagenerationtechniques.SupervisedMLalgorithmswillbeusedtovalidatetheperformanceofthesyntheticdatasets.Statisticaldisclosurecontrol(SDC)methodsthatcanfurtherdecreasethedisclosureriskassociatedwithsyntheticdatawillalsobeconsidered.

Overview

Toinformtheviabilityoftheuseofsyntheticdataasavalidandreliablealternativetorealdatainthehealthcaredomainwewillanswerthefollowingresearchquestions:

WhatisthedifferentialinperformancewhenusingsyntheticdataversusrealdatafortrainingandtestingsupervisedMLmodels?

WhatisthevarianceofabsolutedifferenceofaccuraciesbetweenMLmodelstrainingonrealandsyntheticdatasets?

HowoftendoesthewinningMLtechniquechangewhentrainingusingrealdatatotrainingusingsyntheticdata?

Whatistheimpactofstatisticaldisclosurecontrol(i.e.privacyprotection)measuresontheutilityofsyntheticdata(i.e.similaritytorealdata)?

Toanswerthesequestions,19openhealthcaredatasetscontainingbothcategoricalandnumericaldatahavebeenselectedforexperimentation[19].Syntheticdatasetsaregeneratedforeachofthese19datasetsusingthreepopularsyntheticdatageneratorsthatapplyCART[15,17],parametric[8,17]andBayesiannetwork[16]approaches,respectively,toenablearobustcomparisonofthethreesyntheticdatagenerationtechniquesacrossabroadrangeofdata.

Initiallyweanalysewhetherthemultivariaterelationshipsthatexistintherealdataarepreservedinthesyntheticversionsofthedata,fordatageneratedusingeachofthethreesyntheticdatagenerationtechniques,bycomputingpairwisemutualinformationscoresforeachvariablepaircombinationineachdataset[16].Itisimportantthatsuchrelationshipsareretainedwhendataissynthesised.

ToevaluatetheutilityofsyntheticdataforMachineLearning,wetheninvestigatetheperformanceofsupervisedMLmodelstrainedonsyntheticdataandtestedonrealdata,comparedwithmodelstrainedonrealdataandalsotestedontherealdata.Thisallowsustodetermineifamodeldevelopedusingsyntheticdatacanclassifyrealdataexamplesasaccuratelyandreliablyasamodeldevelopedusingrealdata.Weconsiderfivedifferentsupervisedmachinelearningmodelstocompareperformanceanddetermineiftherearedifferencesinrobustnessacrosseachofthesemodels.Standardevaluationmetricsarecomputedformodelstrainedonrealandsyntheticdata,foreachMLmodel,andforeachdataset[20].Thedifferencesinaccuracyformodelstrainedonsyntheticdataversusmodelstrainedonrealdataarecomputedtoanalysetheextenttowhichsyntheticdatacausesadegradationinmodelperformance,ifany.

ItispertinentthattheoptimalMLmodelbuiltusingsyntheticdatamatchestheoptimalMLmodelthatwouldbeselectedifrealdatawereusedinthemodeltrainingprocess.Thiswouldprovidestakeholdersinhealthcarewithconfidenceintheuseofsyntheticdataformodeldevelopment.Thus,weconsiderhowoftenthebestMLclassifierbuiltusingsyntheticdatamatchesthebestMLmodelbuiltusingrealdata.

Finally,theimpactofanumberofstatisticaldisclosurecontrolmethodsonmodelperformanceisassessed.Statisticaldisclosurecontrolmethodsseektofurtherenhancedataprivacy;however,thiscanleadtoalossinusefulnessofthedata[21]andweconsidertheextenttowhichperformancedegradationoccursasaresultofSDC.

Thislarge-scaleassessmentofthereliabilityofsyntheticdatawhenusedforsupervisedML,utilising19healthcaredatasetsand3syntheticdatagenerationtechniques,providesanimportantcontributioninrelationtothetrustandconfidencethatstakeholdersinhealthcarecanhaveinsyntheticdata.Wealsoproposeapipelinetoillustratehowsyntheticdatacanpotentiallyfitwithinthehealthcareprovidercontext.Thisworkdemonstratesthepromisingperformanceofsyntheticdatawhilsthighlightingitslimitationsandfutureworkdirectionstoovercomethem.

SyntheticData:PresentandFutureUse

ThevalidityanddisclosureriskassociatedwithsyntheticdatahasbeenunderinvestigationbytheU.S.CensusBureausince2003forthepurposeofcreatingpublicusedatafromacombinationofsensitivedatafromtheCensusBureau’sSurveyofIncomeandProgramParticipation(SIPP),theInternalRevenueService’s(IRS)individuallifetimeearningsdata,andtheSocialSecurityAdministration’s(SSA)individualbenefitdata[22,23].Thegoalwastoenablethereleaseofsynthesisedperson-levelrecordscontainingpersonalandfinancialcharacteristicsfromconfidentialdatasets,whilstpreservingprivacy.Successfulresultshaveledtothereleaseofpublicusesyntheticdatafiles.ResearcherscanhavetheirworkvalidatedagainsttheGoldStandard(real)databytheCensusBureau,thusenablingthemtodeterminetheimpactofsyntheticdataontheirexploratoryanalysesandmodeldevelopmentandhaveconfidenceintheirresults,whilstalsoallowingtheCensusBureautocontinuouslyimprovetheirsynthesistechniques.Thepublicreleaseofthisdatahasprovidedsignificantbenefittotheresearchcommunityandgeneralpopulation,enablingmoreextensiveeconomicpolicyresearchtobeperformedbygroupswhocouldnotpreviouslyaccessusefuldata[24-29].ThisworkledtothereleaseoffurthersyntheticdatasetsbytheCensusBureau.TheSyntheticLongitudinalBusinessDatabase(SynLBD)comprisesdatafromanannualeconomiccensusofestablishmentsintheU.S.[30].Thisdatasetprovidesbroadaccesstorichdatathatsupportstheresearchandpolicy-makingcommunitiesinbusinessandemploymentrelatedtopics.OnTheMapisatoolutilisingsyntheticdatatoprovideworkforcerelatedmaps,demographicprofilesandreportsofU.S.citizens,aswellasdisastereventinformationandtheimpactofsucheventsonworkersandemployers[31].Similarly,syntheticdatahasalsobeenunderinvestigationintheUKasameanstoprovidepublicaccesstorichdatafromUKLongitudinalStudies[32-34]thatcontainhighlysensitivedatalinkingnationalcensusdatatoadministrativedataforindividualsandtheirfamilies.

Thesedatasetsenableresearcherstoexploredataanddevelopandtestcodeandmodelsoutsidethesecureenvironmentwhererealdataresideswithnorestrictions,whilstthedataownersprovideavalidationmechanismwhereresults,codeandmodelscanbevalidatedonbehalfofresearchersontherealdatawithinthesecureenvironmentandfeedbackprovided.Thisprocessincreasesresearchproductivitywhilstensuringthedevelopmentofrobustandvalidmodels[35].

Whilstsyntheticdatahasbeenusedtoaccelerateanddemocratisebusinessandeconomicpolicyresearch[22-35],itisnotcurrentlyinuseforhealthcareresearch,anareathatcouldbenefitenormously.Withadvancementsintechnology,particularlyMLandartificialintelligence(AI),thepotentialtodevelopdiagnostictoolsforcliniciansanddatadrivendecision-makingplatformsforhealthpolicy-makersisever-increasing[36,37].Suchtoolsrequireaccesstohealthcaredata,forexample,totrainAIalgorithmsandproducemodelsthatcanidentifyhealthconditionsandhealth-relatedpatternsacrossthepopulation.Currentlyitcantakealengthyperiodoftimeforresearcherstogainaccesstohealthcaredata,arichandunder-utilisedresource,duetoprivacyconcerns[38-42].Forexample,inthecaseofthe40monthMIDASProject[36,43]developingadata-drivendecisionmakingtoolforhealthcarepolicymakers,ittookmorethan20monthstoobtainaccesstotherequireddataduetolegalandethicalconstraints.Inaddition,anumberofimportantdatavariablescouldnotmadeavailablewhichrestrictedtheutilityoftheplatformunderdevelopment.Withthehelpofsyntheticdata,suchdata,withmoreorallvariablesincluded,couldhavebeenmadeavailableinamatterofweeksthusprovidingmoretimefordevelopmentandevaluationoftheplatform.Theplatformcouldthenhavebeeninstalledinhealthcaresitesmorequicklyandconnectedtorealdataforvalidationandcomparisonofperformanceforsyntheticversusrealdata,enablingperformancetweakstomitigatebiasintroducedbysyntheticdata,ifany.Syntheticdatacouldalsoenablecross-siteanalyticsacrossvarioushealthregions,thatwouldenablepolicymakerstoconnecttheirhealthspacesandpotentiallyprovidesignificantenhancementstocross-nationalhealthpolicy.

Theultimategoalofthisworkistofurtherassessthevalidityanddisclosureriskofsyntheticdataunderthestringentconditionsassociatedwithhealthcaredata,withtheviewtosuccessfullydevelopingapipelineforuseinhealthcarethatenablessyntheticdatasetstobereleasedpubliclytoresearcherswhowouldotherwisenotbeabletoaccessthedata,oraccessitinatimelyfashion,inordertoaccelerateresearchbyenablingthewiderresearchcommunitytousethedataforanalysisandmodeldevelopment.Theresultsofsuchanalysesandthemodelsandcodedevelopedcanthenbegiventohealthcaredepartmentsforvalidationontherealdata,andifeffectivecanbeputintousebycliniciansandhealthpolicy-makers.

SyntheticDataPipelineforHealthcare

TounderstandhowhealthcaredepartmentscanbenefitfromsyntheticdataweproposeapipelineshowninFigure1.Thisisaproposedsyntheticdatasharingpipelineprovidedasanillustrationofhowsyntheticdatacanpotentiallyworkwithinarealhealthcaresettingtoexpeditedataanalytics.Infuturework,weplantotestthispipelineinarealsetting.InthispipelinerealdataresideswithintheNationalHealthcareDepartmentinfrastructure.Thedatacannotbesharedexternallyduetoitssensitiveandprivatenature.HealthcaredepartmentsmayonlyhaveasmallnumberofdatasciencestaffwiththeexpertisenecessarytoapplyMLtechniquestomanyoftheirdatasets,andsotheycannotmaximisetheuseoftheirdatanordiscovertheiruseduetolackofresources.ByapplyingasyntheticdatagenerationtechniquetotherealdataalongwithSDCmeasures,asyntheticdatasetcanbeproducedandmadeavailabletotheexternalresearchcommunityinplaceoftherealdata.Externalresearchers,inlargenumbersandwithwiderangingexpertise,canpotentiallydevelopoptimalMLmodelstrainedonthesyntheticdataandsharetheperformanceoftheMLmodel,themodelitselfandthemodelspecificationwiththeNationalHealthcareDepartment.ThehealthcaredepartmentcanthentesttheMLmodelonrealdataorin-housetechnicalstaffcanrebuildthemodelaccordingtothespecificationprovidedbyresearcherswherethespecificationcanincludetheprogramcodewrittenbyresearchers,detailsoftheMLalgorithmtouse,e.g.decisiontree,supportvectormachineetc.,andtheoptimalhyperparametersettingsdeterminedduringdevelopment.Usingthesesettings,themodelcanthenberebuilt,thistimebytrainingontherealdatainsteadofsyntheticdata,whichin-housestaffhaveaccessto.

Figure1Proposedsyntheticdatasharingpipelinetoillustratehowsyntheticdatacouldbeimplementedtoexpeditehealthcaredataanalytics.

Methods

DatasetSelection

Forexperimentation,19openhealthcaredatasetshavebeenselectedfromtheUCIMachineLearningRepository[19].Missingvalueshavebeenremovedfromthedatasetseitherbyremovingfeatureswithahighnumberofmissingvaluesorremovingobservationswhereafeaturecontainsamissingvalue.TheexperimentaldatasetsandtheirpropertiesaresummarisedinTable1.Thesedatasetswereselectedtoenableananalysisofsyntheticdataperformancewhenappliedtodatasetsofdifferingvolumeanddatatypes(categoricalandnumerical).

Table1.Summaryofexperimentaldatasets.a

Dataset

No.ofAttributes

No.ofCategoricalAttributes

No.ofNumericalAttributes

No.Classes/Labels

No.ofObservations

BreastCancerWisconsin(Original)

683

BreastCancer

277

BreastCancerCoimbra

116

BreastTissue

106

ChronicKidneyDisease

209

Cardiotocography(3Class)

2126

Cardiotocography(10Class)

2126

Dermatology

358

DiabeticRetinopathy

1151

Echocardiogram

106

EEGEyeState

14980

HeartDisease

303

Lymphography

148

Post-OperativePatientData

PrimaryTumor

336

Stroke

29072

ThoracicSurgery

470

ThyroidDisease

5786

ThyroidDisease(New)

215

Total

283

144

139

105

58,655

aEachdatasethasbeenencodedwithaletter(column1)andwillbereferencedusingthisletterfortheremainderofthepaper.

GeneratingSyntheticData

Inthiswork,weanalyseandassesstheperformanceofthreepubliclyavailablesyntheticdatagenerationtechniquesthatarebasedonwell-known,seminalworkinthearea[6-10,15,16].Thesemethodsareaparametricdatasynthesistechnique,anon-parametrictree-basedsynthesistechniquethatutilisesCART[15],andasynthesistechniquethatutilisesBayesiannetworks[16].Whilstotherapproachesexist,somearedevelopedforspecificdatasetsandproblems,e.g.SimPopsimulatespopulationsurveydata[44],andSyntheasimulatespatientpopulationandelectronichealthrecorddata[45],whereasthesetechniquesareconsideredtobemoregeneral.TheRpackage,Synthpop,developedbyNowak,RaabandDibben[17],providesapubliclyavailableimplementationoftheparametricandCARTbasedsyntheticdatagenerators.TheDataSynthesizerpythonimplementation,developedbyPing,StoyanovichandHowe[16],providesapubliclyavailableimplementationoftheBayesiannetworkbasedsyntheticdatagenerator.Theseimplementationshavebeenutilisedinthisexperimentalwork.

AttributesaresynthesisedsequentiallyinboththeparametricandCARTmethods.Thesyntheticvaluesforthefirstattributearesynthesisedusingarandomsamplefromtheoriginalobserveddatasinceithasnopredictorsfrompreviouslysynthesisedattributesinthedataset.Whensynthesisingattributes,bothcategoricalandnumerical,withthenon-parametricmethod,theCARTmethodisapplied.CARTisappliedtoallvariablesthathavepredictors,i.e.attributespriortotheminthesequence,anddrawsfromtheconditionaldistributionsfittedtotheoriginaldatausingCARTmodels.Theparametricmethodsynthesisesattributebasedondatatype.Numericalattributesaresynthesisedusingnormallinearregression.Categoricalattributesaresynthesisedusingpolytomouslogisticregressionwheretheattributehasmorethantwolevels,whilstlogisticregressionisappliedtosynthesisebinarycategoricalvariables[17].TheBayesiannetworkmethodofsynthesisingdatalearnsadifferentiallyprivateBayesiannetworkthatcapturescorrelationstructurebetweenattributesintherealdataanddrawssamplesfromthismodeltoproducesyntheticdata[16].

SupervisedMachineLearningwithRealandSyntheticData

AkeymeasureofdatautilityofasyntheticdatasetforthepurposeofMListodeterminehowwellasupervisedMLmodeltrainedonsyntheticdata,performswhentaskedwithclassifyingrealdata.ThiswilldeterminewhethersupervisedMLmodelswillberobustenoughtoclassifyrealdataexamplesifonlysyntheticdataisprovidedforthetrainingofthesemodels.

ToevaluatewhethersyntheticdatasetscanbeusedasavalidalternativetorealdatasetsinML,foreachofthe19datasets(Table1),fivedifferentclassificationmodelsweretrained.Initiallythemodelsweretrainedandtestedontherealdatatoobtainaperformancebenchmark.Subsequently,aclassifierwastrainedoneachofthesyntheticdatasets,generatedusingparametric,CARTandBayesiannetworktechniques,andthentestedwiththerealdata.Modelsaretestedonrealdataonly,todeterminewhetheramodeldevelopedbytrainingonsyntheticdatacanbeputintousebyhealthcaredepartmentsandusedtoaccuratelyclassifynew,realexamples.

Therangeofmodelsappliedtoeachdatasetwere:stochasticgradientdescent(SDG)decisiontree(DT),k-nearestneighbors(KNN),randomforest(RF),andsupportvectormachine(SVM).Thisselectionofalgorithmswasappliedtodeterminehowwelleachperformedwhentrainedwiththerealdatacomparedwiththesyntheticdata,withbothtestedonrealdata.

TheclassifierswereimplementedusingPython’sScikit-Learn0.21.3machinelearninglibraryandareasfollows:

StochasticgradientdescentclassificationwasimplementedusingSGDClassifier,asimplelinearclassifier,withloss=“hinge”,random_state=0andallotherparameterssettotheirdefaults.

DecisiontreeclassificationwasimplementedusingDecisionTreeClassifier,anoptimisedversionofCART,withcriterion=“gini”,max_depth=10andrandom_state=0andallotherparameterssettotheirdefaults.

K-NearestNeighborsclassificationwasimplementedusingKNeighborsClassifierwithn_neighbors=10,weights=‘uniform’,leaf_size=30,p=2,metric=‘minkowski’,n_jobs=2andallotherparameterssettotheirdefaults.

RandomForestclassificationwasimplementedusingRandomForestClassifierwithcriterion=“gini”,max_depth=10,min_samples_split=2,n_estimators=10,random_state=1andallotherparameterssettotheirdefaults.

SupportVectorMachineclassificationwasimplementedusingSVCwithC=1.0,degree=3,kernel=‘rbf’,probability=True,random_state=Noneandallotherparameterssettotheirdefaults.

Fortrainingandtesting,Python’sScikit-Learn0.21.3ShuffleSplitrandompermutationcross-validatorwasusedwith10splittingiterationsandatrain/testsplitof75/25.Categoricalattributesweretransformedintoindicatorattributesusingone-hotencoding.

StatisticalDisclosureControl

Syntheticdataisconsiderednottocontainrealunitsandthereforetheriskofdisclosureofarealpersonisconsideredtobeunlikely[46].Whilstunlikely,thescenariowheresomeofthegeneratedsyntheticdataisverysimilartotherealdata,resultinginpotentialdisclosurerisk,mustbeconsideredandwhereadditionalprotectionscanbeappliedtosyntheticdataitisplausibletodoso.Additionalstatisticaldisclosurecontrol(SDC)measures,beyonddatasynthesis,canbeappliedasaprecautionarymeasuretoaddfurtherprotectionstosyntheticdatabyreducingtheriskofreproducingrealpersonrecordsandreplicatingoutlierdata,thus

人人文库> 全部分类> 行业资料 > 信息产业

温馨提示

1. 本站所有资源如无特殊说明，都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
2. 本站的文档不包含任何第三方提供的附件图纸等，如果需要附件，请联系上传者。文件的所有权益归上传用户所有。
3. 本站RAR压缩包中若带图纸，网页内容里面会有图纸预览，若没有图纸预览就没有图纸。
4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
5. 人人文库网仅提供信息存储空间，仅对用户上传内容的表现方式做保护处理，对用户上传分享的文档内容本身不做任何修改或编辑，并不能对任何下载内容负责。
6. 下载文件中如有侵权或不适当内容，请与我们联系，我们立即纠正。
7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

Reliability of Supervised Machine Learning Using 使用监督机器学习的可靠性

文档简介

温馨提示

最新文档

评论

Reliability of Supervised Machine Learning Using 使用监督机器学习的可靠性

文档简介

温馨提示

最新文档

评论

相关文档