版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
ClusteringClusteringOverviewPartitioningMethodsK-MeansSequentialLeaderModelBasedMethodsDensityBasedMethodsHierarchicalMethods2OverviewPartitioningMethods2Whatisclusteranalysis?FindinggroupsofobjectsObjectssimilartoeachotherareinthesamegroup.Objectsaredifferentfromthoseinothergroups.UnsupervisedLearningNolabelsDatadriven3Whatisclusteranalysis?FindiClustersInter-ClusterIntra-Cluster4ClustersInter-ClusterIntra-CluClusters5Clusters5ApplicationsofClusteringMarketingFindinggroupsofcustomerswithsimilarbehaviours.BiologyFindinggroupsofanimalsorplantswithsimilarfeatures.BioinformaticsClusteringmicroarraydata,genesandsequences.EarthquakeStudiesClusteringobservedearthquakeepicenterstoidentifydangerouszones.WWWClusteringweblogdatatodiscovergroupsofsimilaraccesspatterns.SocialNetworksDiscoveringgroupsofindividualswithclosefriendshipsinternally.6ApplicationsofClusteringMarkEarthquakes7Earthquakes7ImageSegmentation8ImageSegmentation8TheBigPicture9TheBigPicture9RequirementsScalabilityAbilitytodealwithdifferenttypesofattributesAbilitytodiscoverclusterswitharbitraryshapeMinimumrequirementsfordomainknowledgeAbilitytodealwithnoiseandoutliersInsensitivitytoorderofinputrecordsIncorporationofuser-definedconstraintsInterpretabilityandusability10RequirementsScalability10PracticalConsiderationsScalingmatters!11PracticalConsiderationsScalinNormalizationorNot?12NormalizationorNot?121313EvaluationVS.14EvaluationVS.14Evaluation15Evaluation15SilhouetteAmethodofinterpretationandvalidationofclustersofdata.Asuccinctgraphicalrepresentationofhowwelleachdatapointlieswithinitsclustercomparedtootherclusters.a(i):averagedissimilarityofiwithallotherpointsinthesameclusterb(i):thelowestaveragedissimilarityofitootherclusters16SilhouetteAmethodofinterpreSilhouette17Silhouette17K-Means18K-Means18K-Means19K-Means19K-Means20K-Means20K-MeansDeterminethevalueofK.ChooseKclustercentresrandomly.Eachdatapointisassignedtoitsclosestcentroid.Usethemeanofeachclustertoupdateeachcentroid.Repeatuntilnomorenewassignment.ReturntheKcentroids.ReferenceJ.MacQueen(1967):"SomeMethodsforClassificationandAnalysisofMultivariateObservations",Proceedingsofthe5thBerkeleySymposiumonMathematicalStatisticsandProbability,vol.1,pp.281-297.21K-MeansDeterminethevalueofCommentsonK-MeansProsSimpleandworkswellforregulardisjointclusters.Convergesrelativelyfast.RelativelyefficientandscalableO(t·k·n)t:iteration;k:numberofcentroids;n:numberofdatapointsConsNeedtospecifythevalueofKinadvance.Difficultanddomainknowledgemayhelp.Mayconvergetolocaloptima.Inpractice,trydifferentinitialcentroids.Maybesensitivetonoisydataandoutliers.Meanofdatapoints…NotsuitableforclustersofNon-convexshapes22CommentsonK-MeansPros22TheInfluenceofInitialCentroids23TheInfluenceofInitialCentrTheInfluenceofInitialCentroids24TheInfluenceofInitialCentrSequentialLeaderClusteringAveryefficientclusteringalgorithm.NoiterationAsinglepassofthedataNoneedtospecifyKinadvance.Chooseaclusterthresholdvalue.Foreverynewdatapoint:Computethedistancebetweenthenewdatapointandeverycluster'scentre.Iftheminimumdistanceissmallerthanthechosenthreshold,assignthenewdatapointtothecorrespondingclusterandre-computeclustercentre.Otherwise,createanewclusterwiththenewdatapointasitscentre.Clusteringresultsmaybeinfluencedbythesequenceofdatapoints.25SequentialLeaderClusteringA2626GaussianMixture27GaussianMixture27ClusteringbyMixtureModels28ClusteringbyMixtureModels28K-MeansRevisited
modelparameterslatentparameters29K-MeansRevisited
modelparamExpectationMaximization30ExpectationMaximization30
31
31EM:GaussianMixture32EM:GaussianMixture323333DensityBasedMethodsGenerateclustersofarbitraryshapes.Robustagainstnoise.NoKvaluerequiredinadvance.Somewhatsimilartohumanvision.34DensityBasedMethodsGenerateDBSCANDensity-BasedSpatialClusteringofApplicationswithNoiseDensity:numberofpointswithinaspecifiedradiusCorePoint:pointswithhighdensityBorderPoint:pointswithlowdensitybutintheneighbourhoodofacorepointNoisePoint:neitheracorepointnoraborderpointCorePointNoisePointBorderPoint35DBSCANDensity-BasedSpatialClDBSCANpqdirectlydensityreachablepqdensityreachableoqpdensityconnected36DBSCANpqdirectlydensityreachDBSCANAclusterisdefinedasthemaximalsetofdensityconnectedpoints.StartfromarandomlyselectedunseenpointP.IfPisacorepoint,buildaclusterbygraduallyaddingallpointsthataredensityreachabletothecurrentpointset.Noisepointsarediscarded(unlabelled).37DBSCANAclusterisdefinedasHierarchicalClusteringProduceasetofnestedtree-likeclusters.Canbevisualizedasadendrogram.Clusteringisobtainedbycuttingatdesiredlevel.NoneedtospecifyKinadvance.Maycorrespondtomeaningfultaxonomies.38HierarchicalClusteringProduceAgglomerativeMethodsBottom-upMethodAssigneachdatapointtoacluster.Calculatetheproximitymatrix.Mergethepairofclosestclusters.Repeatuntilonlyasingleclusterremains.Howtocalculatethedistancebetweenclusters?SingleLinkMinimumdistancebetweenpointsCompleteLinkMaximumdistancebetweenpoints39AgglomerativeMethodsBottom-upExample
BAFIMINARMTOBA0662877255412996FI6620295468268400MI8772950754564138NA2554687540219869RM4122685642190669TO9964001388696690SingleLink40Example
BAFIMINARMTOBA06628772Example
BAFIMI/TONARMBA0662877255412FI6620295468268MI/TO8772950754564NA2554687540219RM412268564219
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 八年级物理下学期核心知识与素养进阶清单
- 八年级生物上册第五单元第一章动物的主要类群复习课教学设计
- 北师大版小学一年级数学下册第五单元《100以内数加与减(一)》整体教学设计
- 八年级地理(粤人版)上册 第四单元 中国的主要产业
- 初中八年级地理气候第1课时·气温降水与季风核心知识清单
- 初中八年级道德与法治《宪法监督:筑牢法治国家的基石》导学案
- 初中八年级科学(浙教版)核心知识清单:物质在水中的分散状况深度解析
- 初三地理中考一轮复习:专题四 居民、文化与发展合作深度整合教案
- 《医学免疫学与微生物学》整合教案:抗结核感染的免疫屏障-以临床医学专业本科二年级为例
- 八年级物理(上册)核心知识清单:光的反射定律与综合应用
- 2026年上海市初三语文二模试题汇编《综合运用》含答案
- (2026版)《煤矿重大事故隐患判定标准》培训课件
- 2026年北京市西城区初三下学期二模英语试卷和答案
- 社区特殊人群服务管理操作规范
- 体检中心感染工作制度
- T-SZRCA 011-2025 人形机器人专用线缆技术规范
- 汉字造型美学研究报告
- 2026年湖南高考历史真题试卷+解析及答案
- 2026年安徽高考地理真题解析含答案
- 动力卷绕机培训课件
- 2025年心电图高频考题题库及答案(共650题)
评论
0/150
提交评论