最新跨媒体检索介绍主题讲座课件_第1页
最新跨媒体检索介绍主题讲座课件_第2页
最新跨媒体检索介绍主题讲座课件_第3页
最新跨媒体检索介绍主题讲座课件_第4页
最新跨媒体检索介绍主题讲座课件_第5页
已阅读5页,还剩33页未读 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

1、什么是跨媒体?从应用平台方面理解电视机电脑手机报纸Ipad以文字搜文字以图片搜图片以文字搜图片以文字搜视频什么是跨媒体?从检索研究方面理解什么是跨媒体? 2010年1月Nature发表的“2020 Vision”论文指出:文本、图像、语音、视频及其交互属性将紧密混合(mix)在一起,即“跨媒体”。2011年2月Science开灯“Dealing with Data”专辑:数据的组织和使用体现跨媒体计算。趋势:从“多媒体”研究向“跨媒体”发展!什么是跨媒体?跨媒体特性即多媒体数据之间以及用户互动与多媒体数据之间存在着内容跨越与语义关联。吴飞, 庄越挺. 互联网跨媒体分析与检索:理论与算法. 计算

2、机辅助设计与图形学学报,Vol.22, No.1, pp.1-9, 2010.跨媒体的主要研究范畴跨媒体检索:用户向计算机提交一种类型的多媒体对象作为查询例子,系统可以自动找到其它不同类型及语义上相似的多媒体对象。跨媒体推理:跨媒体推理是指从一种类型的多媒体数据,经过问题求解转向另外一种类型的多媒体数据。(OCR等)跨媒体存储:现有处理海量数据的检索技术主要是针对文本信息,如google和百度等搜索引擎。跨媒体存储研究高效压缩、索引和分片等方法,以及对用户行为的个性化索引等技术。惊涛骇浪?AudioVideoWebpageCorrelated multi-modal DataShared sp

3、aceHow to bridge both semantic-gap and heterogeneity gap?Japan Earthquake跨媒体分析的挑战From FeiWu跨媒体的内容鸿沟视觉特征空间听觉特征空间高层语义空间爆炸、海洋、天空、鸟。语义鸿沟内容鸿沟基于线性变换的子空间映射算法视觉特征空间听觉特征空间投影子空间Heterogeneous Metric Learning with Joint Graph Regularizationfor Cross-Media RetrievalXiaohua Zhai, Yuxin Peng and Jianguo XiaoInstit

4、ute of Computer Science & technology, Peking UniversityAAAI 2013Existing metric learning methods have previously been designed primarily for single-media data and cannot be directly applied to cross-media data.Make full use of the structure information of the whole heterogeneous spaces.MotivationHet

5、erogeneous Metric Learning Given two sets of heterogeneous pairwise constraintsS is the set of similarity constraints and D is the set of dissimilarity constraints . Each pairwise constraints (xi,yj) indicates if two heterogeneous media objects xi and yj are relevant or irrelevant inferred from the

6、category label.Joint Graph Regularized Heterogeneous MetricThey propose to learn multiple linear transformation matrices U and V , they can map the heterogeneous media data to a common output spaces.The distance measure is defined as:Joint Graph Regularized Heterogeneous MetricObjective functionThe

7、formulation of the general regularization framework for heterogeneous distance metric learning is defined as:f (U, V) is the loss function defined on the sets of similarity and dissimilarity constraints S and D g(U, V) and r(U, V) are regularizer defined on the target parameter matrices U, V. , are

8、the balancing parameters.Joint Graph Regularized Heterogeneous MetricLoss functionThe minimization of the loss function will result in minimizing (maximizing) the distances between the media objects with the similarity (dissimilarity) constraintsNormalize the elements of Z column by column to make s

9、ure that the sum of each column is zero - to balance the influence of the similarity constraints and dissimilarity constraints.Joint Graph Regularized Heterogeneous MetricScale regularization r(U,V) is used to control the scale of the parameters matrices and reduce overfitting.Joint Graph Regularize

10、d Heterogeneous MetricJoint graph regularizationDefining a joint undirected graph, G = (V, W) on the dataset. Each element wij of the similarity matrix W = wij(m+n)(m+n) means the similarity between the i-th media object and j-th media object. Using label information to construct the symmetric simil

11、arity matrix: whereJoint Graph Regularized Heterogeneous MetricJoint graph regularizationSetting wii = 0 for 1 i m+n to avoid self-reinforcement. And the normalized graph Laplacian L is defined as: Where I is an (m+n)(m+n) identity matrix and D is an (m+n)(m+n) diagonal matrix with . is symmetric an

12、d positive semidefinite, with eigenvalue in the interval 0,2. where O represents for all of media objects in the learned metric space. denotes the normalized graph Laplacian. Joint Graph Regularized Heterogeneous MetricJoint graph regularizationThe formulation of g(U,V) :Minimizing g(U, V) encourage

13、s the smoothness of a mapping over the joint data graph, which is constructed from the initial label informationJoint Graph Regularized Heterogeneous Metric Iterative optimizationObtain orthogonal transformation matrices U and V , they minimize the following object function:where X and Y represent f

14、or two sets of coupled media objects from different media with the same labels. U and V define two orthogonal transformation spaces where media objects in X and Y can be projected as close to each other as possible.Maximize tr(XTUVTY) will minimize function, its singular value decomposition:Joint Gr

15、aph Regularized Heterogeneous MetricFix V and update U Different Q(U,V) with respect to U and V setting it to zero, respectively: Obtain the analytical solution U and V as We alternate between updates to U and V for several iterations to find a locally optimal solution. Here the iteration continues

16、until the cross-validation performance decreases on the training set. In practice, the iteration only repeats several rounds.Joint Graph Regularized Heterogeneous MetricDatasetsWikipedia: 2866 image-text pairs with label from the 10 semantic categories. This dataset is randomly split into a training

17、 set of 2173 documents and a test set of 693 documents. XMedia dataset : 5000 texts, 5000 images, 1000 audio, 500 videos and 500 3D models. This dataset is randomly split into a training set of 9600 media objects and a test set of 2400 media objects.ExperimentsFeatures Images: using bag-of-word mode

18、l. Each image is represented as a histogram of 128-codeword SIFT codebook. texts: each text represented as a 10-topic latent Dirichlet Allocation(LDA) model.Audio: 29-dim MFCC features to represent each clip of audio.Videos: segmenting each clip of video into video shots. Then 128-dimension BoW hist

19、ogram features are extracted for each video keyframe. The final similarity for video is obtained by averaging all of the similarities of the video keyframes. 3D model: Each 3D model is firstly represented as the concatenated 4700-dimension vector of a set of Light-Field descriptors as described in .

20、Then the concatenated vector is reduced to 128-dimension vector based on Principal Component Analysis (PCA)ExperimentsBaseline methods and Evaluation metricsCCA (Canonical correlation analysis): Through CCA we could learn the subspace that maximizes the correlation between two sets of heterogeneous

21、data.CFA(cross-modal factor analysis): it adopts a criterion of minimizing the Frobenius norm between pairwise data in the transformed domainCCA+SMN is current state-of-the-art , since it consider not only correlation analysis but also semantic abstraction for dierent modalities.ExperimentsMAP score

22、sExperimentsPrecision-Recall curvesExperiments多媒体数据的统一表达多媒体数据的表达是指采用哪个一定的数据结构来表示多媒体样本。例如,采用四元组表示web页面中的一幅图像,或者提取图像的底层视觉特征,构成多维向量来表示数据库中的图像。跨媒体检索属于基于内容的多媒体检索范畴,只不过在检索对象上从单一类型的多媒体数据扩充到多种不同类型的多媒体数据,支持数据间的灵活跨越。跨媒体检索的性能很大程度上依赖于相似度匹配算法,而相似度匹配正式以不同类型的多媒体数据所采用的表达方式为依据的。因此数据表达模型的设计师非常基础和重要的。多媒体数据的统一表达设有尚未标注的

23、图像和音频数据集合 ,作为训练数据集合,已知覆盖了Z个语义类别,映射算法描述如下:步骤1 聚类1)对于每一个语义类别Zi,分别提取其中包括的图像和音频数据的底层内容特征,建立相应的特征矩阵SI,SA;2)对于每一个语义类别Z,随机选取m个图像例子Ii进行语义标注;3)计算Ii在底层特征空间上的聚类质心ICri;4)与ICri为起始条件,对数据库中所有的图像数据进行kmeans聚类;5)聚类结果中属于相同类别的图像被赋予与Ii相同的语义标记;6)对音频数据集重复1-4。多媒体数据的统一表达相关性保持映射1)分析图像和音频之间在底层内容特征上的典型相关性,即计算SI和SA对应的子空间基向量Wx和W

24、y; 2)求取视觉和听觉特征响亮映射到子空间中的向量表示:Web环境中的跨媒体相关性推理在具体的应用环境中,如web,往往包含了一些具体的数据特征,这些特征比多媒体数据本身的内容特征蕴含更直接的语义信息,可以用来辅助内容特征进行跨媒体检索,提高检索效率。例如,web连接就可以作为一种辅助特征。跨媒体关联图图模型是一种常用的数据关系表达方式,可以用途模型表达web环境中的图像,以及图像相关的各种特征。这种表达方式不但可以清楚地描述数据之间的各种联系,而且有助于发现数据之间的互补信息。对于多媒体数据而言,多种类型的多媒体数据之间存在着复杂的数据关系,主要可以划分为模态内部(intra-media

25、correlation)和模态之间(cross-media correlation)两种数据关系。链接关系分析分别用V,I,A表示视频、图像和音频数据集,m,n,k分别是数据集V,I,A中的样本个数,用xVi,xIi,xAi分别表示数据库中第i个视频、第i个图像,以及第i个音频数据的特征向量。根据如下两个启发式规则,可以利用web环境中多媒体数据所在网页之间的链接关系,度量不同类型多媒体数据之间的相关性(cross-media distance)大小:规则1:如果两个媒体对象a和b同属于一个web页面,则a和b在语义具有相似性;规则2:如果web页面A指向另一页面B和C,则B中包含的多媒体对象

26、和C中包含的多媒体对象在语义上具有相似性。链接关系分析根据上述启发规则,建立视频-图像、图像-音频和音频-视频的跨媒体关联矩阵LVI,LIA,LAV,以LIA为例,其矩阵元素rij表示多媒体数据 之间的相关值,rij计算方法如下:输入:从web页面获取的图像和音频数据输出:跨媒体相关矩阵LIA1.2. 3.4.5. Construct a symmetric matrix LIA, whose cell lij is the normalized values of rij.基于图模型的全局相关性推理图像音频IaIbIcIdAaAbAc近年来的研究热点Cross-media RetrievalCross-media RankingCross-media HashingCross-collection Topic ModelingFrom FeiWuMission: learn one appropriate metric for ranking multi-modal data to preserve the orders of relevance. For

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

评论

0/150

提交评论