版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
1、Create Photo-Realistic Talking FaceChangbo Hu*This work was done during visiting Microsoft Research China with Baining Guo and Bo ZhangOutlineIntroduction of talking faceMotivationsSystem overviewTechniquesConclusionsIntroductionWhat is a talking faceFace (lip) animation, driven by voiceApplications
2、The process of talking faceFace modelMotion captureMapping between audio and video Rendering, Photo-realistic?LiteraturesWalter,93, DecFace, 2Dwire frame modelTerzopoulos,95, Skin and muscle modelBreglar,97, Video Rewrite, Sample image based TS Huang,98,Mesh model from range dataPoggio,98, MikeTalk,
3、 Viseme morphingGuenter,99, Making face, 3D from multicamera Zhengyou Zhang, 00, 3D face modeling from video through epipolar constraintCosatto,00, Planar quads model Some Face modelsMotivationsAim: a graphics interface for conversation agentPhoto-realisticDriven by ChineseSmooth connection between
4、sentencesExtended from “Video rewrite”System overview:Pipeline of the system(1)Video with SoundImagesSoundPose trackingPhoneme segmentationAnnotationLip motion TrackingTrain databaseSystem overview: Pipeline of the system(2)New textWav soundTTS systemTriphone sequenceSegmentationSynthesized triphone
5、 sequenceTrain databaseLip motion sequenceRewrite to facesBackground sequenceTechniquesAnalysis:Audio processImage processSynthesisLip image Background imageStitch togetherAudio part:Sound SegmentationGiven the wav file and the scriptUsing HMM to train the segment systemSegment wav file to phoneme s
6、equenceExample of the segmentation result:SILOPEN023SILOPEN2442s4361if46274j7580ia18197sh98109ang1110121y122130e4131133y134145in2146154h155164ang2165194Annotation with PhonemeUsing phoneme to annotate video framesEach phoneme in a sentence corresponds to a short time of video sequenceTraining Senten
7、ceAudio FramesVideo FramesPhoneme SequenceFrames for Phoneme1Frames for Phoneme1Phoneme1Frames for Phoneme2Frames for Phoneme2Phoneme2Phoneme Distance Analysis Phoneme&triphone basicsChinese Phoneme vs. English PhonemeDistance Metrics definitionsResultsPhoneme BasicsPhonemes represents the basic ele
8、ments in speech. All possible speech can be represented by combination of phonemes.CH, JH, S, EH, EY, OY, AE, SILTriphone are three consecutive phonemes. It not only represents pronounce characteristics but also contains context information.T-IY-P, IY-P-AA, P-AA-TChinese Phoneme vs. EnglishChinese p
9、honeme has two basic groups: Initials and Finals.Initials: B, P, M, F, Finals: a3, o1, e2, eng3, iang4, ue5, Chinese finals each has 5 tones: 1,2,3,4,5.Different tones: a1, a2, a3, a4, a5.Chinese finals actually is not a basic elements of speech.For example: iang1, iao1, uang1, iong1Chinese phoneme
10、set is much larger than English.Phoneme Distance AnalysisDefine the distance between any two phonemes.Since we only synthesis video but not sound, so tone is ignoredLip shape motion is the core element for distance metrics.Phoneme Distance AnalysisVideo 1Video 2Video 4Video 1Video 2Video 3Phoneme 1:
11、Phoneme 2:Time Align to an uniform lengthVideo 2Video 3Video 4Video 2Video 1Video 1Average the videos to get an average videoVideo AverageVideo AverageBy comparing the two aligned average videos, we generate the distance matrix of the whole phoneme set.Image part: Pose TrackingAssume a plane model f
12、or faceStandard minimization method to find transform matrix (affine transform)Black,95Mask is used to constrain interests part of the faceTemplate PictureMask ImagePose trackingMotion prediction using parameters with physical meaningPose TrackingSome tracking results:Lip Motion TrackingUsing Eigen
13、Points (Covell, 91)Feature Points include Jaw, lip and teethTraining database specified manuallyAuto tracking through all pose-tracked imagesLip motion trackingLip Motion TrackingTrain Database (hand-labeled)Auto Tracking ResultsSynthesis new sentencesNew text converted by TTS system to wavWav is se
14、gmented to phoneme sequenceUsing DP to find an optimal video sequence from the training databaseTime-align triphone videos and stitch them together.Transform the lip sequence and paste them to background faces.Lip sequence synthesisOptimal phoneme sequencesTriphone 1Triphone 2Triphone 5Triphone 3Tri
15、phone 4Triphone 6Triphone 7Triphone 8Triphone BTriphone 9Triphone ATriphone CNew phoneme sequencesNew phoneme sequencesDynamic ProgrammingBeginTriphone1Triphone3Triphone2Triphone4EndTriphone5Edge Cost DefinitionTwo parts: phoneme distance: 3 phonemes distances added togetherLip shape distance for th
16、e overlap portion of triphone videoWeighted add together two partBackground video generationBackground is a video sequence when the virtual character spoke something elseSimilarity measurement of backgroundSelect “standard frame”The frame with maximal number of frames similar to itFilter out the frames with jerkinessStitch the time-aligned result to background facesWrite back with a maskTransform the synthesized lip to the background faceMask image for write-back operationOriginal background fram
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2026年中国非处方药(OTC)市场分析预测研究报告
- 2026年功能性膜材料市场调查报告
- 2026年城管笔模拟题目答案及答案详解
- 2026年密闭鼓风炉备料工标准化作业考核模拟试卷(含答案)
- 托育知识题库及答案详解
- 2026年攀枝花市西区住房和城乡建设局招聘临聘人员备考题库(含答案)
- 2026年摊铺机项目投资分析及可行性报告
- 2026年会计信息系统工程师实践技能测验模拟试卷(含答案)
- 2026年高级工程师职业资格考试模拟题及答案详解
- 我国银行笔试题库(含答案)
- 2025年国家公务员考试证监会历年面试试题及解析
- 《工业机器人系统操作与运维》 课件 第22讲-机器人圆弧编程与焊接
- 挖机破碎合同协议书范本
- 儿童文学概论 课件全套 第1-12章 儿童文学基本理论-儿童文学整本书阅读指导
- 售前工程师笔试题及答案
- 2025年大学生信息素养大赛培训考试题库500题(附答案)
- 中职学校“双师型”教师队伍建设与激励机制的实践与研究
- 2024-2025学年云南省昆明八中八年级(上)期中数学试卷(含答案)
- 安全伴我行-大学生安全教育智慧树知到期末考试答案章节答案2024年哈尔滨工程大学
- 支气管哮喘病例讨论课件
- 2023-2024学年广西百色市高二(上)期末数学试卷(含解析)
评论
0/150
提交评论