版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
1、Machine Perception and Interaction Group (MPIG) .cn 跟我学CS231(7)袁洪慧MPIG Open Seminar 0220公众号:mpig_robotRecurrent Neural NetworkRNN的应用RNN的分类Recurrent Neural NetworkRNN的正向传播Truncated Backpropagation LSTMOther RNN Variantssummary目录RNN的应用对于序列化的特征任务,都适合用RNN来解决:情感分析关键字提取语音识别机器翻译股票分析RNN的分类“Vanilla” Neural N
2、etworkVanilla Neural NetworksRecurrent Neural Networks: Process Sequencese.g. Image Captioning image - sequence of wordse.g. Sentiment Classification sequence of words - sentimente.g. Machine Translation seq of words - seq of wordse.g. Video classification on frame levelRecurrent Neural Networkusual
3、ly want to predict a vector at some time stepsRecurrent Neural NetworkhRecurrent Neural NetworkWe can process a sequence of vectors x by applying a recurrence formula at every time step: new statesome function with parameters Wold state input vector atsome time stepRecurrent Neural NetworkWe can pro
4、cess a sequence of vectors x by applying a recurrence formula at every time step:Notice: the same function and the same set of parameters are used at every time step.(Simple) Recurrent Neural NetworkThe state consists of a single “hidden” vector h:RNN的展开图RNN的正向传播RNN: Computational GraphRe-use the sa
5、me weight matrix at every time-step:RNN: Computational Graph: Many to Many RNN: Computational Graph: Many to OneRNN: Computational Graph: One to ManySequence to Sequence: Many-to-one + one-to-manyMany to one: Encode input sequence in a single vectorOne to many: Produce output sequence from single in
6、put vectorTruncated Backpropagation Backpropagation through time梯度截断(Gradient Clipping)为梯度设置阈值,超过该阈值的梯度值都会被cut,这样更新的幅度就不会过大,因此容易收敛。具体做法:Truncated Backpropagation through timeTruncated Backpropagation through timeVanilla RNN Gradient FlowComputing gradient of h0 involves many factors of W (and repeat
7、ed tanh) Bengio et al, “Learning long-term dependencies with gradient descent is difficult”, IEEE Transactions on Neural Networks, 1994 Pascanu et al, “On the difficulty of training recurrent neural networks”, ICML 2013Largest singular value 1: Exploding gradients Largest singular value 1: Vanishing
8、 gradients Gradient clipping: Scale Computing gradient gradient if its norm is too bigSimple-RNN在实际应用中并不多,原因:如果输入越长的话,展开的网络就越深,对于“深度”网络训练的困难最常见的是 Gradient Explode 和 Gradient Vanish 的问题。Simple-RNN基于先前的词预测下一个词,但在一些更加复杂的场景中,例如,“I grew up in France I speak fluent French” “France”则需要更长时间的预测,而随着上下文之间的间隔不断
9、增大时,Simple-RNN会丧失学习到连接如此远的信息的能力。LSTM(Long Short-Term Memory)Long Short Term Memory (LSTM)RNN和LSTM框图 LSTM的核心思想逐步理解 LSTM之遗忘门逐步理解 LSTM之输入门LSTM还需要记住东西,所以有了图示“记忆门”。逐步理解 LSTM逐步理解 LSTM之输出门Other RNN VariantsGRU(Gated Recurrent Unit)GRU是和LSTM功能几乎一样的另一种网络。最终的模型比标准的 LSTM 模型要简单,也是非常流行的变体SummaryRNNs allow a lot
10、of flexibility in architecture design Vanilla RNNs are simple but dont work very well Common to use LSTM or GRU: their additive interactions improve gradient flow Backward flow of gradients in RNN can explode or vanish. Exploding is controlled with gradient clipping. Vanishing is controlled with additive interactions (LSTM) Better/simpler architectures are a hot
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2026年六月员工培训与发展方案
- 2026年危险化学品试题及答案全套试题及答案
- 2026届湖北省襄阳市第四中学高三下学期考前测试历史试题(含答案)
- 2026年初中历史综合试卷
- 班主任年度思想工作个人总结2026(3篇)
- 等级医院评审自评报告模版(3篇)
- 乙肝检验常见试题及答案解析
- 六年级下册数学北师大含答案 神奇的莫比乌斯带
- 四年级下册数学北师大含答案 栽蒜苗(一)1
- 航天与生物考试题目及答案
- (正式版)DB44∕T 2828-2026 城镇燃气安全检查与评估标准
- 2026 文物保护责任监理师《法律法规与工程管理》考试参考题库 800题 -含答案
- 2026年贵州幼儿园教师进城选调考试试题
- 2025年(数据安全管理员)数据安全管理试题及答案
- 2025年度组织生活会个人发言提纲1
- 龟类科普教学课件
- 2026年及未来5年市场数据中国机械硬盘(HDD)行业市场发展数据监测及投资策略研究报告
- 译林版必修一Unit4重难点词汇讲解
- 术中神经电刺激在神经外科推广策略
- 眼科进修汇报课件
- 2025年艺术衍生品品牌联名合同
评论
0/150
提交评论