版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领
文档简介
1、Kernel-Based Approaches for Sequence Modeling: Connections to Neural Methods Kevin J LiangGuoyin WangYitong LiRicardo HenaoLawrence Carin Department of Electrical and Computer Engineering Duke University kevin.liang, guoyin.wang, yitong.li, ricardo.henao, Abstract We investigate time-
2、dependent data analysis from the perspective of recurrent kernel machines, from which models with hidden units and gated memory cells arise naturally. By considering dynamic gating of the memory cell, a model closely related to the long short-term memory (LSTM) recurrent neural network is derived. E
3、xtending this setup ton -gram fi lters, the convolutional neural network (CNN), Gated CNN, and recurrent additive network (RAN) are also recovered as special cases. Our analysis provides a new perspective on the LSTM, while also extending it ton -gram convolutional fi lters. Experiments1are performe
4、d on natural language processing tasks and on analysis of local fi eld potentials (neuroscience). We demonstrate that the variants we derive from kernels perform on par or even better than traditional neural methods. For the neuroscience application, the new models demonstrate signifi cant improveme
5、nts relative to the prior state of the art. 1Introduction There has been signifi cant recent effort directed at connecting deep learning to kernel machines 1,5,23,36 . Specifi cally, it has been recognized that a deep neural network may be viewed as constituting a feature mappingx (x), for input dat
6、ax Rm. The nonlinear function(x), with model parameters, has an output that corresponds to ad-dimensional feature vector;(x) may be viewed as a mapping ofxto a Hilbert spaceH, whereH Rd . The fi nal layer of deep neural networks typically corresponds to an inner product|(x), with weight vector H; fo
7、r a vector output, there are multiple, with| i(x) defi ning thei-th component of the output. For example, in a deep convolutional neural network (CNN) 19,(x) is a function defi ned by the multiple convolutional layers, the output of which is ad-dimensional feature map;represents the fully-connected
8、layer that imposes inner products on the feature map. Learningand,i.e., the cumulative neural network parameters, may be interpreted as learning within a reproducing kernel Hilbert space (RKHS) 4, withthe function inH;(x)represents the mapping from the space of the input x to H, with associated kern
9、el k(x,x0) = (x)|(x0), where x0is another input. Insights garnered about neural networks from the perspective of kernel machines provide valuable theoretical underpinnings, helping to explain why such models work well in practice. As an example, the RKHS perspective helps explain invariance and stab
10、ility of deep models, as a consequence of the smoothness properties of an appropriate RKHS to variations in the inputx5,23. Further, such insights provide the opportunity for the development of new models. Most prior research on connecting neural networks to kernel machines has assumed a single inpu
11、tx, e.g., image analysis in the context of a CNN 1,5,23. However, the recurrent neural network (RNN) has also received renewed interest for analysis of sequential data. For example, long short-term These authors contributed equally to this work. 1Implementations can be found at 33rd Conference on Ne
12、ural Information Processing Systems (NeurIPS 2019), Vancouver, Canada. memory (LSTM) 15,13 and the gated recurrent unit (GRU) 9 have become fundamental elements in many natural language processing (NLP) pipelines 16,9,12. In this context, a sequence of data vectors(.,xt1,xt,xt+1,.)is analyzed, and t
13、he aforementioned single-input models are inappropriate. In this paper, we extend to recurrent neural networks (RNNs) the concept of analyzing neural networks from the perspective of kernel machines. Leveraging recent work on recurrent kernel machines (RKMs) for sequential data 14, we make new conne
14、ctions between RKMs and RNNs, showing how RNNs may be constructed in terms of recurrent kernel machines, using simple fi lters. We demonstrate that these recurrent kernel machines are composed of a memory cell that is updated sequentially as new data come in, as well as in terms of a (distinct) hidd
15、en unit. A recurrent model that employs a memory cell and a hidden unit evokes ideas from the LSTM. However, within the recurrent kernel machine representation of a basic RNN, the rate at which memory fades with time is fi xed. To impose adaptivity within the recurrent kernel machine, we introduce a
16、daptive gating elements on the updated and prior components of the memory cell, and we also impose a gating network on the output of the model. We demonstrate that the result of this refi nement of the recurrent kernel machine is a model closely related to the LSTM, providing new insights on the LST
17、M and its connection to kernel machines. Continuing with this framework, we also introduce new concepts to models of the LSTM type. The refi ned LSTM framework may be viewed as convolving learned fi lters across the input sequence and using the convolutional output to constitute the time-dependent m
18、emory cell. Multiple fi lters, possibly of different temporal lengths, can be utilized, like in the CNN. One recovers the CNN 18,37,17 and Gated CNN 10 models of sequential data as special cases, by turning off elements of the new LSTM setup. From another perspective, we demonstrate that the new LST
19、M-like model may be viewed as introducing gated memory cells and feedback to a CNN model of sequential data. In addition to developing the aforementioned models for sequential data, we demonstrate them in an extensive set of experiments, focusing on applications in natural language processing (NLP)
20、and in analysis of multi-channel, time-dependent local fi eld potential (LFP) recordings from mouse brains. Concerning the latter, we demonstrate marked improvements in performance of the proposed methods relative to recently-developed alternative approaches 22. 2Recurrent Kernel Network Consider a
21、sequence of vectors(.,xt1,xt,xt+1,.), withxt Rm. For a language model,xt is the embedding vector for thet-th wordwtin a sequence of words. To model this sequence, we introduce yt= Uht, with the recurrent hidden variable satisfying ht= f(W(x)xt+ W(h)ht1+ b)(1) whereht Rd,U RV d,W(x) Rdm,W(h) Rdd, and
22、b Rd. In the context of a language model, the vectoryt RVmay be fed into a nonlinear function to predict the next wordwt+1in the sequence. Specifi cally, the probability thatwt+1corresponds toi 1,.,V in a vocabulary ofV words is defi ned by element i of vector Softmax(yt+ ), with bias RV . In classi
23、fi cation, such as the LFP-analysis example in Section 6, V is the number of classes under consideration. We constitute the factorizationU = AE, whereA RV jandE Rjd, often withj ? V. Hence, we may writeyt= Ah0t, withh0t= Eht; the columns ofAmay be viewed as time-invariant factor loadings, andh0trepr
24、esents a vector of dynamic factor scores. Letzt= xt,ht1represent a column vector corresponding to the concatenation ofxtandht1; thenht= f(W(z)zt+ b)where W(z)= W(x),W(h) Rd(d+m). Computation ofEhtcorresponds to inner products of the rows ofEwith the vectorht. Letei Rdbe a column vector, with element
25、s corresponding to row i 1,.,j of E. Then component i of h0tis h0i,t= e| iht = e| if(W (z)zt + b)(2) We viewf(W(z)zt+b)as mappingztinto a RKHSH, and vectoreiis also assumed to reside within H. We consequently assume ei= f(W(z) zi+ b)(3) 2 (a)(b)(c) Figure 1: a) A traditional recurrent neural network
26、 (RNN), with the factorizationU = AE. b) A recurrent kernel machine (RKM), with an implicit hidden state and recurrence through recursion. c) The recurrent kernel machine expressed in terms of a memory cell. where zi= xi,h0. Note that hereh0also depends on indexi, which we omit for simplicity; as di
27、scussed below, xiwill play the primary role when performing computations. e| iht = e| if(W (z)zt + b) = f(W(z) zi+ b)|f(W(z)zt+ b) = k( zi,zt)(4) wherek( zi,zt) = h( zi)|h(zt)is a Mercer kernel 29. Particular kernel choices correspond to different functionsf(W(z)zt+b), andis meant to represent kerne
28、l parameters that may be adjusted. We initially focus on kernels of the formk( z,zt) = q( z|zt) =h| 1ht,2whereq()is a function of parameters,ht= h(zt), andh1is the implicit latent vector associated with the inner product,i.e., h1= f(W(x) x + W(h)h0+ b). As discussed below, we will not need to explic
29、itly evaluatehtorh1 to evaluate the kernel, taking advantage of the recursive relationship in (1). In fact, depending on the choice ofq() , the hidden vectors may even be infi nite-dimensional. However, because of the relationshipq( z|zt) =h| 1ht, for rigorous analysisq()should satisfy Mercers condi
30、tion 11,29. The vectors(h1,h0,h1,.)are assumed to satisfy the same recurrence setup as (1), with each vector in the associated sequence( xt, xt1,.)assumed to be the same xiat each time,i.e., associated withei,( xt, xt1,.) ( xi, xi,.). Stepping backwards in time three steps, for example, one may show
31、 k( zi,zt) = q x| ixt+ q x | ixt1+ q x | ixt2+ q x | ixt3+ h| 4ht4 (5) The inner producth| 4ht4 encapsulates contributions for all times further backwards, and for a sequence of lengthN,h| NhtN plays a role analogous to a bias. As discussed below, for stability the repeated application ofq()yields d
32、iminishing (fading) contributions from terms earlier in time, and therefore for large N the impact ofh| NhtN on k( zi,zt) is small. The overall model may be expressed as h0t= q(ct) ,ct= ct+ q(ct1) , ct= Xxt(6) wherect Rjis a memory cell at timet, rowiof Xcorresponds to x| i, andq(ct)operates pointwi
33、se on the components ofct(see Figure 1). At the start of the sequence of lengthN,q(ctN)may be seen as a vector of biases, effectively corresponding toh| NhtN; we henceforth omit discussion of this initial bias for notational simplicity, and because for suffi ciently largeNits impact onh0tis small. N
34、ote that via the recursive process by whichct is evaluated in (6), the kernel evaluations refl ected byq(ct) are defi ned entirely by the elements of the sequence( ct, ct1, ct2,.). Let ci,trepre- sent thei-th component in vector ct , and defi next= (xt,xt1,xt2,.). Then the sequence ( ci,t, ci,t1, ci
35、,t2,.) is specifi ed by convolving in time xiwithxt, denoted xi xt. Hence, the jcomponents of the sequence( ct, ct1, ct2,.) are completely specifi ed by convolvingxtwith each of thej fi lters, xi,i 1,.,j,i.e., taking an inner product of xiwith the vector inxtat each time point. In (4) we represented
36、h0i,t= q(ci,t)ash0i,t= k( zi,zt); now, because of the recursive form of the model in (1), and because of the assumptionk( zi,zt) = q( z| izt), we have demonstrated that we 2One may also design recurrent kernels of the formk( z,zt) = q(k z ztk2 2)14, as for a Gaussian kernel, but if vectorsxt and fi
37、lters xiare normalized (e.g.,x| txt= x | i xi= 1), thenq(k zztk 2 2)reduces toq( z |zt). 3 may express the kernel equivalently ask( xi xt) , to underscore that it is defi ned entirely by the elements at the output of the convolution xi xt. Hence, we may express componentiofh0tas h0i,t= k( xi xt). Co
38、mponent l 1,.,V of yt= Ah0tmay be expressed yl,t= j X i=1 Al,ik( xi xt)(7) whereAl,irepresents component(l,i)of matrixA. Considering (7), the connection of an RNN to an RKHS is clear, as made explicit by the kernelk( xi xt) . The RKHS is manifested for the fi nal outputyt, with the hiddenhtnow absor
39、bed within the kernel, via the inner product (4). The feedback imposed via latent vectorhtis constituted via update of the memory cellct= ct+ q(ct1)used to evaluate the kernel. Rather than evaluatingyt as in (7), it will prove convenient to return to (6). Specifi cally, we may consider modifying (6)
40、 by injecting further feedback via h0t, augmenting (6) as h0t= q(ct) ,ct= ct+ q(ct1) , ct= Xxt+ Hh0t1(8) where H Rjj, and recallingyt= Ah0t(see Figure 2a for illustration). In (8) the input to the kernel is dependent on the input elements(xt,xt1,.)and is now also a function of the kernel outputs at
41、the previous time, viah0t1. However, note thath0t is still specifi ed entirely by the elements of xi xt, for i 1,.,j. 3Choice of Recurrent Kernels given2 f 1)-gram fi lters, respectively. This underscores that at the heart of such models, one performs convolutions between the sequence of data(.,xt+1,xt,xt1,.) and fi lters
温馨提示
- 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
- 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
- 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
- 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
- 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
- 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
- 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。
最新文档
- 2026事业单位工勤技能-上海-上海土建施工人员三级(高级工)历年参考题库含答案详解
- 校园卫生保洁常态化管理制度
- 小学单元作业整体设计方案
- 应急物资材料采购管理方案
- 食品生产企业安全生产管理制度
- 有限空间作业安全管理实施细则
- 物流运输企业温室气体排放方案
- 装配式混凝土结构工程技术交底
- 应急物资管护作业管理制度
- 土石方开挖回填专项施工方案
- 2026年非接触式物位仪表行业技术创新与应用报告
- 2026中国新能源电池材料技术突破与市场前景研究报告
- 2026年《中国肺动脉高压诊断治疗指南(2026版)》
- 手术室净化监理细则
- 2026年发展对象培训班考试题库(含完整答案解析)
- 2026年西学中考试题库及答案
- 中国精神:兴国强国之魂
- 大酒店合作经营合同协议模板
- ASCVD一级预防:他汀联合依折麦布策略
- 近年文言文《岳阳楼记》中考真题30套
- 2025年统计学期末考试题库:统计学在法律学中的应用综合案例分析试题集
评论
0/150
提交评论