VCS培训教材(中文)_第1页
VCS培训教材(中文)_第2页
VCS培训教材(中文)_第3页
VCS培训教材(中文)_第4页
VCS培训教材(中文)_第5页
已阅读5页,还剩72页未读 继续免费阅读

下载本文档

版权说明:本文档由用户提供并上传,收益归属内容提供方,若内容存在侵权,请进行举报或认领

文档简介

1、VERITAS Cluster Server,Page 2,CONTENTS,VCS 预备知识 VCS的基本概念和术语 VCS的管理集群服务 VCS常见问题的解决 总结,VCS预备知识,Page 4,Metropolitan HA Disaster Recovery (over SAN, MAN or LAN),Wide Area Disaster Recovery,VCS的几个常用场景,LAN,Local Clustering,MAN,WAN,Cluster Server,Cluster Server, Volume Manager,Cluster Server, Volume Manage

2、r, Volume Replicator, Global Cluster Manager,Page 5,Clustering: Application and Database Failover数据库和应用的失效转移,Page 6,A Cluster View of Applications,A “service group” is a collection of resources that monitor the status of an application (服务组是各种监控应用状态的资源的集合) Application failover is controlled by the s

3、ervice group(应用的失效转移是由服务组来控制的),1.,1.,2.,2.,3.,4.,5.,Page 7,Active/Passive Clustering(主备方式)“asymmetric configuration”(非对称配置),Primary server hosts application(主机提供服务,Primary server FAILS,Secondary server hosts primary application,备机处于等待状态,一旦主机发生故障,接管服务),Page 8,Active/Active Clustering (互备方式)“symmetric

4、 configuration”(对称配置),Primary server hosts primary application,Primary server FAILS,Secondary server hosts both primary and secondary applications,Secondary server hosts secondary application 两个节点提供不同的服务,互相备用,当一个节点故障,服务马上有第二个节点接管服务,VCS的基本概念和术语,Page 10,集群,SCSI JBODS,Several networked systems几个节点 Shar

5、ed storage共享存储 Single administrative entity单个管理节点 Peer monitoring相互监控,Fibre Switches,Page 11,systems 系统,Members of a cluster集群的一个成员 Referred to as nodes也称之为节点 Contain copies of: 包括如下内容 Communication protocol configuration files通信协议的配置文件 VCS configuration files VCS的配置文件 VCS libraries and directories

6、VCS的安装目录 VCS scripts and daemons VCS的脚本和后台程序 Share a single dynamic cluster configuration 共享一个动态的集群配置 Provide application services 提供应用的服务,Page 12,Service Groups 服务组,A service group is a related collection of resources.服务组是资源的一个集合 Resources in a service group must be available to the system.服务组中的资源在

7、系统中必须是可用的 Resources and service groups have interdependencies.服务组和资源存在相互依赖关系,NFS Service Group,NFS,IP,Disk,Mount,Share,NIC,Page 13,Service Group Types 服务组的类型,Failover失效转移 Can be partially or fully online on only one server at a time同一时间只能在一台机器上运行 VCS controls stopping and restarting the service grou

8、p when components fail当服务组某个资源出错时,VCS控制它的停止和重启 Parallel并行 Can be partially or fully online on multiple servers simultaneously可以同时在多台机器上运行 Examples: Oracle Parallel Server Web, FTP servers,Page 14,Resources 资源,VCS objects that correspond to hardware or software components包括软件和硬件组件 Monitored and contr

9、olled by VCS通过VCS来监控和控制 Classified by type通过资源类型分类 Identified by unique names and attributes通过唯一的名称和属性来标识 Can depend on other resources within the same service group在同一服务组中可依赖其他资源,Page 15,Resource Types 资源类型,General description of the attributes of a resource通常描述一种资源的属性 Example Mount resource type a

10、ttributes:例如mount资源类型的属性 MountPoint 挂载点 BlockDevice 挂载设备 Other example resource types:其他类型的资源 Disk磁盘 Share共享 IP浮动IP NIC网卡,Page 16,Agents 代理,Processes that control resources 控制资源的程序 One agent per resource type每种类型的资源对应一个代理 Agent controls all resources of that type.一个代理控制对应类型的所有资源 Agents can be added

11、into VCS agent framework.用户可以加入自己的代理到VCS的框架中,Page 17,Dependencies依赖关系,Resources can depend on other resources. 资源可以依赖其他资源 Parent resources depend on child resources. 父资源依赖子资源 Service groups can depend on other service groups.服务组可以依赖其他服务组 Resource types can depend on other resource types.资源类型之间也存在依赖,

12、比如IP类型必须依赖NIC类型 Rules govern service group and resource dependencies.资源和服务组之间的依赖关系由规则管理 No cyclic dependencies are allowed.不允许出现循环依赖,Mount,Disk,(Parent),(Child),Page 18,Private Network 私有网络,Minimum two communication channels with separate infrastructure:至少需要两条独立的通信链路 Multiple NICs (not just ports)多块

13、网卡 Separate hubs, if used独立的hub Heartbeat communication determines which systems are members of the cluster.心跳之间的通信决定哪些系统是集群的成员 Cluster configuration broadcast updates cluster systems with status of each resource and service group.集群中的资源和服务组的状态信息通过广播更新到各个节点,Page 19,Low Latency Transport (LLT)低时延传输协议

14、,Provides fast, kernel-to-kernel communications提供快速,内核到内核的通信 Is connection oriented Is not routable 不需要路由 Uses Data Link Provider Interface (DLPI) over Ethernet 使用以太网的链路层,Page 20,Group Membership Services/Atomic Broadcast (GAB),Manages cluster membership 管理集群成员 Maintains cluster state 维护集群状态 Uses br

15、oadcasts 使用广播 Runs in kernel over Low Latency Transport (LLT) 运行在llt之上,Page 21,VCS Engine (had)VCS的引擎,Maintains configuration and state information for all cluster resources维护整个集群的所有资源的配置和状态信息 Uses GAB to communicate among cluster systems通过gab与集群的其他成员通信 Is monitored by hashadow process由后台进程hashadow来

16、监控,hashadow,SystemA,SystemB,LLT,LLT,Hardware,Kernel,Private Network,had,had,Page 22,VCS Architecture总体架构,SystemA,SystemB,Shared Cluster Configuration in Memory,Hardware,Kernel,Resources,Agents,Mount,hashadow,hashadow,/v,Disk,c1d0t0s0,hme0,NIC,IP,,had,/v,Disk,c1d0t0s0,hme0,NIC,IP,had,Mount,LL

17、T,LLT,GAB,GAB,VCS管理集群服务,Page 24,Cluster Configuration集群配置,Page 25,Starting VCS 启动VCS,main.cf,Cluster Conf,Private Network,System2,System3,System1,Page 26,Starting VCS: Second System,main.cf,had hashadow,Private Network,had hashadow,System2,System3,Cluster Conf,Cluster Conf,System1,Page 27,Starting V

18、CS: Third System,main.cf,had hashadow,had hashadow,main.cf,main.cf,had hashadow,System1,System2,System3,Shared Cluster Configuration in Memory,Private Network,Page 28,Stopping VCS 停止VCS,Page 29,The hastop Command 停止命令,The hastop command stops the VCS engine. Syntax: hastop option arg -option Options

19、: -local -force | -evacuate -sys sys_name -force | -evacuate -all -force Example: hastop -sys train4 -evacuate,Page 30,Displaying Cluster Status 显示集群的状态,The hastatus Command Displays status of items in the cluster. Syntax: hastatus -option arg -option arg Options: -group service_group -summary Examp

20、le: hastatus -group OracleSG,Page 31,Protecting the Cluster Configuration 保护集群的配置,Cluster configuration opened; .stale file created Resources added to cluster configuration in memory; main.cf out of sync with memory configuration Changes saved to disk; .stale removed,haconf -makerw,Cluster Conf,hare

21、s add ,haconf dump -makero,main.cf ,main.cf,.stale,Page 32,Opening and Saving the Cluster Configuration打开和保存集群配置,The haconf command opens, closes, and saves the cluster configuration. Syntax: haconf option -option Options: -makerwOpens configuration -dump Saves configuration -dump makero Saves and c

22、loses configuration Example: haconf -dump -makero,Page 33,Starting VCS with a Stale Configuration,main.cf,had hashadow,Private Network,hastart,had hashadow,System2,System3,main.cf,.stale,main.cf,Page 34,Forcing VCS to Start on the Local System,System1,main.cf,Private Network,hastart -force,had hasha

23、dow,System2,System3,main.cf,.stale,Cluster Conf,main.cf,Page 35,Forcing a System to Start,Page 36,The hasys Command,Alters or queries state of had Syntax: hasys option arg Options: -force system_name -list -display system_name -delete system_name -add system_name Example: hasys -force train11,Page 3

24、7,Propagating a Specific Configuration配置文件的传播,Stop VCS on all systems in the cluster and leave applications running: hastop -all force Start VCS stale on all other systems: hastart stale The -stale option causes these systems to wait until a running configuration is available from which they can bui

25、ld. Start VCS on the system with the main.cf that you are propagating: hastart,Page 38,Summary of Start Options启动总结,The hastart command starts the had and hashadow daemons. Syntax: hastart -option Options: -stale -force Example: hastart -force,Page 39,Validating the Cluster Configuration验证集群配置,The h

26、acf utility checks the syntax of the main.cf file. Syntax: hacf -verify config_directory Example: hacf -verify /etc/VRTSvcs/conf/config,Page 40,Modifying Cluster Attributes修改集群属性,The haclus command is used to view and change cluster attributes. Syntax: haclus option arg Options: -display -help -modi

27、fy -modify modify_options -value attribute -notes Example: haclus value ClusterLocation,Page 41,Startup States and Transitions启动的状态和迁移,Page 42,Shutdown States and Transitions停止的状态和迁移,RUNNING,LEAVING,EXITING,EXITED,EXITING_FORCIBLY,FAULTED,hastop,hastop -force,Resources offlined, agents stopped,Unexp

28、ected exit,VCSTroubleshooting,Page 44,从以下几个方面来监控VCS:,VCS的日志文件 系统的日志文件 使用hastatus命令查看VCS的状态 SNMP 事件告警机制 集群管理图形界面cluster manager,Page 45,VCS Log Entries,VCS引擎日志: /var/VRTSvcs/log/engine_A.log 通过GUI图形界面查看日志或者 hamsg 命令: hamsg engine_A Example entries: TAG_D 2001/04/03 12:17:44 VCS:11022:VCS engine (had)

29、 started TAG_D 2001/04/03 12:17:44 VCS:10114:opening GAB library TAG_C 2001/04/03 12:17:45 VCS:10526:IpmHandle:recv peer exited errno 10054 TAG_E 2001/04/03 12:17:52 VCS:10077:received new cluster membership TAG_E 2001/04/03 12:17:52 VCS:10080:Membership: 0 x3, Jeopardy: 0 x0,Page 46,代理日志:Agent Log

30、Entries,代理日志在 /var/VRTSvcs/log目录下面 日志文件用 AgentName_A.log来命名,如:IP_A.log 日志级别的设置: none error (默认设置) info debug all 通过命令来改变日志级别: hatype -modify res_type LogLevel debug,Page 47,集群通信问题解决:,使用命令 hastatus summary检查VCS 如果输出类似如下,则表明集群之间的通信有问题 VCS:11307:Node has not received cluster membership yet, cannot proc

31、ess HA command 如果输出类似如下,则表明VCS的引擎启动有问题 hatest1 STALE ADMIN WAIT: all system stale 首先用lltconfig命令检查llt模块是否是running状态,如果不是检查/etc/llttab文件,Page 48,LLT模块问题解决:,检查/etc/llthost文件,主机名必须与/etc/llttab中的主机名保持一致,主机序列号必须在0-31范围内 如果llt的状态是running,用命令lltstat n检查是否所有的心跳线都是好的 ,请先确认在/etc/llttab中配置的网卡是否都是UP状态的,可以用ifcon

32、fig查看,类似输出如下: LLT node information: Node State Links * 0 test-smc3 OPEN 3 1 storage-1 OPEN 3,Page 49,GAB模块问题解决:,首先检查GAB模块是否已经运行,gabconfig a 如果输出如下,则表明GAB模块有问题,请检查/etc/gabtab文件, GAB Port Memberships 如果GAB一起动马上关闭了,请检查LLT模块是否有问题 如果没有h端口的输出则表明HAD 有问题,正常的输出如下: GAB Port Memberships = Port a gen a76401 mem

33、bership 01 Port h gen a76404 membership 01,Page 50,HAD模块问题解决,首先确认LLT模块和GAB模块已经正确启动 使用hacf verify /etc/VRTSvcs/conf/config检查VCS的配置文件是否配置正确,无输出则表明是正确的 确认VCS的license是否是正确的:vxlicrep,如果输出类似如下,则需要重新输入license vxlicrep ERROR V-21-3-1003 There are no valid VERITAS License keys installed in the system. 重新输入有效

34、的license,使用命令vxlicinst,按照提示输入license 使用命令hastatus -sum 查看状态 STALE_ADMIN_WAIT: The system has a stale configuration and no other system is in a RUNNING state. ADMIN_WAIT: The system cannot build or obtain a valid configuration.,Page 51,STALE_ADMIN_WAIT,To recover from STALE_ADMIN_WAIT state:从这个状态恢复 V

35、isually inspect the main.cf file to determine whether it is valid.验证配置文件是否正确 Edit the main.cf file, if necessary.如有必要修改该文件 Verify the syntax of main.cf, if modified. 修改之后验证语法的正确性 hacf verify config_dir Start VCS on the system with the valid main.cf file:强制启动VCS使用有效的配置文件 hasys -force system_name All

36、other systems perform a remote build from the system now running.其他的节点可以通过这个启动的节点进行远程启动,Page 52,ADMIN_WAIT,A system can be in the ADMIN_WAIT state under these circumstances:下列情形之一可能会出现这个状态 A .stale flag exists and the main.cf file has a syntax problem. 配置文件有问题 A disk error occurs affecting main.cf d

37、uring a local build.本地启动的时候硬盘有问题 The system is performing a remote build and last running system fails.该节点正在远程启动,结果那个节点失效了 Restore main.cf and use the procedure for STALE_ADMIN_WAIT.,Page 53,Identifying Other Problems 其他问题的确定,After verifying that HAD, LLT, and GAB are functioning properly, run hasta

38、tus sum to identify problems in other areas:在检查了HAD,LLT和GAB正确之后就要使用hastatus sum 来确定其他区域的问题 Service groups 服务组 Resources 资源 Agents and resource types代理和资源类型,Page 54,Service Group Problems: Group Not Configured to Start or Run 服务组的问题,没有配置为自动启动,Service group not onlined automatically when VCS starts:Ch

39、eck AutoStart and AutoStartList attributes: VCS启动的时候服务没有自动online,先检查AutoStart 和AutoStartList 这两个属性 hagrp display service_group Service group not configured to run on the system:服务组没有配置为在这个节点上运行 Check the SystemList attribute. 检查SystemList 属性 Verify that the system name is included.确认这个节点属于这个集群,Page

40、55,Service Group AutoDisabled 服务组自动失效,Autodisable occurs when:由下列情形会发生自动失效 GAB sees a system but had is not running on the system.节点已经运行gab,但是没有启动VCS的had Resources of the service group are not fully probed on all systems in the SystemList.在所有的检点上服务组的资源没有全部探测到 A particular system is visible through d

41、isk heartbeat only.通过磁盘心跳只有部分节点是可见的 Make sure that the service group is offline on all systems in SystemList attribute.确认这个服务组在所有的节点上都是offline的 Clear the AutoDisabled attribute: 清除自动失效属性 hagrp autoenable service_group -sys system Bring the service group online.将这个服务组online,Page 56,Service Group Not

42、Fully Probed 服务组没有全部探测到,Usually a result of improperly configured resource attributes: 通常是资源的属性没有正确的配置 Check ProbesPending attribute:检查这个属性 hagrp -display service_group Check which resources are not probed:查看哪个资源没有探测到 hastatus -sum Check Probes attribute for resources:检查资源的属性 hares -display To probe

43、 resources: 探测这个资源 hares probe resource -sys system,Page 57,Service Group Frozen 服务组冻结,Verify value of Frozen and TFrozen attributes: 确认这两个属性的值 hagrp -display service_group Unfreeze the service group: 解冻这个服务组 hagrp -unfreeze group -persistent If you freeze persistently, you must unfreeze persistentl

44、y.如果是持久冻结,解冻的时候必须要是持久解冻,Page 58,Service Group Is Not Offline Elsewhere服务组在任何地方都没有offline,Determine which resources are online/offline:确定哪些资源是online和offline的 hastatus -sum Verify the State attribute: 确认状态属性 hagrp -display service_group Offline the group on the other system:在其他节点offline这个服务组 hagrp -of

45、fline Flush the service group:使这个服务组可以被部分拉起 hagrp -flush service_group -sys system,Page 59,Service Group Waiting for Resource服务组在等待某个资源,Review Istate attribute of all resources to determine which resource is waiting to go online.查看哪个资源正在等待online的过程中 Use hastatus to identify the resource.使用hastauts来确

46、认这个资源 Make sure the resource is offline (at the operating system level). Clear the internal state of the service group: hagrp flush service_group -sys system Bring all other resources in the service group offline and try to bring these resources online on another system. Verify that the resource wor

47、ks properly outside VCS. Check for errors in attribute values.,Page 60,Incorrect Local Name主机名不一致,A service group cannot be brought online if the system name is inconsistent in llthosts, llttab, or main.cf files. 如果在llthosts,llttab和main.cf中的主机名不一致则这个服务组不会被online Check each file for consistent use of

48、 system names.检查这些文件 Correct any discrepancies. 修改成一致的 If main.cf is changed, stop and restart VCS. 如果main.cf 被修改了,停止和重启VCS If ltthosts or ltttab is changed:如果llthosts和llttab修改了,停止VCS,gab,和llt,重新启动llt,gab和VCS Stop VCS, GAB, and LLT. Restart LLT, GAB, and VCS.,Page 61,Concurrency Violations 网络冲突,Occu

49、rs when a failover service group is online or partially online on more than one system失效转移类型的服务组在多个节点上运行就会导致冲突 Notification provided by the Violation trigger: Invoked on the system that caused the concurrency violation Notifies the administrator and takes the service group offline on the system caus

50、ing the violation Configured by default with the violation script in /opt/VRTSvcs/bin/triggers Can be customized: Send message to the system log. Display warning on all cluster systems. Send e-mail messages.,Page 62,Service Group Waiting for Resource to Go Offline服务组等待资源offline,Identify which resour

51、ce is not offline:确定哪个资源没有offline hastatus summary Check logs.检查日志 Manually bring the resource offline, if necessary.必要的时候手动offline这个资源 Configure ResNotOff trigger for notification or action. 可以配置ResNotOff trigger 这个处罚脚本,一旦发生这种情况可以报告给管理员,Page 63,Resource Problems: Unable to Bring Resources Online 资源

52、问题:不能将某个资源online,Possible causes of failure while bringing resources online:不能将资源online的原因 Waiting for child resources 等待子资源 Stuck in a WAIT state 在一个等待状态 Agent not running 代理没有运行,Page 64,Problems Bringing Resources Offline 资源offline的问题,Waiting for parent resources to come offline等待父资源offline Waitin

53、g for a resource to respond等待这个资源的响应 Agent not running代理没有运行,Page 65,Critical Resource Faults 严重资源错误,Determine which critical resource has faulted:查看严重资源错误 hastatus summary Make sure that the resource is offline. 确认这个资源已经offline Examine the engine log. 检查日志 Fix the problem. 修复问题 Verify that the reso

54、urces work properly outside of VCS.确认这个资源可以在VCS之外正确运行 Clear fault in VCS.在VCS中清除fault状态,Page 66,Clearing Faults 清除faults,After external problems are fixed:在外部的错误修正后 Clear any faults on nonpersistent resources.清除非持久的资源的错误 hares -clear resource -sys system Check attribute fields for incorrect or missi

55、ng data.检查不正确的配置属性 If service group is partially online: Flush wait states: hagrp -flush service_group -sys system Bring resources offline first before bringing them online.,Page 67,Agent Problems: Agent Not Running代理的问题:代理没有运行,Determine whether the agent for that resource is FAULTED:确认那个代理的资源是否使FAULTED状态的 hastatus summary Use the ps

温馨提示

  • 1. 本站所有资源如无特殊说明,都需要本地电脑安装OFFICE2007和PDF阅读器。图纸软件为CAD,CAXA,PROE,UG,SolidWorks等.压缩文件请下载最新的WinRAR软件解压。
  • 2. 本站的文档不包含任何第三方提供的附件图纸等,如果需要附件,请联系上传者。文件的所有权益归上传用户所有。
  • 3. 本站RAR压缩包中若带图纸,网页内容里面会有图纸预览,若没有图纸预览就没有图纸。
  • 4. 未经权益所有人同意不得将文件中的内容挪作商业或盈利用途。
  • 5. 人人文库网仅提供信息存储空间,仅对用户上传内容的表现方式做保护处理,对用户上传分享的文档内容本身不做任何修改或编辑,并不能对任何下载内容负责。
  • 6. 下载文件中如有侵权或不适当内容,请与我们联系,我们立即纠正。
  • 7. 本站不保证下载资源的准确性、安全性和完整性, 同时也不承担用户因使用这些下载资源对自己和他人造成任何形式的伤害或损失。

最新文档

评论

0/150

提交评论