Automation with Power

Power Business Continuity and Automation

Connect, learn, and share your experiences using the business continuity and automation technologies and practices designed to ensure uninterrupted operations and rapid recovery for workloads running on IBM Power systems. 


#Power
#TechXchangeConferenceLab


#Servers
 View Only
  • 1.  Dead Man Switch

    Posted 04/17/10 11:33 AM

    Originally posted by: SystemAdmin


    Is the RSCT 2.4.12.0 that comes with the latest AIX 5300-11-03-1013 stable ?

    I updated my system AIX 5300-11-03-1013 with HACMP 5.4.1.7. I did a failover test with halt -q on the Node1. Node2 crashed with the error: TS_DMS_EXPIRING_EM (Dead Man Switch) and dump.

    There is no activity on my database cluster. It's is in testing.

    Attached: errpt output
    #PowerHA-(Formerly-known-as-HACMP)-Technical-Forum
    #PowerHAforAIX


  • 2.  Re: Dead Man Switch

    Posted 04/17/10 11:21 PM

    Originally posted by: Casey_B


    There is an rsct apar that sounds applicable.

    IZ66768: SERIAL NETWORK THREAD MONITORING ERROR CAN CAUSE DMS TIMEOUT DURING HACMP FAILOVER

    Upgrade to include at least this APAR...Or to the latest RSCT level.

    Hope this helps,
    Casey
    #PowerHAforAIX
    #PowerHA-(Formerly-known-as-HACMP)-Technical-Forum


  • 3.  Re: Dead Man Switch

    Posted 04/20/10 09:53 AM

    Originally posted by: SystemAdmin


    Thank's Casey B

    I upgraded to RSCT 2.4.12.2
    The problem is solved.
    #PowerHA-(Formerly-known-as-HACMP)-Technical-Forum
    #PowerHAforAIX


  • 4.  Re: Dead Man Switch

    Posted 05/30/10 07:32 AM

    Originally posted by: parsona


    Hi

    Having the same problem. Have the following installed:

    AIX 5.3 ML11 SP3
    PowerHA 5.4 with SP8
    rsct 2.4.13.1

    Cluster has a shared concurrent volume group that is mirrored between two buildings. (AIX mirrorvg).

    If i issue halt -q, failover works but if I physically remove power from one site, i.e server and storage, I get DMS on the second node while it is aquiring the resources.
    Cluster has two disk heartbeats, one in each building.
    Is this normal behaviour.

    Thank for your assistance.
    #PowerHA-(Formerly-known-as-HACMP)-Technical-Forum
    #PowerHAforAIX


  • 5.  Re: Dead Man Switch

    Posted 06/01/10 08:24 AM

    Originally posted by: Casey_B


    Hello,

    The problem described above is very specific. It only occurs during a node down, and at specific
    levels of RSCT.

    Thanks,
    Casey
    #PowerHAforAIX
    #PowerHA-(Formerly-known-as-HACMP)-Technical-Forum