This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

Can Watchdog interrupt be used if system is hanged in ISR

Hi ,

I am using watchdog to know the stack size and all if system hangs.

For that I tried an example of keeping while(1); in ISR . So if after watchdog reaches it timeout I am reseting the device .. this works fine.

But I don't want to reset the device and I need the an LED to glow if watchdog timesout. So I am using watchdog to generate NMI and I have an ISR to glow LED. But I didn't obsered any LED glow.

Can anyone please help me whether I am wrong in using watchdog interrupts properly?

Thanks

Sravan

  • Hello Sravan,

    Can you attach the code that shows how you are configuring the watchdog and how is the NMI Handler set up?

    Regards
    Amit
  • Hi Amit ,

    Please go through the following snippet of code:

    void WD_Timer0_Init(void)
    {

    Error_Block eb;
    Hwi_Params hwiParams_1;

    Error_init(&eb);
    Hwi_Params_init(&hwiParams_1);
    Hwi_construct(&(wd_timer), INT_WATCHDOG_TM4C129,
    watchdog_isr, &hwiParams_1, &eb);
    hwiParams_1.priority = 0;
    if (Error_check(&eb))
    {
    Display(("Eror in Initialization of Watchdog.\n"));
    }
    // Check whether the Watchdog timer0 is locked
    if(WatchdogLockState(WATCHDOG0_BASE) == TRUE)
    {
    // Unlock Watchdog timer0 if it is locked
    WatchdogUnlock(WATCHDOG0_BASE);
    }

    WatchdogReloadSet(WATCHDOG0_BASE, WD_TIMEOUT_5S);
    WatchdogIntEnable(WATCHDOG0_BASE);
    WatchdogIntTypeSet(WATCHDOG0_BASE, WATCHDOG_INT_TYPE_NMI);

    // Enable the WD timer0
    WatchdogEnable(WATCHDOG0_BASE);

    }

    and

    void watchdog_isr(UArg dummy)
    {

    WatchdogIntClear(WATCHDOG0_BASE);
    GPIOPinWrite(LED0_BASE,(LED_WHITE|LED_RED),(LED_RED));
    printf("\n3\n");
    }


    And in the following I have used to generate watchdog interrupt.

    void Switch(void)
    {
    while(1);
    }
    This above " switch " function is an another GPIO interrupt.


    Summary:

    I am using switch (GPIO) interrupt to wait continuously in the ISR. So watch dog will come into action and LED has to blink .

    Regards
    Sravan
  • Hi,

    If you use CCS, then instead watchdog you may use the  __stack variable added to your watch window to monitor the stack. You will get interrupted automatically if your stack overwrites this limit.

    (take care , that variable is with two underscores).

    Petrei

  • While one can certainly do that, I would strongly dicourage you, since it is actually bad design.
    When your application hangs in an interrupt service routine, or runs into a stack overflow, you have an application bug. Just resetting the system is really bad idea. A watchdog is really the LAST line of defense, to catch unexpected problems (e.g. possibly acceptable to recover from EMI issues (static discharges, mains surges).
    Imagine your smartphone, internet router, car, alarm clock (etc.) would just restart on the slightest problem. No customer would accept this behavior. I would rather focus on fixing bugs, not circumventing them.
  • Hi f.m ,
    Yes what you said is true.
    Actually my application is rebooting the device after certain time with the help of watchdog timer.So I want to nail down this issue.
    Instead of rebooting device I want this watchdog interrupt to print the stack size of all the tasks that I am using in my application.

    For testing I am going to creating this issue and then I want to display the stack size and the other registers value () etc., to find the actuall problem.

    At first I am trying to test the implementaion of watchdog and displaying stack sizes in those case was true or not. That why I want to do a dummy test before actually integrating it to my application
  • Hi Petrei,

    Thanks for commenting
    Actually I don't know the root cause and I suspect that stack overflow ..

    And there may be any other reason I have to find it.. I am using this watchdog so if my system hangs , then I wil display some register values and stacksize ..


    Regards
    Sravan
  • Hi,

    You claim the purpose to trigger the NMI and you configured for that. But did you replaced the NMI interrupt vector with that of your watchdog_isr? Take care any interrupt vector should be declared as void my_int(void), without any parameters.

    Second, since you mention/use the word "task" do you use any form of RTOS? Because in these cases a common source of errors is the limited heap, too abused maloc like functions can trigger some hard to find problems...

    Petrei

  • The watchdog (the implementations I know) do not create an interrupt, but execute an immediate reset of the MCU.
    So, activating a watchdog is not really supposed to help you.

    Instead, I would instrument the code, either with debug messages over any appropriate channel (most often UART), or, in the simplest case, GPIO toggles, which could be watched with the scope. Then, you can start locate your problem. This usually takes several turns (remove instrumentation code on non-affected locations, add new one near the "fault line").
    Keep in mind, however, that this method is not un-intrusive. Especially debug messages could aggravate looming stacksize issues.
  • Hi Petrei,

    Yes I tried both with NMI and regular interrupt, in both case it is not served.

    Yes I am using multiple tasks. And the problem is when system is in idle task i.e., all the task are in suspended state... the sustem reboots automatically by itself. It is happening at random and not always but very frequently.

    So I am not sure whether it is heap or stack or something else.

    Thanks
    Sravan
  • Hi f.m. ,

    Sorry to mention this . If I use any debug messge to print then I am not observing any reboot. But if those are commented then it reboots .

    My main reason to use watchdog is not to reset ..

    When this unknown hang occurs watchdog comes into action . If we use reset it will reset the device otherwise it has to serve what ever code I write in ISR . Isn't that the behavoiur ??

    Please comment if I am wrong.

    I am using watchdog as one of my way to find the issue.If you have any other ideas just motivate me

    Thanks
    Sravan
  • Hello All,

    The way it is NMI or Interrupt is handled for the watchdog in TI RTOS is different. While an Interrupt can be logged, the NMI which resides in 0-15 of the CPU Interrupt vector has a quirk when do RTOS.

    Sravan,

    Please check on the TI-RTOS forum for NMI in Tiva. (while I do a parallel check)

    Regards
    Amit
  • If I use any debug messge to print then I am not observing any reboot. But if those are commented then it reboots .

    That called a Heisenberg bug, AFAIK. In this case, it might be worth trying to reduce the debug output size (send only single characters, for instance), or try the GPIO toggle method.

    But if additional load (debugging output) alleviates your problem, it's probably not stack related. I would guess it is a timing issue, possibly in connection to interrupts.

    Do you catch all interrupts, i.e. resetting all interrupt flags ? If, for instance, you get an unexpected overflow interrupt once in a while, and don't reset the corresponding flag, your code cycles endlessly in that handler routine.

    I am using watchdog as one of my way to find the issue.

    You can try to use the watchdog as debug method, that's of course no problem. I'm just not conviced it is the best method. Actually, I never used the TM4C watchdog, so I don't know it's capabilities, and what the interrupt is good for. I know from other vendor's MCUs, which allow to configure an interrupt shortly before the watchdog times out. However, there is a good chance your application is already hopelessly lost at that time ...

  • Just a thought...

    You say that the watchdog is causing the reset. Are you sure that is what caused the reset? Would it still reset if you completely disable the watchdog?

    Just something to try in case a wrong suspicion is leading you the wrong path.
  • A few comments on your approach Sravan, I've not looked in detail at your code.

    First using a watchdog is not a bad idea for debugging such Heisenbugs, I've used similar appproaches in the past myself. There are consequences you need to consider though.

    • First using an NMI is dangerous.  There is also little need for it on a system with programmable interrupt priorities (IMHO there is seldom need for an NMI and it's presence should always be viewed with a jaundiced eye).  Rather than NMI I'd consider using the highest priority interrupt.I suspect, though, that the use of NMI is not causing a problem.
    • Second you should not be using any of the printf family of functions in any interrupt. They will have both re-entrancy and stack issues.
    • Third, since you suspect stack you should avoid doing any processing that affects the stack or is dependent on the stack as far as you can.  Another reason to avoid the printf family.

    What I would do in this case is reserve an area of memory for forensic information.  Whenever the processor boots print this information.  Note that this means you will have to reserve an area so that the compiler will not initialize it.  When you get your watchdog interrupt copy the information (Such as stack pointers) to the reserved are and force a reboot.  Hopefully you will not have modified the behaviour of the program and the problem still occurs.

    This even has a chance of working in the presence of a problematic NMI.  And since the printing now occurs after restoring the system to a good state you have less to worry about in terms of overflowing stacks and other corrupt internals.

    If you do not do so already, one of the things you should print is the reset source.  It is very easy to think you have a watchdog issue when you have an infrequent power issue and vice-versa.

    Robert

  • f. m. said:
    But if additional load (debugging output) alleviates your problem, it's probably not stack related. I would guess it is a timing issue, possibly in connection to interrupts.

    Good point.

    As well as races, resource deadlock could be an issue in a tasking environment.

    Robert

  • Hi Quark,

    No actually system hangs after some time . I used watchdog to recover so it is resetting after that unknown hang
  • Hi f.m,

    Even I printing single character through UART , but it is not happening. I want to make this interrupt has high priority than any other..

    I have a doubt :

    When something hangs due to stack over flow and all.. will UART or other API's will work ..??
  • Hi Robert,

    Exactly I too want to do this... But I need some help in making watchdog interrupt as high priority.

    First I am verifying will watchdog serves purpose in all types of faults , For this as a trail method I am keeping an infinite loop in one ISR and hoping for watchdog interrupt to serve.

    But I don't know how to test ...And I am litting one LED and also I am printing on console. I agree you that if it is a stack issue "printf " won't work but atleast LED has to glow right??

    In my case it isn't

    Please can you help me that making " priority = 0x00 " for watchdog regular interrupt will serve this purpose .


    Thank you.