This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

how to find the dabort occur reason in TMS570LS3137?

Other Parts Discussed in Thread: DP83640, TMS570LS3137

I've encounterd dabort errors frequently. I want to know how to make sure the place where the dabort error happens.

Now, I encountered two dabort errors:

1. I made a board with myself, maybe the hardareware is not so much stable, the board maybe not receive the packet from the network.When I initialize the board's EMAC function , I try to set a flag which indicate the board can receive the data packet from the network(the receive program can be entered), when the board cannnot receive the packets at first, then reinitialize the EMAC function(set the DP83640 again and creat the udp 'pcb' with the api of the lwip ). At this time,the dabort occurs in the etharp_send_ip() function in etharp.c file.

2.When the board can send and receive packets successfully, I send UDP packets frome the PC to the board every 100ms, after several minutes later, the dabort error happens, I don't know the reason led to that. How to qurify the place or variable that error happens and then fix it?

The attach is the project file I used.

Regards,

yong

  • Yong,

    I don't have access to a board at the moment and didn't check your .rar.

    But you can easily find the address of the instruction that caused the data abort.

    There are coprocessor registers in CP15 that record the address, and you can easily inspect these in CCS.

    Take a look at the section "Fault Status and Address Registers" in the Cortex R4 CPU Technical reference manual (From ARM's website).

    The contents of this register should help you zoom in on where in your program you are accessing the address that causes the abort.   This could be a pointer problem or something simple.

     

  • Hi, Anthony

    I just try what you said,but seems still cannot find the reason about the second dabort mentioned above,the snapshot of the picture is below

    the data fault address recorded by the CP15 register is 0xEE03A66C(the CP15_DATA_FAULT_ADDRESS register value), where it said that memory map prevented reading of target memory at  0xEE03A66C,while the CP15_DATA_FAULT_STATUS value is 0x00000808.

    I look the Cortex-R4 CPU Technical Reference Manual, it said that the R14 records the last position,while here the R14 value is 0x00013E2C, the program here is the " _esmCcmErrorsClear_" , I don't know how the program can enter into the sys_core.asm file.

    I set a breakpoint where the _esmCcmErrorsClear_() function is called in the sys_startup.c(the program already run to the main() function ),but the program never stopped here, I still  cannot find the reason led the program dabort.

    Hope you can help!

    Regards,

    yong

  • I put  the lwIPRxIntHandler() back to the EMACCore0RxIsr() function as an interrupt service function.

    I noticed that when the lwip receive the data in the callback function(locatorReceive()), the orginal stack is not big enough to send the data(the orginal is 0x100, from the 0x08001300 to 0x08001200 of the IRQ stack ),then the data defined in the main would be changed by the Enthernet receive interrupt function(the IRQ stack is not big enough so it will change the FIQ and the Supervisor even to the User stack), I expanded the IRQ satck to 0x300(from 0x08001500 to 0x08001200), the program can work.But when the program receive data continuously, the SP move to the Abort section and to the Undefined at last(SP is 0x08001700,just as defined in the HCG),I don't know where the SP changed.

    At what circumstance, the SP would move to the Abort section and the Undefined section?How to qurify the place that the program made it?

    Regards,

    yong

  • Hi Yong,

    Regarding

    yong zhang2 said:
    Cortex-R4 CPU Technical Reference Manual, it said that the R14 records the last position,while here the R14 value is 0x00013E2C, the program here is the " _esmCcmErrorsClear_" , I don't know how the program can enter into the sys_core.asm file.

    You would need to trap at the exception vector (put hardware breakpoint at address 0x0C) to catch the address in R14.  Once the CPU starts calling other functions, it will reuse R14 and so the real 'answer' will be on the stack somewhere (hopefully).   

    R14 is probably fine to start with.  There's some other note in the ARM Cortex R4 TRM that discourages it as not being precise and tells you to do the CP15 register instead but it's probably good enough for now.

    yong zhang2 said:
    At what circumstance, the SP would move to the Abort section and the Undefined section?

    It sounds like you might have found your problem but you're saying just making the stack size larger isn't robust for continual transmit and receive, is that right?    And by "SP would move to the Abort section" I think you just mean that the IRQ stack grew so large that it started writing over the stack area of Abort mode, is that what you mean.  

    So the real question is 'how bit do I need to make the IRQ stack' so that there are no more overflows - correct?

    If so I don't know the answer to that question.  I think this is where another person's comment on the forum about the lwIP demo running the IP processing in IRQ mode might be applicable.   The 'right' answer is probably to move the processing back to the task level, and from there if you have an RTOS running you could use the RTOS's tools to check task stack usage and even perhaps set the MPU up so that it would catch a stack overflow from that particular task. 

    Otherwise, you'd need to analyze the way that the lwIP code uses the stack to figure out how big the stack needs to be.   I don't know enough about that code to tell you a good answer, unfortunately.   And it would take more work & a deeper understanding of the lwIP code than I'm capable of at the moment.

     

  • Hi, Anthony:

    Glad to see your reply, the SP in the TMS570LS3137 will minus from the top(for example, the IRQ stack is from 0x08001200-0x08001500, the SP will start from 1500, then minus proper value when needed until 0x08001200,if there is still some space needed, it will enter to the FIQ stack(0x08001100-0x08001200)), while the dabort stack address is 0x08001500-0x08001600,  I don't know why  the SP move there?

    I thought that the TMS570LS3137 may have some internal rules that when a dabort exception happens, the SP will move to the dabort area and if undefined entry exception happens, the undefined stack entered. Is that correct?

    Regards,

    yong

  • Hi Yong,

    Ok, I think it's two separate things we're talking about.

    First, if the stack pointer for the IRQ stack keeps decrementing, such that it decrements past it's bottom address of 0x08001200 and into the FIQ stack space, this is a stack overflow and it's a common source of abort type problems, because variables on the stack can be corrupted during the overflow and when they're used again they might be 'pointing' to unmapped memory.   So this might be a *cause* of the problem - the stack decrementing past it's limit.   To solve that problem you need to know how much stack is needed in the worst case which isn't always an easy thing to determine. Then you can either reduce this # by changing the structure of the code or increase the space allocated to the IRQ stack.

    The 2nd thing is when the SP switches from one area to another by mode change - not by simply decrementing.   This happens because there are multiple SP registers,  the exception modes have their own copy.  So when you enter an exception mode the effective SP changes;  but the value from the user-mode or system mode SP still is available, it will be there again when you switch back to user/system mode.   There is a table in the Cortex R4 TRM that explains this.  I'll paste an image here - I'd send a link but I haven't figured out how to directly link to a particular section of ARM's online documentation.   You can probably find it by the section # though.

     

  • Thanks a lot for your detailed explanation! I've learned a lot.

    Best Regards,

    yong