This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

CCS/LAUNCHXL-CC2640R2: Multi_role very weird error with devList[]

Part Number: LAUNCHXL-CC2640R2

Tool/software: Code Composer Studio

Hello,

I have this weird error, I think there might be some SDK error (can't explain it any other way).

I am using 2 LaunchXL CC2640r2f, my code is based on multi_role.

1st device scans, 2nd device advertises. I don't use devList[]:

1st device finds 2nd device in GAP_DEVICE_INFO_EVENT, and then connects to the 2nd. My svc is found, devices paired and bond save success. 2nd device do GATT_write to 1st device.

Now 2nd device gets stuck, unless I keep this exact line in the code:

and another line using devList[], for instance:

Even though these 2 lines do nothing (there's no other place using devList[] in the code).

With these 2 lines everything works fine. If I delete it, the code crashes right after 2nd device tries to do GATT_write to 1st device.

Any idea what could be the problem?

I am using sdk_3_20_00_21, compiler version 18.12.3.LTS.

Thanks,

Amit

  • Hi Amit, 

    Can you reproduce this issue using the default and unmodified multi_role project?

    Thanks,
    Elin

  • Amit,

    It sounds like you have modified the sample app a bit to remove devList, perhaps there is some lingering affects from the code not being completely removed.

    You mention that the code crashes right after doing a GATT write, Are you using mr_doGattRw to achieve this write?

    Can you follow the steps in this guide (subsection = Deciphering CPU Exceptions) to find the stack trace at the time of the crash

    http://dev.ti.com/tirex/explore/node?node=ANPdcnhbDh6Eu5rLf2U9uA__krol.2c__LATEST

  • Hey Elin,

    Of course that is not possible, since if I'd removed devList[] I'd have to modify:

    multi_role_processRoleEvent()

    multi_role_addDeviceInfo()

    mr_doConnect()

  • Hey Sean,

    I have modified it quite a lot. Everything worked and my code didn't use devList[], but when I removed it the code got stuck. Took me a while to figure that's the cause (which is very odd since, as mentioned, my code doesn't even use it!).

    I've written a function similar to mr_doGattRw (though I in my case req.len = 4 instead of 1 byte). As mentioned before, it worked.

    One more thing, the return value of GATT_WriteCharValue() is SUCCESS but it still stuck right after the send.

    I've entered the link but I don't understand what to do exactly.. I only use CCS, not sure what's RTOS Object View etc. Been doing debugs with prints to UART. Can you please further explain what to do?

    Thanks,

    Amit

  • Amit,

    Stuck usually means crashing.

    1. Attach debugger

    2. Reproduce issue w/ debugger attached and debug view open in CCS

    3. Open ROV using Tools--> Runtime Object View.

    4. With ROV open, follow the steps above to determine the exception type and stack trace

  • Hey Sean,

    I've opened ROV.

    What file is the TI-RTOS configuration file that relates to M3Hwi?

    Thanks,

    Amit

  • Amit,

    The entirety of TI-RTOS (including m3Hwi) is configured by app_ble.cfg.

    Before getting into modifying that. Can you you share a capture of tools --> ROV --> HWI --> exception

    Here is a video on how to use ROV

    https://www.youtube.com/watch?v=MI_2iM2WbU8

  • Hey Sean,

    Hwi exception: 

    Error: java.lang.RuntimeException: Target is running

     

    To the left you can see UART log (putty). Gets stuck in the middle of the print, right after GATT_WriteCharValue() is done:

    Still have no idea what's wrong..

    Thanks,

    Amit

  • Amit,

    Please hit the pause button.

  • Hey Sean,

    Thanks,

    Amit

  • Hello,

    Anything new about this?

    Thanks,

    Amit

  • Amit,

    You have an imprecise error, force it to precise as detailed by:

    http://dev.ti.com/tirex/content/simplelink_cc2640r2_sdk_3_30_00_20/docs/blestack/ble_user_guide/html/ble-stack-3.x-guide/debugging-index.html#exception-cause

     


    Then attach to debugger and re-run the experiment.

  • Hey,

    Is this ok:

    Thanks,

    Amit

  • Hey Amit,

    Something in the code is attempting to write to address 0x01. This is in Flash (thus not writable) and is causing a bus fault.

    I recommend setting a hardware watch point on address 0x01 and seeing what code is writing to that address.

  • Hey Sean,

    How do I set a hardware watch point? I am sorry, using this debugger is kind of new to me.

    Thanks,

    Amit

  • Hello,

    I'm not sure I've done this right. I've added this to the expressions:

    "(uint8_t*)0x01"

    But then I got data access error in some other address 0x3475:

    So I've added another expression like this:

    Then I got a different address again:

    So I'm not sure I've added the hardware watch point right. Can you please tell me whats wrong?

    And again I remind you that if only I keep those two useless lines in my code:

    It works without errors and doesn't get stuck.

    I have no idea whats going on..

    Amit

  • Hi Amit,

    We have a guide on how to set watchpoints in CCS here:

    http://dev.ti.com/tirex/content/simplelink_cc2640r2_sdk_3_30_00_20/docs/blestack/ble_user_guide/html/ble-stack-3.x-guide/debugging-index.html#watchpoints-in-ccs

    I have linked it for your benefit. However, since as you have observed the address is changing we need to do a bit more clever debugging.

    Here is what I can summarize:

    1. You are reaching a hard fault because the device is attempting to write to flash memory

    2. The address at which the attempted write occurs is random

    3. The presence of the "useless" devList structure in memory is likely masking the problem because it is changing the layout of variables in SRAM

    Can you attempt the following:

    1. Reproduced the issue, pause the debugger, open ROV, select BIOS -> scan for errors and paste the results on this thread

    2. Try to reduce the compiler optimizations of the project (O3 for example) or the application file only (e.g. multi_role.c) to see if you can reproduce the problem more repeatedly. Compiler optimizations changes the way memory is indexed and addressed and may make the issue easier to see/debug. Steps below

    http://dev.ti.com/tirex/content/simplelink_cc2640r2_sdk_3_30_00_20/docs/blestack/ble_user_guide/html/ble-stack-3.x-guide/debugging-index.html#optimizations

  • Hey Sean,

    I've tried both ideas and here's what I've found.

    Note - the error also occurs if I only comment this line

    Leaving devList[] definition not commented:

    So unless the compiler removes devList[] for no usage, it should be in the memory. And of course I still get the same error.

    To the attempts:

    1. Here's what the BIOS scan for errors produced:

    2. I've reduced optimization level to 3 - Interprocedure Optimizations and let it run again, here's what it produced:

    Still clueless...

    Thanks,

    Amit

  • Hi Amit,

    This almost certainly confirms the case of memory corruption. The ROV cannot read the internal structures related to the kernel because they have been written unexpectedly by some other process or pointer. Which is why the ROV is unable to show information correctly.

    Memory corruption usually happens in the following conditions:

    1. Heap issues e.g. FREEing a pointer twice or using a pointer after it has been free'd
    2. Size issues, an API is called with an improper length causing an operation to touch memory it shouldn't 

    Since the bad operation (triggering the exception) is happening at a random address, the watch point approach won't work.

    However, you mentioend that GATT write is triggering the exception.

    1. Do you ever encounter an exception before calling GATT write?

    2.  Does the exception occur 100% of the time after executing the GATT write?

    3. If the GATT write is removed, is there still an exception?

  • Hey Sean,

    1. How can I check for exceptions before the GATT write, before it collapses?

    2. As long as remove devList[], yes. The exception happens 100% of the time. When devList[] stays, it doesn't happen at all.

    3. If GATT write is removed, there is no collapse. Even when I remove devList[].

    That is why I believe there's something wrong in the GATT or SDK, parts of the code I can't reach...

    Thanks,

    Amit

  • Hi Amit,

    If you set a breakpoint on GATT_write, and are able to run to the breakpoint without issue, then there are no crashes before it.

    If stepping over the GATT write with the debugger causes a hang in an exception, then likely the issue is with the write.

    Can you confirm this?

  • Additionally, can you post all the modifications to the code surrounding the GATT_WriteCharValue and mr_doGattRw that you have made?

  • Hey,

    Confirmed:

    No errors before GATT_WriteCharValue.

    Amit

  • Hi Amit,

    Thanks for the test

    Your screenshot reveals a potential issue with the code:

    In line 213 you allocate space for 1 byte inside req.pValue, then you pass that value to the stack with req.len = 4; (line207). 

    Also in line 208 there is a memcpy that will write 4B into a pointer that only has 1 byte allocated for it which will corrupt the structure of the heap.

    In summary, if you wish to write 4 bytes in GATT then you should modify the GATT_bm_alloc to allocate 4B instead of 1.

    req.pValue = GATT_bm_alloc(conn_handle, ATT_WRITE_REQ, 4, NULL);

    The 3rd parameter is of interest here which controls the size of the memory to allocate.

  • Hey Sean,

    For a minute I thought this was it, fixed this mistake and changed the size to 4 but it still gets stuck in the exact same manner :/

    So close... but it still seems to me like something in GATT_WriteCharValue.

    Just saw your previous reply about modifications surrounding the GATT_WriteCharValue and in mr_doGattRw.

    This function makes the writes to a BT profile similar to simple_profile:

    This function replaces mr_doGattRw().

    Everything worked fine until I've deleted devList[] since I didn't use it.

    Amit

  • Amit,

    The code snippet you shared looks OK.

    Can you share the following:

    1. The changes you made to the profile related to changing this characteristic from 1  byte to 4
    2. The encodeMsg() and printDataToDisplay functions()
    3. Can you validate that task_id and char_handle are correct?
      1. Task ID should be identical to return of ICall_getEntityId()
      2. char_handle should be the same as if you discover the peer's service with BTool or LightBlue
    4. Can you comment out encodeMsg() and printDataToDisplay() and try to send a hardcoded value such as codedMsg = 0xAABBCCDD, does this work?
    5. You pass the entire msg struct (Message type) by value to station_sendCodedMsg32(), do you know if there is enough runtime stack space for this? what is the size of the Message struct?

  • Hey Sean,

    1. Changes were made in ReadAttrCB():

    WriteAttrCB():

    2. encodeMsg(): simply put msg info in a 32 bit uint (see Message struct below):

    printDataToDisplay():

    3.Task_id passed to the send function:

    as defined in multi_role. I've passed it as an argument to the function since I've put the function in another file.

    Handle:

    This is a call to the send function. Handles are ok. Value is 0, which is the first index in connHandleMap. Haven't changed that mechanism yet (I plan to).

    4. Tried that, same error still happens.

    By the way, either way the message never reaches the other unit, which terminates connection after the sending unit gets stuck.

    Also, don't forget that everything works fine when devList[] is not commented out...

    5. Msg struct:

    size is 32 bits. I don't see any problem with runtime stack space, as it works good with devList[] not commented.

    Thanks,

    Amit

  • Hello,

    Is there anything new?

    Amit

  • Hey Amit,

    Apologies for delay in repsonse, most of TI was on vacation last two weeks for holidays.

    Anyway the code snippets you pasted look okay from a first glance.

    It's very clear that when devList is commented out the issue doesn't occur, however, this is most likely just hiding a memory corruption issue that still exists in your application, just instead occurring at different memory.

    >>I don't see any problem with runtime stack space, as it works good with devList[] not commented.

    I don't think this justifies that the stacks aren't overrun, can you check them using ROV before the error happens?

    Can you share the place the you actually define the stationProfileCodedMsg variable? How about where you pass it into the profile (e.g. gattAttribute_t [])

  • Hey Sean,

    Hope you had a good vacation.

    This is the definition of stationProfileCodedMsg in station_profile.c:

    This is codedMsg definitions in stationProfileAttrTbl[]:

    How do I check the stacks aren't overrun with ROV?

    Thanks,

    Amit

  • Anyone?

    Its been over a month and I got no answer. Is there anything new?

    Thanks,

    Amit

  • Hi Amit, 

    I have looked through the thread and it looks like you have received a lot of help from Sean. 

    Please refer to the User's Guide for help on how to use the ROV. There is also a video that Sean has linked in a previous post.

    Thanks, 
    Elin