TMS320C6657: DSP Rom Bootloader (RBL) NAND Bootloader support ECC?

Part Number: TMS320C6657

We are currently seeing some issues where there is "bit rot" of the intermediate bootloader (IBL) programmed into a NAND flash device attached to the DSP.   From the KeyStone Architecture DSP Bootloader User Guide  section 3.11 it seems to imply that the RBL should be doing some ECC but we are not seeing that.

In our investigation (pulling the data from the NAND device) we see that one bit has been flipped (corrupted) in the intermediate bootloder image residing in teh NAND device.  This one bit corruption prevents the IBL from running and our sytem from booting.

So the question is, "does the RBL strapped to NAND boot actually run ECC?"   

  • Hello,
    Thank you for your inquiry. **x1125363@ti.com** is currently on leave and will be available on **2026-08-28**.
    Your query will be addressed upon their return. We appreciate your patience and understanding.
    Best regards,
    TI E2E Support Team
    ---
    *This is an automated notification.*

  • does the RBL strapped to NAND boot actually run ECC?

    Hi Tom,
    From my understanding RBL should run ECC. Could you please refer the thread "TMS320C6678: NAND Error Management (ECC - EMIF interface) - Processors forum - Processors - TI E2E support forums" regarding timing issue to perform ECC in nand configuration.

    Regards,

    Ben

  • Hello,

    The issue is more focused on the ROM bootloader that is strapped for direct NAND boot.  The ECC in teh RBL does not appear to be able to correct a single bit error when booting from the the NAND.

    We were looking through the ROM Bootloader source code found at the following location: https://software-dl.ti.com/sdoemb/sdoemb_public_sw/rbl/1_0_C6657/index_FDS.html. Could you confirm whether this is indeed the code that the RBL is running during NAND boot? If so, then we have a question around lines 296-316 of BootROM_c6657_PG1.0\C665x_bootROM_src\hw\nand\device_nand.c. (see image below).

    Specifically, if an error is detected on line 296, then bit [13] of the NANDFCR register is written (line 304) so that the addresses and values can be computed. Then there is a do-while loop (lines 308-313) that waits until NANDFSR[11:10] are non-zero or there is a timeout. Assuming there isn’t a timeout, then these two bits of NANDFSR (which are the two MSBs of CORR_STATE[3:0]) will become non-zero almost immediately, because these states correspond to when the hardware is performing the actual error search, calculating the addresses and values, etc. Then, what is assumed to be the final correction state is read from NANDFSR (line 316) and if it’s either equal to 1 (errors cannot be corrected) or greater than 3 (searching/calculating), then the function returns a fault. We have seen it take up to seven reads of the NANDFSR register before a final CORR_STATE (such as 0x1, 0x2 or 0x3) is actually achieved.

     Based on the code comment (“Loop until timeout or the ECC calculations are complete (bit 11:10 == 00b”), it seems like the condition on line 313 should read “while ( (i > 0) && (temp == 0) )”, not “while ( (i > 0) && (temp != 0) )”. Could you confirm whether the released RBL code has a different do-while loop condition.

     SNIP FROM "device_nand.c" (lines 271 to 330):

    --------------------------------------------------------------------------------------------------------------------------------

    UINT32 DEVICE_NAND_ECC_correct(UINT8 *data, UINT8 *readECC)
    {
    VUINT32 temp, corrState, numE;
    UINT32 i, j;
    UINT16 addOffset, corrValue;
    UINT16 syndrome10[8];

    /* Convert the byte array to a 16 bit array. The data arrived in little
    * endian order */
    for (i = j = 0; i < 8; i++, j += 2)
    syndrome10[i] = (readECC[j+1] << 8) | readECC[j];

    // Clear bit13 of NANDFCR
    temp = AEMIF->NANDERRADD1;

    // Load the syndrome10 (from 7 to 0) values
    for(i=8;i>0;i--)
    {
    AEMIF->NAND4BITECCLOAD = (syndrome10[i-1] & 0x000003FF);
    }

    // Read the EMIF status and version (dummy call)
    temp = AEMIF->ERCSR;

    // Check if error is detected
    temp = (AEMIF->NAND4BITECC1 & 0x03FF03FF) | (AEMIF->NAND4BITECC2 & 0x03FF03FF) |   <--- line 296
    (AEMIF->NAND4BITECC3 & 0x03FF03FF) | (AEMIF->NAND4BITECC4 & 0x03FF03FF);
    if(temp == 0)
    {
    return NAND_RET_OK;
    }

    // Start calcuating the correction addresses and values
    AEMIF->NANDFCR |= (0x1U << DEVICE_EMIF_NANDFCR_4BITECC_ADD_CALC_START_SHIFT);   <--- line 304

    // Loop until timeout or the ECC calculations are complete (bit 11:10 == 00b)
    i = NAND_TIMEOUT;
    do
    {
    temp = (AEMIF->NANDFSR & DEVICE_EMIF_NANDFSR_ECC_STATE_MASK)>>10;
    i--;
    }
    while((i>0) && (temp != 0x0));  <--- line 313

    // Read final correction state (should be 0x0, 0x1, 0x2, or 0x3)
    corrState = (AEMIF->NANDFSR & DEVICE_EMIF_NANDFSR_ECC_STATE_MASK) >> DEVICE_EMIF_NANDFSR_ECC_STATE_SHIFT; <--- line 316

    if ((corrState == 1) || (corrState > 3))
    {
    temp = AEMIF->NANDERRVAL1;
    return NAND_BOOT_ERROR_ECC_CORRECT_STATE;
    }
    else if (corrState == 0)
    {
    return NAND_RET_OK;
    }
    else
    {
    // Error detected and address calculated
    // Number of errors corrected 17:16

    --------------------------------------------------------------------------------------

  • Hi Tom, 

    I have to check this with the team and will provide you a response within a day.

    Regards,

    Ben

  • Hi Ben,

     

    Great, thank you. In the interim, we have one clarification, plus some additional information that should be useful. The actual issue with the do-while loop (lines 308-313) is that there isn’t enough time after writing to the NANDFCR register before entering into the loop, such that the first read of the NANDFSR occurs before the calculation of the correction addresses and values has even begun. So, the CORR_STATE[11:10] bits have not yet become non-zero when entering into the loop, which causes the loop to break out immediately. This is where the additional seven reads of the NANDFSR register come into play before a final state of 0x1, 0x2 or 0x3 are reached. This appears to be the very issue that you originally linked us to (https://e2e.ti.com/support/processors-group/processors/f/processors-forum/1087529/tms320c6678-nand-error-management-ecc---emif-interface/4030371?tisearch=e2e-sitesearch&keymatch=EMIF25_FLASH_CTL_REG), but we’re encountering it in the RBL code instead of the IBL code.

     

    This race condition leads to bit errors going uncorrected. For example, we have a NAND flash with two known bit errors, one at block/page/segment 0/7/2 and the other at 0/12/1. Knowing what spare bytes will be read for these two segments allows us to connect a debugger and set a hardware watchpoint while the RBL is running. We placed watchpoints in order to halt the processor on the write of the spare (OOB) bytes into the NANDF4BECCLR register (line 289), then repeatedly (about 38 times) click the “Assembly Step Into” button through Code Composer until we see the NANDFCR register written (line 304). Then we step through the assembly code another ~10 times in order to get on the other side of the do-while loop. We repeat this process for the second known NAND bit error, then simply hit “Resume” and the RBL successfully hands over control to the IBL that has been corrected/copied into RAM.

     

    It appears that we need to implement the same sort of delay mentioned for the IBL, but somehow in the RBL. Is this truly ROM code, or is there some way that it can be updated? Or perhaps there’s something that we could change in our boot table parameters that would introduce a delay during this interaction with the ECC hardware?

     

    Thanks

  • Hi Tom,

    The source code you referenced is published by TI as developer reference code used for the reference for customers to implement on custom boards and would not be the exact code of the boot rom loader. The updated version of the reference is provided in this link Keystone Device Architecture - Texas Instruments Wiki.
    On the boot parameter table: the NAND Mode Boot Parameter Table is documented in the C6655/C6657 data manual, Section 6.28.3.4 (Table 6-89): https://www.ti.com/lit/ds/symlink/tms320c6657.pdf

    Looking at the full table, there's no field that reaches the ECC timing it's limited to I2C bus parameters, chip-select region, and first block number. So this table can't be used to introduce the delay you're looking for.

    Regards,
    Ben

  • So then why is the RBL not running properly and correcting a single bit error detected in the NAND flash?  We have proven that if you pause the the RBL when it detects the error and allow for the ECC calculations to complete then resume the RBL running it can correct the error.

    Is there a timing bug in the RBL running on the DSP?   (the RBL continues without waiting on the ECC calc to finish) 

    Is there anyway we can update the RBL in the DSP?

  • Is there a timing bug in the RBL running on the DSP?   (the RBL continues without waiting on the ECC calc to finish) 

    I have reviewed the errata document and found no reported issues related to RBL timing.

  • We did not see any errata against this issue either ...  but we are still seeing the RBL during direct NAND boot still fails to correct a single bit error when it copies the IBL to RAM and thus the DSP fails to boot.

    ????

  • Hi Tom,

    Is there anyway we can update the RBL in the DSP?

    We cannot modify the RBL source code, as the ROM Boot Loader (RBL) is programmed into ROM during device manufacturing and is not modifiable, as noted in document provided here:  Slide 1.

    Regards,

    Ben

  • So is there any specific strapping or other configurations settings or file formatting that we should check/verify that might influence the ability of RBL's EMIF interface to run the ECC during a direct NAND boot?  If the NAND image is error free the RBL and our IBL run flawlessly.

  • Hi Tom,

    As the issue appears to be with the IBL boot image, have you tried rebuilding the boot image and checking whether the issue still persists.

    Additionally, could you confirm whether the boot image was generated by following the steps outlined in the FAQ below.

     [FAQ] PROCESSOR-SDK-C665X: IBL Build Steps on Linux (Ubuntu) (Processor family: C665x /C667x) - Pro…. 

    Regards,
    Ben

  • Our image is fine.

    When there is not bit error (from bit rot) in the NAND the RBL successfully copies the image in the NAND (RBL=Direct NAND boot) to RAM and the image starts running successfully on the DSP. No issues.

    The problem is when there is a "bit rot" bit error in the image data located in the NAND, in this case the RBL "fails" to successfully copy the image from NAND to RAM (the RBL does not do ECC), thus the DSP is not running.

    Is there a known bug in the RBL boot code for "Direct NAND boot" that doesn't allow the RBL code to run ECC?  

  • Hi Tom,

    I couldn't find any reported bugs related to this.


    Regards,
    Ben

  • So is it TI's recommended solution to NOT use "Direct NAND boot"  hardware strapping option for the RBL to load the DSP image?  Since the DSP's RBL code is not capable of correcting (ECC) a single bit error from bit rot effects to the image stored in the NAND? 

  • Hi Tom,

    Let me check this internally and provide you with an update within one day.

    Regards,
    Ben

  • Hi Tom,

    The expert has suggested reprogramming the flash at run time to prevent data bit errors.

    This could help avoid bit errors during the boot process.

    Regards,
    Ben