This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

TMS570LS3137: nERROR/ECC LED activates when calling a function from UART ISR

Part Number: TMS570LS3137
Other Parts Discussed in Thread: HALCOGEN,

Hello,

I am using a TMS570LS3137 with HALCoGen 4.07.01, CCS 12.2 and TI ARM Compiler 20.2.7.LTS.

The project is compiled with optimization disabled:

    -mv7R4
    --code_state=32
    -Ooff
    --abi=eabi

Two UART receive paths are used in the application:

1. sciREG:

  • VIM channel: 74
  • ISR: comp_pod_sci_low_level_interrupt
  • Receive interrupt vector: INTVECT1 == 11

2. scilinREG:

  • VIM channel: 27
  • ISR: comp_maintenance_sci_low_level_interrupt
  • Receive interrupt vector: INTVECT1 == 11

Both peripherals are configured with UART_INT_LEVEL_1. The same issue is observed on both interrupt paths.

I have UART receive ISRs declared as follows:

    #pragma CODE_STATE(comp_pod_sci_low_level_interrupt, 32)
    #pragma INTERRUPT(comp_pod_sci_low_level_interrupt, IRQ)

    void comp_pod_sci_low_level_interrupt(void)
    {
        if (g_s_cp != NULL)
        {
            uint32_t vec = g_s_cp->pod_com.instance->INTVECT1;

            if (11U == vec)
            {
                uint8_t received_data =(uint8_t)(g_s_cp->pod_com.instance->RD & 0x00FFU);

                comp_pod_ring_buf_put(g_s_cp, received_data);
            }
        }
    }

The called function is:

    static result_enum_t comp_pod_ring_buf_put(comp_pod_struct_t *const p_cp, uint8_t data)
    {
        if (NULL == p_cp)
        {
            return RESULT_ERR_NULL;
        }

        uint16_t next = (p_cp->ring_buf.head + 1U) & POD_RING_BUF_MASK;

        if (next != p_cp->ring_buf.tail)
        {
            p_cp->ring_buf.buffer[p_cp->ring_buf.head] = data;
            p_cp->ring_buf.head = next;
        }

        return RESULT_SUCCESS;
    }

When comp_pod_ring_buf_put() is called from the ISR, the board's ECC/nERROR LED becomes active.

However, the application continues running normally. It does not remain in ramErrorReal, flashErrorReal or another data-abort loop.

If I copy the same ring-buffer operations directly into the ISR, the LED does not become active:

    uint16_t next = (g_s_cp->ring_buf.head + 1U) & POD_RING_BUF_MASK;

    if (next != g_s_cp->ring_buf.tail)
    {
        g_s_cp->ring_buf.buffer[g_s_cp->ring_buf.head] = received_data;

        g_s_cp->ring_buf.head = next;
    }

The same behavior is observed in another UART component with an equivalent ring-buffer function.

I also increased the IRQ stack from 0x100 bytes to 0x400 bytes and adjusted the linker RAM boundaries accordingly. This did not change the behavior, so a simple IRQ stack overflow does not appear to be the cause.

Because optimization is disabled, the function is not inlined. Calling it adds a BL instruction, executes code from another Flash address and creates an additional function stack frame. Copying the function body into the ISR avoids these operations.

What could be causing this problem?

Best regards,
Hasan

  • Hi Hasan,

    I would suggest you verify which ESM bit got set. You can refer below e2e to understand how to verify these bits:

    (+) TMS570LC4357: ESM High Interrupt occurs with no pending interrupt in ESMIOFFHR - Arm-based microcontrollers forum - Arm-based microcontrollers - TI E2E support forums

    We have one internal AI which can analyze all our internal database, and it is providing below suggestion, please refer this one as well:

    ---

    ## Root Cause Analysis

    ### The Core Issue: Uninitialized IRQ Stack Memory and ECC

    The TMS570LS3137 uses **TCRAM (SRAM) with hardware ECC**. The ECC is computed and checked on every **64-bit (double-word) read**. If a memory location has never been written since power-on, its ECC bits are in an indeterminate state. Reading such a location will trigger an **ECC single-bit correction or double-bit detection error**, which asserts the `nERROR` pin and lights the ECC LED.

    The key distinction in your observation is:

    | Scenario | Behavior |
    |---|---|
    | Ring-buffer code **inlined** in ISR | No ECC error |
    | Ring-buffer code in a **called function** | ECC error triggered |

    This difference is explained entirely by **stack frame usage**:

    - When the code is **inlined**, the compiler uses only the registers already allocated for the ISR's stack frame — no new stack region is touched.
    - When `comp_pod_ring_buf_put()` is **called**, the compiler generates a new stack frame for that function. This pushes the link register (`LR`), possibly `r4`–`r11`, and local variables onto the **IRQ stack** at addresses that have **never been written before**. The subsequent function epilogue **reads back** those stack locations (to restore registers via `POP`/`LDMFD`), and if those addresses were never initialized with valid ECC, the read triggers an ECC fault.

    > **In short: The function call causes a stack `PUSH` followed by a `POP`. The `POP` reads memory that was never ECC-initialized, triggering the fault.**

    Increasing the stack size to `0x400` did not help because the problem is not *overflow* — it is that the **newly allocated stack pages were never written** (and thus have no valid ECC syndrome).

    ---

    ## Why the Application Continues Running

    The TMS570 ECC logic distinguishes between:
    - **Single-bit errors (correctable):** The SECDED hardware corrects the data and asserts `nERROR`, but execution continues. This matches your observation exactly — the LED lights, but the application keeps running.
    - **Double-bit errors (uncorrectable):** These cause a data abort.

    You are hitting **single-bit ECC errors on uninitialized TCRAM**, which are corrected in hardware but still assert `nERROR`.

    ---

    ## Solutions

    ### White check mark Solution 1: Initialize the Entire IRQ Stack Region at Startup (Recommended)

    The standard fix is to **write a known pattern to the entire IRQ stack** during startup, before any interrupts are enabled. This initializes the ECC for every address in the stack region.

    HALCoGen-generated startup code (`sys_startup.c`) typically does this for the full SRAM via the `_memInit_` routine, but **it may not cover the IRQ stack region if it was recently enlarged or if the linker script was modified**.

    Check your `sys_startup.c` for a call like:

    ```c
    /* Initialize TCRAM ECC */
    _memInit_();
    ```

    And verify that the **entire IRQ stack region** (as defined in your linker `.cmd` file) falls within the range covered by `_memInit_`. If you enlarged the stack from `0x100` to `0x400` and adjusted the linker RAM boundaries, confirm that the new boundaries are reflected in the startup ECC initialization range.

    ---

    ### White check mark Solution 2: Explicitly Zero-Initialize the Stack in Your Startup Code

    Add an explicit stack scrub in your startup assembly or C startup code before enabling VIM/interrupts:

    ```asm
    ; In sys_startup.asm or equivalent
    ; Switch to IRQ mode and zero the stack
    MSR CPSR_c, #0xD2 ; IRQ mode, IRQ/FIQ disabled
    MOV r0, #0
    MOV r1, #0
    ; ... fill IRQ stack region with STM/STMIA loop
    ```

    Or in C (before `main()` or before enabling interrupts):

    ```c
    extern uint32_t __IRQ_STACK_START__;
    extern uint32_t __IRQ_STACK_END__;
    uint32_t *p = &__IRQ_STACK_START__;
    while (p < &__IRQ_STACK_END__) {
    *p++ = 0xDEADBEEFU; /* any write initializes ECC */
    }
    ```

    ---

    ### White check mark Solution 3: Verify HALCoGen `_memInit_` Coverage

    In HALCoGen-generated projects, open `sys_startup.c` and look for the TCRAM initialization block. It typically uses the **MINITGCR** and **MSIENA** registers to trigger a hardware memory initialization. Confirm:

    1. The hardware init covers the full SRAM range including your stack.
    2. The init completes (polls the status bit) **before** `vimInit()` or any interrupt enable.

    ---

    ### White check mark Solution 4: Use `__attribute__((noinline))` Carefully / Restructure

    This is a **workaround, not a fix**, but confirms the diagnosis: if you mark `comp_pod_ring_buf_put` as `__attribute__((always_inline))`, the ECC error disappears — because no new stack frame is created. You already observed this with manual inlining.

    ---

    ## Summary Checklist

    - [ ] Confirm `_memInit_()` in `sys_startup.c` covers the **full IRQ stack** address range after your linker script change.
    - [ ] Confirm memory initialization completes **before** `vimInit()` / interrupt enable.
    - [ ] If using a custom linker `.cmd`, verify the IRQ stack section is placed in a region covered by the ECC scrub.
    - [ ] Check TI's errata for TMS570LS3137 (SPNZ193) for any related silicon errata on TCRAM ECC initialization.

    ---

    > Warning️ **Note:** This analysis is based on general expert knowledge of the TMS570LS3137 architecture, ARM Cortex-R4 ECC behavior, and TI ARM Compiler behavior, as this specific issue falls outside the scope of the EDA tool knowledge bases available to me. I recommend cross-referencing with the [TI E2E Community forums](https://e2e.ti.com) and the TMS570LS3137 Technical Reference Manual (SPNU489) for authoritative confirmation.

    --
    Thanks & regards,
    Jagadish.