This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

AM6442: The legacy interrupt from the FPGA stops being called.

Genius 3215 points

Part Number: AM6442

Hi All,

We are currently verifying the PCIe Root Complex functionality on a custom board equipped with AM64x, using Linux built based on Processor SDK Linux v9.2.

An FPGA is connected as the endpoint device.
On Linux, the FPGA is correctly recognized as a PCIe endpoint device, and we are able to access its PCIe configuration space and BARs.
Next, we are attempting to verify whether legacy interrupts from the FPGA can be detected.

To verify legacy interrupt detection, we used the following approach:
We assigned a legacy interrupt vector to the FPGA using the pci_alloc_irq_vectors() function,
retrieved the corresponding Linux IRQ number using pci_irq_vector(),
and registered a handler function with request_irq() to be called when the IRQ occurs.
When we triggered a legacy interrupt once from the FPGA, we confirmed that the registered handler function was successfully called.

However, when we repeatedly triggered legacy interrupts from the FPGA at approximately one-second intervals,
the handler function was called for a certain number of times,
but after that, it stopped being called.
This “certain number” is not consistent and varies with each repetition.

Upon checking the Linux source code, we found that when a legacy interrupt occurs,
the j721e_pcie_legacy_irq_handler() function in pci-j721e-host.c is called,
and within that function, the IRQ corresponding to the interrupt vector assigned to the FPGA is triggered.
Reference:
https://git.ti.com/cgit/ti-linux-kernel/ti-linux-kernel/tree/drivers/pci/controller/cadence/pci-j721e-host.c?h=ti-rt-linux-6.1.y

We added log output to the j721e_pcie_legacy_irq_handler() function,
and observed that while the legacy interrupts from the FPGA were being triggered repeatedly,
log output appeared up to the point where the handler function stopped being called.
After that point, no further log output was observed.
Therefore, we believe the reason the registered handler function stopped being called
is that the j721e_pcie_legacy_irq_handler() function itself was no longer being invoked.

We would like to ask the following two questions:

[Question 1]
When a legacy interrupt is triggered by a PCIe endpoint device,
is it correct to expect that the j721e_pcie_legacy_irq_handler() function in pci-j721e-host.c will be called
in Linux built based on PSDK Linux v9.2?

[Question 2]
When legacy interrupts are repeatedly triggered from the endpoint device,
the j721e_pcie_legacy_irq_handler() function eventually stops being called.
What could be the possible causes of this phenomenon?

Best Regards,

Ito

  • Hi Ito-san,

    I will review the corresponding pcie code and get back to you.

  • Hi BIn,

    Thank you for your help.

    Best Regards,

    Ito

  • Hi All,

    How is the progress on this question?

    Best Regards,

    Ito

  • Hi Ito-san,

    [Question 1]
    When a legacy interrupt is triggered by a PCIe endpoint device,
    is it correct to expect that the j721e_pcie_legacy_irq_handler() function in pci-j721e-host.c will be called
    in Linux built based on PSDK Linux v9.2?

    Yes, this is expected.

    [Question 2]
    When legacy interrupts are repeatedly triggered from the endpoint device,
    the j721e_pcie_legacy_irq_handler() function eventually stops being called.
    What could be the possible causes of this phenomenon?

    If you wait longer then let the endpoint device to deassert and re-assert the legacy interrupt, does it change the behavior of the problem?

  • Hi Ito-san,

    We added log output to the j721e_pcie_legacy_irq_handler() function,
    and observed that while the legacy interrupts from the FPGA were being triggered repeatedly,
    log output appeared up to the point where the handler function stopped being called.
    After that point, no further log output was observed.

    When the FPGA repeatedly trigger the legacy interrupts, have you checked the PCIe EP always sent ASSERT_INTA packet followed by DEASSERT_INTA packet and never missed a DEASSERT_INTA?

  • Hello Ito-san,

    When the issue occurs and the interrupts are no longer triggered i.e. j721e_pcie_legacy_irq_handler() is no longer invoked, can you please run the following command at the Linux prompt and see if the interrupts get triggered again?

    devmem2 0x0F1000C8 w 0x2

    Regards,
    Siddharth.

  • Hi Bin-san,

    Thank you for your reply.

    If you wait longer then let the endpoint device to deassert and re-assert the legacy interrupt, does it change the behavior of the problem?

    We tried extending the interval for repeatedly generating legacy interrupts from the FPGA from 1 second to 10 seconds,
    but the behavior did not change.
    After a certain number of occurrences, the j721e_pcie_legacy_irq_handler() function stopped being called.

    When the FPGA repeatedly trigger the legacy interrupts, have you checked the PCIe EP always sent ASSERT_INTA packet followed by DEASSERT_INTA packet and never missed a DEASSERT_INTA?

    We are using an FPGA from Altera and utilizing Intel’s PCIe Hard IP.
    We have confirmed that the FPGA can consistently issue instructions to send packets to Intel’s PCIe Hard IP.
    Although we have not directly verified the packets themselves, we believe there are no transmission errors.

    Best Regards,

    Ito

  • Hi Siddharth-san,

    Thank you for your reply.

    I executed the command you provided, but the interrupt was not retriggered.

    Best Regards,

    Ito

  • I executed the command you provided, but the interrupt was not retriggered.

    Ito-san, it's strange that the interrupt wasn't triggered because I am able to recreate the issue in my setup and the devmem2 command triggered the interrupt. Can you please try raising the interrupt again from the FPGA after running the devmem2 command on the AM64 SoC and see if the interrupt is registered?

    Regards,
    Siddharth.

  • Ito-san,

    Can you please apply the following 'diff' to the pci-j721e-host.c driver in the SDK and check if the issue is resolved?

    diff --git a/drivers/pci/controller/cadence/pci-j721e-host.c b/drivers/pci/controller/cadence/pci-j721e-host.c
    index 624ae38c38f3..275bfe658d36 100644
    --- a/drivers/pci/controller/cadence/pci-j721e-host.c
    +++ b/drivers/pci/controller/cadence/pci-j721e-host.c
    @@ -125,6 +125,7 @@ static void j721e_pcie_legacy_irq_handler(struct irq_desc *desc)
            }
    
            chained_irq_exit(chip, desc);
    +       j721e_pcie_user_writel(pcie, USER_EOI_REG, EOI_LEGACY_INTERRUPT);
     }
    
     static void j721e_pcie_irq_eoi(struct irq_data *data)

    Regards,
    Siddharth.

  • Hi Siddharth-san,

    Thank you for your reply.

    After checking the interrupt registration, I confirmed that the j721e_pcie_legacy_irq_handler() function,
    which had stopped being called, is now being invoked again.

    Also, by applying the patch you provided to pci-j721e-host.c,
    I confirmed that the j721e_pcie_legacy_irq_handler() function is called even without executing the devmem2 command.

    However, within the processing of the j721e_pcie_legacy_irq_handler() function,
    although it should trigger the IRQ corresponding to the interrupt vector assigned to the FPGA, the registered function is still not being called.

    The FPGA is configured to generate a legacy interrupt when accessing a specific address in its PCIe BAR space.
    The interval between the FPGA instructing the Intel PCIe Hard IP to send an ASSERT_INTA packet and then a DEASSERT_INTA packet is set to 20 microseconds.

    I suspected that the registered function might not be called because more than 20 microseconds elapse between the legacy interrupt occurring and the j721e_pcie_legacy_irq_handler() function being invoked. To investigate, I configured the kernel to output logs just before accessing the BAR space and when the j721e_pcie_legacy_irq_handler() function is called.

    As a result, the difference in timestamps between the log output before BAR access and when the handler is called was usually around 10 microseconds, but in about 1 or 2 cases out of 100, it exceeded 20 microseconds. When the difference was over 20 microseconds, the j721e_pcie_legacy_irq_handler() function was called, but the registered function was not.

    When I first asked my question, I mentioned that Linux was built based on Processor SDK Linux v9.2. To be more precise, it was built based on PROCESSOR-SDK-LINUX-RT-AM64X version of PSDK Linux v9.2.

    Therefore, I have an additional question:

    [Question 3]
    In Processor SDK Linux-RT v9.2, is there any benchmark information regarding the time it takes from receiving a PCIe legacy interrupt to invoking the j721e_pcie_legacy_irq_handler() function?
    For example, if the time from asserting to deasserting the legacy interrupt on the FPGA is shorter than the maximum time it takes for the handler to be called, the handler will be invoked, but the registered function will not.
    I am concerned that the occurrence of a legacy interrupt may not be properly notified to the user program.

    Best Regards,

    Ito

  • Ito-san,

    After checking the interrupt registration, I confirmed that the j721e_pcie_legacy_irq_handler() function,
    which had stopped being called, is now being invoked again.

    Also, by applying the patch you provided to pci-j721e-host.c,
    I confirmed that the j721e_pcie_legacy_irq_handler() function is called even without executing the devmem2 command.

    Thank you for confirming that the patch fixes the issue of the interrupt handler not being invoked.

    However, within the processing of the j721e_pcie_legacy_irq_handler() function,
    although it should trigger the IRQ corresponding to the interrupt vector assigned to the FPGA, the registered function is still not being called.

    The FPGA is configured to generate a legacy interrupt when accessing a specific address in its PCIe BAR space.
    The interval between the FPGA instructing the Intel PCIe Hard IP to send an ASSERT_INTA packet and then a DEASSERT_INTA packet is set to 20 microseconds.

    Does the FPGA send a DEASSERT_INTA unconditionally after 20 microseconds? Ideally, the DEASSERT_INTA should be sent only after the interrupt handler for the FPGA running on Linux clears the interrupt on the FPGA.

    As a result, the difference in timestamps between the log output before BAR access and when the handler is called was usually around 10 microseconds, but in about 1 or 2 cases out of 100, it exceeded 20 microseconds. When the difference was over 20 microseconds, the j721e_pcie_legacy_irq_handler() function was called, but the registered function was not.

    This is expected behavior. Please note the following:
    i) The j721e_pcie_legacy_irq_handler() function will be invoked when the FPGA sends an ASSERT_INTA.
    ii) During the execution of j721e_pcie_legacy_irq_handler(), the STATUS register is read to identify which among INTA, INTB, INTC and INTD were asserted.
    iii) If the DEASSERT_INTA is sent by the FPGA before the STATUS register is read, the STATUS register will indicate that none of INTA, INTB, INTC or INTD were asserted.

    The issue is due to the FPGA unconditionally sending a DEASSERT_INTA after 20 microseconds. The fix therefore is the following:
    i) Do not send a DEASSERT_INTA unconditionally from the FPGA
    ii) The Interrupt handler for the FPGA running on Linux (not j721e_pcie_legacy_irq_handler(), but the interrupt handler for the FPGA endpoint) should service the interrupt and 'CLEAR' the interrupt on the FPGA Endpoint
    iii) Only when the Interrupt on the FPGA Endpoint is cleared, it should send DEASSERT_INTA

    [Question 3]
    In Processor SDK Linux-RT v9.2, is there any benchmark information regarding the time it takes from receiving a PCIe legacy interrupt to invoking the j721e_pcie_legacy_irq_handler() function?
    For example, if the time from asserting to deasserting the legacy interrupt on the FPGA is shorter than the maximum time it takes for the handler to be called, the handler will be invoked, but the registered function will not.
    I am concerned that the occurrence of a legacy interrupt may not be properly notified to the user program.

    The interrupt latency isn't benchmarked and as you have rightly noticed, it is variable by nature. Since the interrupt is being handled by an ARM Cortex-A which is an application core and is optimized for performance rather than latency, if the requirement is to have tight constraint on the interrupt handling latency, the ARM Cortex-R is well suited for such cases.

    Nevertheless, irrespective of a bound on the interrupt latency, the approach being taken where the DEASSERT_INTA is sent unconditionally is incorrect. DEASSERT_INTA should only be sent once the interrupt handler services it and Clears the interrupt on the FPGA Endpoint.

    Regards,
    Siddharth.

  • Hi Siddharth-san,

    Instead of sending DEASSERT_INTA unconditionally from the FPGA, we modified the FPGA so that writing 1 to a specific address in the FPGA's PCIe BAR space issues an ASSERT_INTA packet, and writing 0 to the same address issues a DEASSERT_INTA packet.

    Furthermore, in the function registered as the FPGA endpoint interrupt handler, we implemented a mechanism to repeatedly write 0 to that address until a read-back confirms that 0 has been successfully written.

    After applying these changes, we confirmed that the following sequence (1–4) works as intended:

    1. Write 1 to the specific address in the BAR space to trigger ASSERT_INTA from the FPGA
    2. Execute the j721e_pcie_legacy_irq_handler() function
    3. Execute the function registered as the FPGA endpoint interrupt handler (writes 0 to the specific address)
    4. Trigger DEASSERT_INTA from the FPGA

    However, upon checking the log output, we observed that for a single ASSERT_INTA issuance, the j721e_pcie_legacy_irq_handler() function was called 4–5 times. The function registered from j721e_pcie_legacy_irq_handler() was also invoked twice.

    Additionally, we repeatedly performed the operation “write 1 to the specific BAR space address to trigger ASSERT_INTA from the FPGA” at intervals of about one second. After approximately 500 repetitions, the SSH window connected to the Linux environment via Ethernet closed, and we were unable to log in again afterward.
    It appeared as though Linux had “crashed.”

    From these results, it seems that the j721e_pcie_legacy_irq_handler() function operates independently of the registered function, and that DEASSERT_INTA is issued by the FPGA only after the registered function successfully writes 0. Until the INTA status is deasserted, the handler appears to be repeatedly invoked.

    If, for some reason, the registered function fails to write 0 correctly and the INTA status remains asserted, we suspect that j721e_pcie_legacy_irq_handler() will keep being called repeatedly, eventually causing Linux to “crash.”

    Therefore, we have two additional questions:

    [Question 4]
    We assume that j721e_pcie_legacy_irq_handler() calls the IRQ registered function. Does j721e_pcie_legacy_irq_handler() wait for the called function to complete its processing?

    [Question 5]
    If the INTA status remains asserted, is it possible that j721e_pcie_legacy_irq_handler() will be repeatedly invoked indefinitely, leading to a situation where Linux “crashes”?

    Best Regards,

    Ito

  • Hello Ito-san,

    After applying these changes, we confirmed that the following sequence (1–4) works as intended:

    Thank you for testing the suggested change and confirming that it works.

    However, upon checking the log output, we observed that for a single ASSERT_INTA issuance, the j721e_pcie_legacy_irq_handler() function was called 4–5 times. The function registered from j721e_pcie_legacy_irq_handler() was also invoked twice.

    This occurs because the DEASSERT_INTA from the FPGA hasn't reached the PCIe Controller on the AM64 SoC yet. Please note that PCIe emulates Legacy Interrupts (INTx) using Messages which travel over the PCIe link and will take a finite duration to propagate.

    If, for some reason, the registered function fails to write 0 correctly and the INTA status remains asserted, we suspect that j721e_pcie_legacy_irq_handler() will keep being called repeatedly, eventually causing Linux to “crash.”

    As stated above, it is not the case that the registered function is failing to write a 0. DEASSERT_INTA is taking time to propagate and until it is received, the j721e_pcie_legacy_irq_handler() will continue to be invoked.

    The fix therefore is to account for the finite propagation delay for the DEASSERT_INTA message sent from the FPGA after the 'registered function' clears it on the FPGA. In the 'registered function', please add a delay of around 1 millisecond before exiting the function. This way, the complete sequence will be:

    1. Write 1 to the specific address in the BAR space to trigger ASSERT_INTA from the FPGA
    2. The j721e_pcie_legacy_irq_handler() function is invoked
    3. The function registered as the FPGA endpoint interrupt handler is invoked
    4. The registered interrupt handler writes '0' to the specific address in the BAR space to clear the interrupt and cause the FPGA to send DEASSERT_INTA
    5. After writing '0' to the BAR space, wait for 1 millisecond using 'mdelay(1)'.
    6. Return from the registered interrupt handler back to its caller which is j721e_pcie_legacy_irq_handler().
    7. j721e_pcie_legacy_irq_handler() writes EOI_LEGACY_INTERRUPT to the USER_EOI_REG register and exits

    In the sequence above, acknowledging an existing interrupt and preparing for new interrupts is done by writing EOI_LEGACY_INTERRUPT (0x2) to the USER_EOI_REG register. After an existing interrupt is acknowledged by this mechanism, if a DEASSERT_INTA wasn't received yet by this point in time, an interrupt will be triggered again for the previous interrupt itself. Therefore, the suggestion described in step-5 above will prevent this from occurring.

    Since the delay depends on the setup being used, please experiment with the delay values starting with 1 millisecond and increasing it in units of 1 millisecond until the issue of spurious interrupts no longer occurs.

    [Question 4]
    We assume that j721e_pcie_legacy_irq_handler() calls the IRQ registered function. Does j721e_pcie_legacy_irq_handler() wait for the called function to complete its processing?

    Yes, it waits for the 'registered function for the FPGA endpoint interrupt handler' to run to completion.

    [Question 5]
    If the INTA status remains asserted, is it possible that j721e_pcie_legacy_irq_handler() will be repeatedly invoked indefinitely, leading to a situation where Linux “crashes”?

    Yes, that is correct. Please implement 'step-5' suggested above and check if spurious interrupts are no longer seen.

    Regards,
    Siddharth.

  • Hi Siddharth-san,

    Thank you for your help.

    Yes, it waits for the 'registered function for the FPGA endpoint interrupt handler' to run to completion.

    From the log outputs, it appeared that the j721e_pcie_legacy_irq_handler() function did not wait for the registered handler to complete.
    When we added log printing at steps 2, 4, and 7 of the proposed sequence, the logs were printed in the order:
    “Step 2 log (j721e… function) → Step 7 log (j721e… function) → Step 4 log (registered handler).”

    Therefore, when registering the FPGA endpoint interrupt handler via request_irq(), we had previously set the third argument flags to 0.
    As a trial, we specified IRQF_NO_THREAD in flags, and the logs were then printed in the order:
    “Step 2 log (j721e… function) → Step 4 log (registered handler) → Step 7 log (j721e… function).”
    The timestamps of the Step 4 log and Step 7 log differed by more than 1 ms.
    We were also able to confirm behavior consistent with the proposed sequence.

    Next, we modified the request_irq() call for the FPGA endpoint interrupt handler to always specify IRQF_NO_THREAD in flags, and ran a user program that writes to a particular address in the BAR space every 100 ms.
    Under these conditions, we confirmed correct behavior per the sequence for 10,000 consecutive iterations.

    Based on this, we believe the issue where the legacy interrupts from the FPGA eventually stop being detected has been resolved.

    [Question 6]
    We are building Linux based on PSDK Linux v9.2 from PROCESSOR-SDK-LINUX-RT-AM64X.
    Is it possible that, if IRQF_NO_THREAD is not specified in flags when registering a handler via request_irq(), the j721e_pcie_legacy_irq_handler() function would not wait for the registered handler to finish?

    Best Regards,

    Ito

  • From the log outputs, it appeared that the j721e_pcie_legacy_irq_handler() function did not wait for the registered handler to complete.
    When we added log printing at steps 2, 4, and 7 of the proposed sequence, the logs were printed in the order:
    “Step 2 log (j721e… function) → Step 7 log (j721e… function) → Step 4 log (registered handler).”

    In general, a 'threaded_irq' is used when the interrupt handlers (registered handler in this case) may take time to execute, AND, there are potentially multiple interrupts and their respective handlers (INTA, INTB, INTC and INTD). Since we don't want to block the interrupt handlers of INTB, INTC and INTD until the handler(s) for INTA finish, we use the 'threaded_irq' which schedules the 'registered handler' to run at a later point in time.

    Next, we modified the request_irq() call for the FPGA endpoint interrupt handler to always specify IRQF_NO_THREAD in flags, and ran a user program that writes to a particular address in the BAR space every 100 ms.
    Under these conditions, we confirmed correct behavior per the sequence for 10,000 consecutive iterations.

    Based on this, we believe the issue where the legacy interrupts from the FPGA eventually stop being detected has been resolved.

    Thank you for identifying a potential fix and sharing it. Although this does fix the issue, please note that 'IRQF_NO_THREAD' is not preferred in Real-Time systems. By marking the handler using 'IRQF_NO_THREAD', the handler not only disables all IRQs, but it also delays execution of other interrupt handlers. Therefore, if it satisfies your use-case given the potential issues in Real-Time systems with this approach, it is alright to go ahead while being aware of the drawbacks.

    We are building Linux based on PSDK Linux v9.2 from PROCESSOR-SDK-LINUX-RT-AM64X.
    Is it possible that, if IRQF_NO_THREAD is not specified in flags when registering a handler via request_irq(), the j721e_pcie_legacy_irq_handler() function would not wait for the registered handler to finish?

    Yes, that indeed seems to be the case given that we are using an 'RT' (Real-Time) system and the preference is to have Threaded IRQ handlers.

    Since the 'registered handler' will run after 'Step 7' (with the default driver), one way to still make it work in the updated sequence is:

    1. Write 1 to the specific address in the BAR space to trigger ASSERT_INTA from the FPGA
    2. The j721e_pcie_legacy_irq_handler() function is invoked
    3. j721e_pcie_legacy_irq_handler() schedules the 'registered endpoint interrupt handler' to be executed in the future
    4. j721e_pcie_legacy_irq_handler() writes EOI_LEGACY_INTERRUPT to the USER_EOI_REG register and exits
    5. The function registered as the FPGA endpoint interrupt handler is invoked
    6. The registered interrupt handler writes '0' to the specific address in the BAR space to clear the interrupt and cause the FPGA to send DEASSERT_INTA
    7. The user-space application which triggers new interrupts from the Endpoint should be aware of 'Step 6' being executed before it writes to the BAR to trigger the next interrupt from the Endpoint - New interrupts shouldn't be generated without the previous one being serviced / cleared.

    For steps 6 and 7 to be synchronized, a 'status' bit can be used in the BAR space with the sequence being:
    - Endpoint interrupt handler writes to the BAR space to clear the interrupt
    - Endpoint interrupt handler writes to the 'status' bit in the BAR space to indicate that the interrupt has been cleared
    - User-space application waits for the 'status' bit in the BAR space to be set
    - After verifying that the 'status' bit in the BAR space is set, the user-space application clears the 'status' bit and then writes to the BAR space to raise a new interrupt

    Regards,
    Siddharth.