This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

TMS320C6746: UPP and random excessive clock cycle consumption.

Part Number: TMS320C6746

I am running a real-time audio processing algorithm on the c6746. When I run the code in externally cached DDR2 memory it intermittently consumes many more clock cycles then it should. By intermittently more than it should, I mean the main processing loop consumes what I consider to be the correct number of clock cycles on the majority of its main loop iterations and then every several iterations (randomly between 2 and 10) the code consumes 5 or 6 times the standard clock cycles. I can get this behavior to stop by disabling the UPP. So it appears to be some kind of interaction with the UPP running. When I run the same code out of all internal memory then the code consumes the correct clocks consistently even with the UPP running.

The UPP is exchanging audio samples with an external FPGA, which is I obviously can’t do without. I have to run the code in external memory because when fully configured the code will not fit (not even close) in internal memory.  I am running the same paired down version in both internal and external memory for this test.  The UPP uses its own DMA to write/read from external cache memory as well.  For this test case there are no other peripherals running, just the UPP and DDR2 cached memory accesses.

Is it possible that this type of behavior is what you would expect from conflicts in the Switched Central Resource? Any ideas on what I may be able to do to get more consistent performance or to alleviate this problem?

  • Hi Dana
    Can the uPP payload reside on internal L2 memory?
    Have you considered c6748 , that has a bigger on chip memory with another 128KB of L3/Shared RAM?

    What speed are you running the core and DDR2?

    Regards
    Mukul
  • Yes, I did a version of this program using 128KB of L2 as cache and 128KB of L2 as simply internal memory with the UPP payload (the In & Out Ping Pong buffers) residing in L2. It behaves the same way, a few of the iterations consume clocks correctly and then you get a single iteration that is significantly more clocks. This repeats over and over again.

    The core is runing at 456 MHz and the DDR2 at 150 MHz.

    Using the c6748 is not an option as the hardware has already been built. Since running the UPP payload in internal memory did not help, it would not appear to be much of a help anyways.

    Any other suggestions?
  • >>It behaves the same way, a few of the iterations consume clocks correctly and then you get a single iteration that is significantly more clocks. This repeats over and over again.

    So my assumption is that it is not the data traffic brought in over upp from external memory that is causing contention with program execution from external memory?
    There is still a possibility that your program code is stalled on some cache updates due to data being dumped in DDR? Do you have further granularity on what portions of your code is taking the extra cycles? Is caching enabled for DDR via MAR bits , perhaps you can check what enabling/disabling the caching for external memory via MAR bits does?
    Need to understand what is the dependency on the audio code vs the upp data - are they completely independent from your code flow/execution standpoint?
    What is the upp data rate / clock rate - does reducing the frequency make any difference (it is plausible that it is not acceptable to reduce frequency in your real application scenario - but perhaps as a debug step it helps figure out what is causing the bottleneck?)
  • Can you also check DDR PBBPR register bprio value. If it is set to default try changing to 0x20
  • Hi Dana
    Any update on this?

    Regards
    Mukul
  • I have been out on vacation, but I have some more information now. With the clock speeds mentioned above the UPP has been receiving 256, 32 bit words from an External FPGA every 48 KHz sample period these would be gathered up by the UPP DMA and written into DSP cached memory. The UPP also writes 2, 32 bit words back to the FPGA (also using the UPP DMA) every 48 KHz sample period.

    I modified the UPP DMA to only receive 4 words from the FPGA instead of the normal 256 words. The occasional excessive clock consumption went away. I then incrementally increased the reads from 4 words back up to 256 to see where the problem returned with the following results:
    At 256 words received the occasional excessive clock consumption was around 5 to 6 times the normal clocks consumed. With only 4 received words via UPP DMA from the FPGA, the clocks consumed were solid (no occasional excessive consumption). It appears that around 32 words the occasional excessive clock consumption begins. By 64 words the occasionally excessive clock cycles is around 1.2 times the normal clock consumption. When I go back to the 256 words it is back to the occasional 5 to 6 times the normal consumption. So there appears to be some internal conflict at the higher UPP DMA receive rate. I am assuming I would see a similar problem on sending at a higher data rate as well (I have not tried it though). Any insight would be appreciated.

    There is a dependency in the code to the audio data (the code is executing an adaptive filter with a decision on the adaptation being the primary dependence). However, I see the occasional excessive clocks regardless of whether the code is adapting the filter or just executing it. This particular version of the code only needs to receive 4, 32 bit words from the FPGA and send 2, However future versions will need to read/send more than that, so we need to know why the occasional excessive clock consumption occurs and if there is a way to reduce the problem. Has this particular problem been observed before?

    The DDR PBBPR register bprio value was already changed from it’s default to 0x30. I tried the suggested value of 0x20 but with no difference.

    Additionally the external DDR2 code resides in the 0xC0000000-0xC0FFFFFF range and uses the MAR[192] bit. When I disable that MAR bit clock consumption goes through the roof as expected, so the cache appears to be functioning as it should.
  • Hi Dana

    Thanks for the additional details.

    We have not seen  issues with what you are reporting here but then again your use-case maybe different then others.

    I am not sure what to make of the 1024 bytes (256 32 bit words) vs 16 bytes transfer differences - apart from the fact that it maybe allowing better interleaving with DSP MDMA accesses  to DDR memory for program/cache  fetch.

    What is your actual uPP clock settings?

    You have played with the L2 cache settings, have you tried L2 cache split between RAM vs cache , even when everything is in DDR ? Is your default L2 all cache?

    If you are using default master priorities - DSP MDMA should be at higher priority (default 2) vs UPP DMA (default 4) - try to to see if setting them at equal priority (MSTPRI) makes the situation worst? This will again point to a bandwidth/interleaving type issue.

    Some concepts on the following wiki

    (Note default priorities listed in the table on the wiki seems incorrect )

    What is the configuration for UPTCR register - and does tweaking that make any difference?

    The internal DMA controller always writes data in bursts of 64 bytes. However, DMA read operations have configurable burst size, which may be set per channel using the RDSIZEI and RDSIZEQ bits in the uPP threshold configuration register (UPTCR). A DMA channel waits until the specified number of bytes leaves its internal buffer before performing another burst read from memory

  • I believe I have found the culprit for the occasional excessive clock cycle consumption. My code performs large FIR computations in the frequency domain. Some routines hold off interrupts for a while. I tried setting the interrupt_threshold to a little less then the DSP clocks between I/O interrupts and the problem went away entirely. The clocks consumed are now as I would expect consistently. Apparently the execution time while running out of internal memory was just fast enough to avoid the symptom, and executing the code out of DDR2 crossed some execution time threshold. Also reducing the words read and stored into memory (from 256 to 4) using the UPP was also enough of a change to avoid seeing the entire problem. I set the threshold and all works as it should.

    Thanks for your ideas and assistance. I will mark this as solved.
  • Dana
    Thanks for taking the time to update the post and confirming that the issue is fixed.
    Always happy to hear uPP success stories on this device - given practically no support in the EVM and SW - I am always encouraged hearing customers make the peripheral work as needed with just documentation.

    Regards
    Mukul