This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

The EMIF16 take too much time before write and read

Hi TI engineer

         Our board is based on C6678.Its EMIF16 is coonnected to 2 devices through CE0 and CE2(name it device A and B).

         In our application , we need to access both 2 devices time to time.But we found that:

          If I only read 20 words from device B,It take 500ns;

          If I write some data to device A first , then I read 20 words from device B.It will take 2000ns.

          So it seems like it take them 1500ns to switch a device to another device?

          Could you tell me if it is true?And how can I avoid this to descend my cost on  EMIF16?Thank you very much.

  • Hi Yuchao Wang,

    I've forwarded this to the EMIF experts. Their feedback should be posted here.

    BR
    Tsvetolin Shulev
  • I found this information in one of our e2e post:

    The minimum time for entire read cycle i.e. setup + strobe + hold is 50ns.

    At 1.25GHz, EMIF16 clock = CPU/6 =>  4.8ns.

    So r_setup + r_strobe + r_hold = 50/4.8 = 11 rounded up.

    So it takes a minimum of 11 complete CPU/6 clock cycles for one read access, so the actual time for the entire read cycle = 4.8 * 11 = 52.8ns.

    For 16-bit wide interface, the theoretical max read throughput = 16/52.8 = 303.03 Mbps

    At 1GHz, EMIF16 clock period = 6ns.

    Using the same calculation as above, we get theoretical max read throughput = 16/54 = 296.29 Mbps

    The minimum time for entire write cycle is 45ns.

    Using the same procedure as for reads,

    Write throughput at 1.25GHz = 16/48 (10 CPU/6 clocks) = 333.33 Mbps

    Write throughput at 1GHz = 16/48 (8 CPU/6 clocks) = 333.33 Mbps

    The above theoretical max throughputs will be half for an 8-bit interface.

    In practice the path through the SOC to/from the EMIF16 module introduces some delays of its own, so observed throughput for reads and writes will be lower.

    To give you an example of silicon testing - At CPU = 900MHz, the EMIF16 setup/strobe/hold parameters were configured to give following:

    Theoretical read throughput = 240 Mbps

    Theoretical write throughput = 300 Mbps

    The observed throughput was closer to 236Mbps for reads and 274Mbps for writes.

    • I should add that the 45ns and 50ns numbers are based on the latency of available NOR flash components in the market since NOR flash is the fastest memory. It is not a limitation of the DSP. Accordingly, please ignore my last statement about "We will clarify the minimum read/write timing parameters in the documentation". We will address your concern over extended wait. A new rev of the EMIF16 user's guide will be released soon with a lot more clarity on how to program the EMIF16 registers.
    • The max throughput supported by the EMIF16 interface is (16b/8b) * CPU/6.
    • For write performance the bottleneck is the memory device. So EMIF16 will be able to send data at (16b/8b) * CPU/6 if memory device is fast enough.
    • I mentioned in my earlier post that the SOC introduces some delays in the switch fabric for reads. So for reads, the bandwidth is limited by the memory device and the internal delays in read return path.
    • The max bandwidth at the EMIF16 pins is (16b/8b) * CPU/18.

  • Hi Raja:
    Well I could make it clear for you:
    The EMIF16 of C6678 is connected to FPGA in our system; C6678 works at 1G Hz.
    We test the time schedule from the FPGA.It seems like the time of access is right, just as you say, setup + strobe + hold.
    So now we hesitate that: after we send the read/write access command( like a = *(short*)0x70000000 ), the transfer didn't start immediately.
    Could you tell us it takes C6678 how much cycles to send the command to EMIF16,and active the "CS" signal after software execute the access code? Is there any delay in this?
    Otherwise, if I active the CS0 first, the active CS2, is it any influence on the reacted time?
  • Could you tell me after the execution , say, a = *(int*)0x70000000, it will take how much cycle for C6678 to send command to EMIF16, active the CS and start the access? Could the delay influenced by the switch from one CS to another CS?
    Thank you very much.

    Regards,
    Yuchao
  • Hi Raja
    Have you received my new message?We have made some new test about this.It seems like the data
    cost too much time in the internal bus in C6678.

    First,our EMIF initial param like that:
    Main Freq:1GHz
    CE 1:
    read setup = 2 cycle; strobe = 16 cycle; hold = 1 cycle;
    write setup = 4 cycle; strobe = 8 cycle; hold = 2 cycle;
    no_wait
    CE 2:
    read setup = 8 cycle; strobe = 40 cycle; hold = 4 cycle;
    write setup = 4 cycle; strobe = 40 cycle; hold = 4 cycle;
    wait0
    Second,we create the test code as followed:
    Emif16_init();
    // CE1
    tick_start = _itoll(TSCH,TSCL);
    *(unbsigned short*)0x74000000 = j;
    tick1 = _itoll(TSCH,TSCL) - tick_start;
    //CE2
    tick_start = _itoll(TSCH,TSCL);
    j = *(unbsigned short*)0x7C000000;
    tick2 = _itoll(TSCH,TSCL) - tick_start;

    test result:
    tick1:25
    tick2:2814
    tick2(without CE1 operation): 500

    At the same time,we check the schedule from the FPGA.The signal of CE/OE/WE shows no error.The data is sended back in less that 100 cycle after the active of CE.It seems like most of the 2814 cycle is costed after data has been transfered to EMIF16.

    Well,I hope this could help you understand what we met.Looking reward for you reply.
    Regards,
    Yuchao
  • Hi Yuchao,

    Do you setup the EMIF16 hidden register bit as suggested in SPRZ334H Silicon Errata Usage Note 25 ?

    Usage Note 25 Performance Degradation for Asynchronous Access Caused by An Unused Feature Enabled in EMIF16 Usage Note
    Revision(s) Affected: 1.0, 2.0
    Details: Although it supports only asynchronous mode operation on the device, the EMIF16 module has a legacy 'synchronous mode' feature that is enabled by default. While this synchronous mode is enabled, EMIF16 issues periodic refresh commands that take precedence over asynchronous accesses commands and stall the execution of the latter until the refresh command is executed. This stall results in reduced throughput of asynchronous accesses when EMIF16 tries to read or write to the asynchronous memory. The stall will manifest itself as a long delay between asynchronous accesses.
    Workaround: Programming bit 31 at the 32-bit address 0x20C00008 to 1 will disable the synchronous mode feature.
    *(Uint32*) 0x20C00008 |= 0x80000000; //Disable synchronous mode feature
    When the synchronous mode is disabled, EMIF16 will not issue any refresh commands. This will no longer result in stall cycles between asynchronous accesses and thus the performance will be improved. This bit affects only the refreshes issued by EMIF16. It does not affect the rest of the device.

    Regards,

    Sei Kato

  • Hi Sei
    I will try to do it.Could you tell me if disabling the synchronous mode, will bring any side effects? Thank you very much.
    Regards,Yuchao
  • Hi Sei
    We have made it! The cofiguration to 0x20c00008 seems to work out this problem. Now the delay has disappeared.
    But still I want to know that what will happen if we forbid the "refresh commands"?Is it safe?
    Regards,
    Yuchao.
  • Hi Yuchao,

    Glad to hear your issue has been solved.
    The Errata says it does not affect other than the refresh commands, and as long as I'm applying this on C6655 no significant side effect can be observed.

    Regards,
    Sei Kato