This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

Performance of XDS200?

Other Parts Discussed in Thread: OMAP-L138, MSP-FET

Hello!

My name is Yeun-Hun Choi.

I would like to buy emulator for TMS320F28x.

Recently, XDS200 emulator was launched by Spectrum Digital.

I hope to know it's performance when comparing with XDS510 series.(Clock, Real time monitoring, Flashing etc..)

 

  • I'm moving this to the CCS forum for better assistance.


    Thanks,
    Brett

  • On a related note, does the XDS200 perform better than an XDS100v2 when run from CCS hosted in a virtual machine?

  • The XDS200 claims to be much more performant than the XDS100, but that doesn't necessarily say much about the XDS200.   My experience so far is that XDS100v2 is so slow as to be completely unusable and XDS510-USB is quite good except for transferring large amounts of data to/from memory.  We need to purchase another emulator for an OMAP-L138 board and I'm undecided between the XDS560 (same price as the XDS510) and the XDS200 (about 1/3 the price of an XDS510).

    Any info regarding performance of either of these two emulators relative to an XDS510 would be appreciated.

  • Hi,

    In my experience and the other colleagues, we found out that:

    - The XDS560v2 shines for program loads of large binaries (as it has larger internal buffers) when compared to any other emulator, but this advantage is concentrated when used with C6000 and ARM cores - not much difference with C28x cores, as its executables are not very large anyways. Debugging operations (step, console I/O) are noticeably faster than a XDS100, but not by a severe margin when compared to the XDS200 - again, C28x cores do not show a large difference.

    - The XDS200 is about the same speed as a XDS510USB in load and debugging operations (step, console I/O), therefore it is a lot faster than a XDS100 for the C6000 and ARM cores. Again, the C28x cores show only a minor difference, mostly noted when doing step operations.

    We don't have any formal testing procedure, therefore no exact numbers are available.

    Hope this helps,

    Rafael

  • Great Rafael ! I was only looking for informal information, not exact test numbers.

    Cheers,

    Andrew

  • Rafael,

    This is exactly the info I was looking for.  I was worried that the XDS200 might be closer to the XDS100 than the XDS510, but you've convinced me to give it a try. 

    Thanks,

    John

  • Thank you very much Rafael.

    It is very helpful information to me.

     

    Best regards,

    Y.H Choi

  • One followup. Does an XDS200 serve as a license key to CCS in the same way that an XDS100 does?

    Thanks,

    Andrew

  • Andrew,

    The standalone XDS200 emulators require a full CCS license. A board with an integrated XDS200 can use the free license.

    You can think of the 200 as a replacement for the 510.

    Cheers,

    Rafael

  • Sorry to revive an old topic, but I stumbled across this remark:

    john filo said:
    My experience so far is that XDS100v2 is so slow as to be completely unusable

    I think that the XDS100v2 is getting unduely harsh treatment here:  the problem appears to be in no small part due to how CCS drives it.  For instance, it limits the bitrate to 3 MHz even though the hardware is capable of 30 MHz.  I've done some preliminary testing with some custom software driving the XDS100v2 and targeting a DM814x, and even though 30 MHz is far in excess of the max value specified in the datasheet (10 MHz), it appears to work reliably (at least when only ICEPick and one or more DAPs are in the chain).

    Of course quite a bit of that will be lost in overhead, mostly due to the (mis)design of DAP and the fact that the FT2232H is not very well suited to JTAG (a small microcontroller would be able to do a much better job), but especially for bulk transfers I still don't see any fundamental obstacles to getting at least to within an order of magnitude of the theoretical max speed.

    Which leaves the question, even acceping the 3 MHz limit, why on earth does it take CCS more than 20 seconds to upload and a whopping 40 seconds to download 50 KB of raw data, which means an average transfer rate of little over 10 kilobit/s, 0.34% of the JTAG bitrate?!

  • I wasn't necessarily blaming the hardware.  I was just relaying what I've observed using the LogicPD OMAP-L138's eval board's builtin in XDS100.  A cynic might say this performance is intentional; give users the bare minimum functionality for free, but make it so unbearably slow that anyone trying to get any work done will gladly purchase an emulator and a CCS license.  But the bottom line for me was that it was completely unusable.  I don't recall how long simple operations (loading a program, single stepping), but it was *painfully* slow.

  • Matthijs van Duin said:
     the problem appears to be in no small part due to how CCS drives it.

    Old thread. Improvements were made. Check:

    Matthijs van Duin said:
    (a small microcontroller would be able to do a much better job)

    XDS100 has been on the road for several years now, therefore there are changes ahead. Looking at the microcontroller world, for example, there are several new JTAG debuggers based off of their own microcontrollers (all the new TM4C and MSP430 launchpads and the new MSP-FET debugger, for example).

    --Cheers

  • john filo said:

    A cynic might say this performance is intentional; give users the bare minimum functionality for free, but make it so unbearably slow that anyone trying to get any work done will gladly purchase an emulator and a CCS license.

    The thought occurred to me too, though it may be prudent to apply Hanlon's razor here, especially since getting good performance out of an FT2232H-based JTAG adapter is non-trivial.  The interaction between the properties of all the layers involved (USB, FT2232H/MPSSE, JTAG, DAP) are such that if each layer gets a nice abstraction without considering the whole stack, basically all hope of good performance is already lost.

     

    176671 said:

     the problem appears to be in no small part due to how CCS drives it.

    Improvements were made.

    [/quote]

    I had tried a CCSv6 beta a while back but didn't notice any obvious performance improvements, but also didn't explicitly test for any.  I just installed the latest version and did some comparisons, same file (~50 KB) and target (DM814x):

    Upload went from 20 seconds to 8 seconds.

    Download went from 40 seconds to 6 seconds!

    So there's definitely been major improvement, and it's interesting to note that download has now become faster than upload, even though it used to be twice as slow.  Still, that's only 6.5 KB/s up and 8.5 KB/s down, or around 2% of the (artificially limited) 3 MHz JTAG bitrate.  In principle a bulk transfer should consist of back-to-back 32-bit DAP transfers (ignoring the occasional check for errors which shouldn't affect the average much) which take

    3 DAP command
    32 data transfer
    1 ICEPick (bypass register)
    4 JTAG overhead (0 run cycles)

    = 40 JTAG cycles total.  (In theory ICEPick can be commanded to temporarily hide itself from the chain, but that's probably more trouble than saving that one bit is worth.)  Especially at high clock rates the command processing overhead of the FT2232H MPSSE also becomes significant, which is 5/30 μs for a bitwise (partial byte) transfer and (5+n)/30 μs for an n-byte transfer. Assuming two bitwise transfers and one 4-byte transfer are used that yields 19/30 μs, which finally yields

    JTAG clock Max DAP throughput Efficiency
    3 MHz 280 KB/s 76%
    10 MHz 843 KB/s 69%
    30 MHz 1.94 MB/s 54%

    (As the JTAG clock rate goes down, MPSSE overhead becomes relatively negligible and efficiency converges to 32 bits / 40 cycles = 80%)

    So on one hand, *cheers* for the improvements done so far, but it seems there's still work to be done.

    (It is also possible of course that I'm overlooking something and reality isn't quite willing to conform itself to these calculations.  If I ever get around to finishing my tool I'll report back on that.)

     

    Out of curiosity, I also tried a similar up/download via the Cortex-A8 instead of DAP.  In either direction, my patience ran out after 5 minutes and I cancelled the transfer (by killing CCS since no "cancel" is actually available).  I tried a 1 KB download instead, this took an insane 21 seconds (!).  I don't know if it's the new CCS version or if some settings are wonky in my project, I don't think core-transfers have ever been this slow for me.

    Even if this bizarre slowness is some kind of settings-problem, transfers through the cortex-a8 always have been and always will be a lot slower than through DAP.  I'm fully aware of the horrible sequence of accesses required to transfer data through the core, but I'm not sure it's that obvious to all users.  In many screenshots that float around on wiki I don't even see DAP listed as debug target, which means people will be doing data transfers through the core instead.  In some cases CCS makes using DAP impossible, e.g. for the transfers used for variables and disassembly when single-stepping the core.

    Of course access through DAP has its own subtleties (especially when MMU/cache is enabled), ideally CCS would be aware of the interconnect topology and (unless explicitly told otherwise) use the fastest way to get data there, regardless of which memory space is formally being used (e.g. if viewing the A8 do an MMU lookup and use DAP on the physical address whenever safely possible), or -- especially since the core will be halted anyway -- load a debug monitor onto the core that can assist in performing a fast transfer.

     

    176671 said:

    (a small microcontroller would be able to do a much better job)

    XDS100 has been on the road for several years now, therefore there are changes ahead. Looking at the microcontroller world, for example, there are several new JTAG debuggers based off of their own microcontrollers (all the new TM4C and MSP430 launchpads and the new MSP-FET debugger, for example).

    [/quote]

    I'm not sure why every μC series seems to need its own JTAG debugger, but I'm all in favor of a microcontroller-based low-cost debugger for the big SoCs as well.  I hope it will still have the Xilinx CoolRunner-II CPLD that's on the XDS100v2 though, I've come to appreciate that little chip and its very tolerant electrical properties a great deal.  For comparison, we have an STM32 processor on an extension board, and an ST-LINK/V2 debugger to go with it (again, not sure why every μC series seems to want its own debugger, but oh well).  If I try to hotplug it, regardless of whether I connect USB first or JTAG first, the processor resets. *sigh*  If I use the XDS100v2 instead (and some software which doesn't care when you mix&match stuff from different vendors, e.g. OpenOCD), no problem hotplugging whatsoever.  I also commonly leave the XDS100v2 attached to the target while not connected to USB, while I only later realized that many chips really don't appreciate getting a voltage on their I/Os while being unpowered themselves -- the CPLD however doesn't care, you're pretty much allowed to power/unpower any part of it in any order/combination.  Very nice.

  • Der Matthijs,

    Would you elaborate how you have driven the XDSv100 faster?
    We're going to evaluate AM335x, and would of course like a bit better performance, without spending ~1k for the XDDDSv560.

    Best regards, Florian