This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

MSP430 HArdware MPY with 32 bit result

Genius 4170 points

Hello,

i am wondering if there is a possibility to directly save a 32 bit result into a long ( 32bit) variable.

The reaons why I ask is the follwoing:

I have a serious timing issue on my application. I am reading out a detector array with 512 pixels.

For every Pixel i am doing an AD-conversion which right now takes about 13 ADC-clock cycles of 5 MHzw hich is around 2,6 µs.

The timing for the detector array read out process is set to 250 kHz each Pixel, so i have about 4µs to do an conversion and afterwords some mathematics like MPY untill the new Pixel is to be read out.

Do the math: 4µs - 2,6 µs = 1,4 µs

In this time i am doing some MPY action, which, regarding the datasheet, sould not take longer than max. 7x MCLK.

Here is part of my Code:

    do
    {
        while ((ADC12IFG & ADC12IFG2) == 0);
        TEMP = ADC12MEM2;
        MPY = TEMP;
        OP2 = array[Schleifenzaehler];
        
        //uint32_detector_array[Schleifenzaehler] = RESHI;
        //uint32_detector_array[Schleifenzaehler] = uint32_detector_array[Schleifenzaehler]<<16


        uint32_detector_array[Schleifenzaehler] |= RESLO;
        Schleifenzaehler ++;
    }while (Schleifenzaehler < 513);

As you migth realize i already tried out some different settnigs for a 32 bit result. Now i have meassured the Clock cycles in Debug mode via CCS:

With the above setting I get around 29 MCLK cycles = 1,16µs

When I comment in my 32bit shifting operation with RESHI and << operator it will take 44 MCLK cycles = 1,6 µs, which theoretically wont fit into my spare time of 1,4 µs.

So finally coming to my question, is there any way of doing this calculation directly into an 32 bit register so i do not have to do the shifting?

Is the cycle counter in CCS trutable or should I assume that bitshifting is way faster then i actually think it is?

There is the option to longer my timings, but I would rather like to do faster math of course.

Thanks for reading and helping.

Best wishes, Seb

  • Hi,

    seb said:
    So finally coming to my question, is there any way of doing this calculation directly into an 32 bit register so i do not have to do the shifting?

    I suggest you to use structures and unions to solve this. 

    See the example below.

    typedef union object
      {
         long int long_value;
          struct value
            {
              int high;
              int low;
            } MYVALUE;
      } MYOBJECT;

    MYOBJECT a;

    void main(void)
    {

     a.MYVALUE.high=10;
     a.MYVALUE.low =0;
     a.long_value = 0x12345678;

    }

    Best regards,

    AES

  • Thanks a lot for that smart tip with example.

    I consider myself still as a newbie in C programming. Soemtime ago i read about structures and unions which use the same physical memory placement in the RAM.

    But it would take me a lot longer to create one by myself, so thank you very much.

    Fortunatly i found out that i dont have 2 µs for solving but 4 µs so my assmued longer lasting Bitshifting for the 23bit result will do it anyway.

    But when i got time i will try your solution, as i am sure i will get i to work, i am looking forward to counting the clock cycles :)

    Greetings, Seb.

  • A trick to speed up things is to precalculate the storage positions.

    unsigned int *ptr;
    do{
      ptr = (unsigned int *)(uint_32_detector_array+Schleifenzaehler);
      ...
      *(ptr++)=RESLO;
      *ptr=RESHI;
      ...
    }while...

    It removes all the array index arithmetic and shifting. The two assignments are in best case just two single assembly instructions. 2*5 cycles for the move plus the initial pointer calculation (copy counter, multiply ounter with size of long == shift twice, add array start position = 5 cycles).

    (I didn't test it, even didn't try to compile it, but it should work)

    But it's possible that the compiler already did this optimization internally.

    I see that in your code, it already takes only 15 cycles to do the assignments.
    But maybe I can drop it to 10 cycles:

      ptr = (unsigned int *)uint_32_detector_array;
    unsigned int *ptr;
    do{
      ...
      *(ptr++)=RESLO;
      *(ptr++)=RESHI;
      ...
    }while...

    The ptr++ operation on short ints or bytes can be doen without additional costs (using the postincrement addressig mode).
    And since ptr is just incremented twice by 2 bytes in each loop, starting at the beginning of the array, the calculation of the target positions is 'free' and only the 2*5 cycles for the move are required (on devices with the older MSP430 core, it takes 2*6 cycles, only the MSP430X core can do a move in 5 cycles)

    If you're asking why it taks so long:

    One cycle reading the instruction, two cycles for reading source and destination address, two cycles for reading and writing the value.

**Attention** This is a public forum