This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

MSP430F2272: reflow guidelines for pre-programmed devices, seeing flash instability post reflow

Part Number: MSP430F2272
Other Parts Discussed in Thread: TPS3808

Hi,

a customer using the MSP430F2272 is pre-programming the devices via a known process at one of the distributors. They are seeing occasional corrupted FLASH after the reflow.

Are there any reflow guidelines for pre-programmed devices?

Thanks!

--Gunter

  • Hi Gunter,

    I understand you are talking about proper reflow soldering, correct?
    My assumption is that the issues your customer is seeing aren't due to the reflow soldering but to ESD events, maybe caused by the machinery placing the devices on the PCB. Please check with your customer if ESD events can be prevented and let me know if the issues are still present after that.

    Best regards,
    Britta
  • Hi Britta,

    I'm the customer to which Gunter referred. The parts in question are built on high speed pick and place machines, which are designed for electronics assembly. The contracted manufacturer has an ISO system which includes ESD practice, and is regularly audited. To my knowledge, they have an excellent if not spotless record for compliance, so I don't believe that a systematic failure in ESD Protection at SMT assembly is the likely root of my issue. As further colloquial evidence of this, there are many other devices on this same assembly side that are much more sensitive to ESD events with no field failures.

    The reason for my concern is that the MSP430F2272's are pre-programmed at distribution, then they are reflow soldered. When they arrive at my plant for final assembly, we re-write everything in flash except for one sector of flash that decides if it can boot the other sector or if it has to copy an onboard backup before booting. We have seen some examples of memory corruption in this area of flash in units deployed in the field.

    As a bit of background, on what I mean by in the field: The units are tested here on through 4 distinct processes one at -45C, one at +70C.The units are subjected to network insertion and removal though multiple power cycles. Our end customer re-tests the units on reception and prior to deployment. The units that fail do so after all the afore mentioned testing is complete and the unit is deployed in the middle of nowhere. All units were reflow soldered between April and June, 2018 and would have been initially programmed at distribution no sooner than December 2017. There appears to be no lot, or serial number(issued sequentially) correlation in the failure.

    I'm aware of the issue raised here:

    https://e2e.ti.com/support/microcontrollers/msp430/f/166/p/60886/861606

    My circuit for the MSP430F2272 has a TPS3808 reset monitor, which seems to be the fix for issue raised.

    At this point, I'm trying to rule out all possible sources of this corruption. I understand that pre-programming is a common practice, however what I want to know is does TI's specified data retention still apply after a single IR reflow cycle or must we re-write the flash contents to refresh the retention?

    as a follow up to that, does this also apply to TLV structures in INFOA?

    I didn't notice an appnote or a specification in the family user guide or the data sheet for flash retention after reflow, if you could direct me to one, that might be enough.

    Thank-you for your consideration in this matter,

    -Jesse

  • Hi Jesse,

    thank you for the detailed description of your flow and the issue you're seeing.

    I have some additional questions which will help me to troubleshoot the root cause of this issue:

    1) Does your application include high temperature during operation in the field?

    2) Memory corruption: Does this affect single bits, bytes or words or an extended range in the memory? Do you see a flip from 0 (non-corrupted) to 1(corrupted) or the other way round?

    3) How are the devices pre-programmed at distribution?

    I'm asking this as one reason for your issue could be that the data bits are not fully erased in the first place which could result in worse data retention.

    You could have a look at the Understanding MSP430 Flash Data Retention Application Report which discusses data retention in Flash memory and its link to higher temperature. In general terms you can say that data retention decreases with increasing temperature.

    To your general question: Neither TLV structure nor the data should need to be re-programmed after reflow but it could be that a combination of short erase time, reflow soldering and high operation temperature could massively decrease data retention and show failure in the field after a certain time.

    Another option might be malfunctioning of the application due to EMV, overclocking of the CPU among others which can cause unintentional writes to Flash (in case such Flash write routines are implemented in your code). Additional safety measures against unintentional write access can prevent this though.

    Could you check the questions above and also review your code with regards to the Flash write operations and come back to me with the information?

    Thanks and best regards,

    Britta

  • Hi Britta,

    1) My application is outdoors, so I do see a relatively large temperature range.  The system consists of two primary types of unit, one is a power receiving unit, one is a power delivery unit, the power receiving units out number the power transmitting units by a 7:1 ratio.The power receiving units in this particular installation run between 30C and 70C inside the box, as measured by the internal temperatures sensor in the MSP430. The power delivery units run up to 30C higher, and max out at near 100C. We do a one point calibration on all units at room temperature before the unit leaves the plant, so the measurement is not terribly accurate, but it gives a decent ball park approximation as there are thousands of units reading similar numbers. That said, I can make our test equipment here in the lab at 20C fail as well, so I'm not overly hopeful that application temperature is the root of the issue.

    2) Let me preface my next statement by saying that reading out and processing all the failed units is a very time consuming process, so I only have a representative sample of the total number of failed units. In every case, the bit goes from 0->1, in most cases, the failure is single bits in multiple words, usually less than 10 words per sector, occasionally most of a sector is 1. In every case of the pre-programmed memory location corruption that I've logged (only one sector) its single bits in multiple words. All failures in this region that I've noticed so far have been in the power receiving boxes.

    3) The devices are pre-programmed using a BPM microsystems programmer (www.bpmmicro.com) I'm not sure which model, but given how long a batch takes to program, I would expect a 3xxx or 4xxx model, if it made a difference, I could probably get detail down to the serial number of the unit. BPM claims that they follow SLAU320AD for their algorithm.

    My The parts are purchased new from TI through our franchised distributor, I'm not really sure what the default ship state of the memory is, I assume erased? I'm not aware of exactly how the BPM algorithm works, so I can't comment on if another erase is done before writing, or if the units are just blank checked. We haven't ever re-programmed units before, so I'm unable to comment on the efficacy of the erase process directly.

    The application runs at 8MHz, and uses a serial port to relay telemetry information to a network processor, so if the clock were not 8MHz, the network processor would not be able to receive the telemetry data, as its not auto baud rate, so I don't think over clock rate is likely. We also had suspicions about internal flash write operations, even though we only ever write to INFO B in application, but nevertheless we stubbed out the part of the code that does all flash erase and all flash write, without success. We then removed the network processor's JTAG port to the micro, that didn't help. We removed the boot loader's ability to copy images, and that didn't help. All that removed, nothing in the device could programmatically write flash and we still end up with failures. My units deployed in the field currently have the internal flash erase and write blocks stubbed out. I did not roll out the release that blocks the network processors JTAG port or the bootloader, as either could potentially have potentially catastrophic consequences to my end customer. I'm not sure what you mean by EMV.

    Thank-you for the flash retention application report, given its findings, its not likely that reflow soldering is my issue. Your response fully answers my inquiry.

  • Hi Jesse,

    thank you for sharing the detailed information with me.

    Based on that quite some of my suspicions are ruled out.

    I see you've already done a lot of work to rule out most of the scenarios. It might be worth to re-look at the "erase" process, as you see Bits flip from 0 to 1 only. In case the erase wasn't given enough time it might be that the cells are not charged to the full extend and start slowly discharging over time and temperature which then at some point results in the flipped Bits.

    Jesse Goetting said:
    Your response fully answers my inquiry.

    Do you want to have further assistance on this issue or does this mean that I should go ahead and close this thread. You can of course always open a new one in case you have further questions.

    Thanks and best regards,

    Britta

  • You can close this thread as resolved. I'm currently involved in another thread regarding the JTAG process, hopefully its as fruitful as this one has been. Thank-you for your assistance.

  • Thanks for the positive feedback!
    Let me know in case there is anything else I could help you with.

    Best regards,
    Britta

**Attention** This is a public forum