LMK04828: LMK04828B: Part-to-part output skew when two devices share the same reference / input-to-output skew in zero-delay mode

Part Number: LMK04828
Other Parts Discussed in Thread: LMK04832,

I'm will be using several LMK04828B devices (on CLK104 module) driven from the same reference clock and need to bound the skew between an output on device A and the equivalent output on device B.

The datasheet's "Clock Skew and Delay" table only specifies on-chip, output-to-output skew (|TSKEW| = 25 ps max same pair / 50 ps max any pair, same format, at fCLK = 245.76 MHz). I don't see a part-to-part (device-to-device) output skew spec, nor a general input-to-output skew figure for zero-delay mode — only the single tP CLKin0 → SDCLKout1 = 0.65 ns entry for one specific configuration.

Could you help with the following:

  1. Does TI publish, or can you provide, a typical and worst-case part-to-part output skew for two LMK04828B devices at the same VCC and temperature, driven from a common reference?
  2. For zero-delay mode (with internal feedback), is there a specified or characterized input-to-output skew (and its part-to-part variation), including the contribution from phase-detector phase-error variation?
  3. If these aren't specified, what's the recommended way to bound or minimize device-to-device skew — e.g., matching feedback paths, output format, divider/analog-delay settings, SYNC procedure — and what residual should I budget for?
  • Hi Gilles,

    Does TI publish, or can you provide, a typical and worst-case part-to-part output skew for two LMK04828B devices at the same VCC and temperature, driven from a common reference?

    Unfortunately, we do not specify this. We recommend multi-device synchronization to reduce any part-to-part skew, which does so by phase aligning the rising edges of all output dividers.

    For zero-delay mode (with internal feedback), is there a specified or characterized input-to-output skew (and its part-to-part variation), including the contribution from phase-detector phase-error variation?

    You can refer to the 650ps typical propagation delay figure.

    If these aren't specified, what's the recommended way to bound or minimize device-to-device skew — e.g., matching feedback paths, output format, divider/analog-delay settings, SYNC procedure — and what residual should I budget for?

    A SYNC event phase aligns all outputs. The setup-and-hold time (no more than one half of the period of the VCO) must be met - which can impose a fairly challenging obstacle to multi-device synchronization. The LMK04832 offers PLL1 R divider SYNC, which only requires a 100ns setup and hold time, which is much easier to meet.  Zero-delay mode (nested ZDM, specifically) will guarantee a deterministic phase relationship between the looped back output and the input to the device. There should still be budgeting for a skew/drift of roughly ~7.5ps in between power cycles.

    Thanks,

    Michael

  • Hi Michael,

    Thank you for the earlier response — the ~7.5 ps post-power-cycle drift figure and the confirmation that nested zero-delay mode guarantees a deterministic input-to-output phase relationship were both very helpful.

    I'd like to follow up specifically on the setup/hold requirement for the divider-reset SYNC, because it drives a multi-board architecture decision.

    Context / constraints:

    • I'm using the AMD CLK104 module (LMK04828B, U2) on ZCU208 boards, 8 boards total, driven from a common 10 MHz reference distributed to each board.
    • On the stock CLK104, CLKin0 is occupied by the 10 MHz INPUT_REF_CLK (via the J11 SMA), and CLKin1 by the on-board TCXO. That leaves the CMOS SYNC pin (via the J42 SMA) as the only available input for the external SYNC/divider-reset signal — I cannot use the high-speed CLKin0 path you'd normally recommend for SYNC, since it's taken by the reference.
    • The LMK04828B is run in nested 0-delay mode, PLL2 locked, VCO at 2.4576 GHz (so ½ VCO period ≈ 200 ps).
    • The goal is a deterministic, repeatable divider reset that is coherent across all 8 boards, so that after MTS the inter-board sample-clock phase alignment holds to a tight budget.

    Questions:

    1. You noted the SYNC setup/hold must be met to "no more than one half of the period of the VCO." For the LMK04828B specifically, is that ½-VCO-period (~200 ps) setup/hold referenced to the CMOS SYNC pin the correct number to design to, or is there a different figure when SYNC arrives on the CMOS pin versus CLKin0?
    2. Does any re-clocking configuration relax this? I've seen the SYSREF re-clocking / D-flip-flop mechanism described in the multi-LMK sync app notes (e.g., SYSREF_MUX = re-clocked, SYNC_1SHOT_EN). My question is whether re-clocking the incoming SYNC by the SYSREF divider actually relaxes the input setup/hold requirement (e.g., to the SYSREF-divider period rather than the VCO period), or whether the setup/hold to the internal capturing clock still has to be met at the ½-VCO level regardless — with re-clocking only affecting output determinism, not the input timing window. I want to make sure I don't design around a relaxation that doesn't actually exist.
    3. Practically, how do multi-board ZCU208/CLK104 MTS designs meet the ½-VCO SYNC setup/hold on the CMOS SYNC pin? Since the SYNC is asynchronous to each board's VCO, is the intended approach to (a) generate the SYNC coherently with the 10 MHz reference and place its edge to maximize setup/hold margin, (b) accept that some boards may reset one VCO cycle off and rely on a subsequent step to resolve it, or (c) something else? If the edge must be placed relative to the reference, what alignment tolerance is acceptable?
    4. How many SYNC pulses are needed for a deterministic reset in a one-shot configuration (SYNC_1SHOT_EN = 1) — a single edge, or is a multi-pulse sequence recommended for robustness across devices? And does the SYSREF divider need to be synchronized separately from the device-clock dividers, or does a single SYNC event handle both?
    5. Reconciling the two competing needs on the LMK04828B. I understand the LMK04832 offers an easier PLL1 R-divider SYNC (100 ns setup/hold), but it lacks analog delay on the DCLK outputs — which I need for per-output fine skew trim to meet the inter-board budget. The LMK04828B gives me the DCLK analog delay (25 ps step) but appears to have the harder ½-VCO-period SYNC requirement on the CMOS pin. Is there a way to get both on the LMK04828B — i.e., a SYNC/divider-reset approach that achieves deterministic multi-board reset without requiring the ~200 ps setup/hold on the CMOS SYNC pin, while retaining DCLK analog delay for skew trimming? For example, does using the SYSREF-request / continuous-SYSREF path, or a particular SYNC_MODE/SYSREF_MUX configuration, provide an easier timing path while keeping DCLK ADLY available?

    Any measured or characterized data on device-to-device divider-reset repeatability under a common reference would also be very helpful for our timing budget.

    Thanks again,
    Gilles

  • Hi Gilles,

    I will get to your questions tomorrow.

    Thanks,

    Michael

  • Does TI publish, or can you provide, a typical and worst-case part-to-part output skew for two LMK04828B devices at the same VCC and temperature, driven from a common reference?
    For zero-delay mode (with internal feedback), is there a specified or characterized input-to-output skew (and its part-to-part variation), including the contribution from phase-detector phase-error variation?

    It's possible to give typical and worst case skew in distribution mode. Comprehensively specifying even typical skew, let alone worst-case skew, for non-distribution mode uses is impossible:

    • Different input and output frequencies will necessarily result in different perceived input-to-output skew. In principle we could provide one example case and some frequency coefficient of propagation delay, but it would need restrictions on R and N divider values that limit its usefulness.
    • In zero-delay mode, if max(N, R) % min(N, R) != 0, it is not possible to establish a single deterministic phase relationship between the input and the output.
    • Even when max(N, R) % min(N, R) == 0 in zero-delay mode, or when zero-delay mode is not used, output dividers not in the feedback path are challenging to align repeatably when the clock distribution path is several GHz. There are options to re-time the SYNC even to the SYSREF divider, though this requires putting the SYSREF divider in the feedback path in zero-delay mode which can limit performance if the SYSREF divider frequency is low.

    If we can restrict the configuration to a known configuration with a single deterministic input-to-output phase, we can probably offer some typical skew numbers for that condition. Worst case is difficult for a different reason due to significant phase detector timing variation across VTUNE voltage (something like 2ns/V) - we could restrict this too, but because VTUNE is a degree of freedom on the PLL over which the user has limited control, I don't know how helpful it would be.

    If these aren't specified, what's the recommended way to bound or minimize device-to-device skew — e.g., matching feedback paths, output format, divider/analog-delay settings, SYNC procedure — and what residual should I budget for?
    • Matching feedback paths is a good idea. Different feedback paths have different lengths, even if their divider settings are otherwise identical.
    • Matching output format is mostly unimportant - all output drivers are built with BJTs that share similar tempcos. If this were LMK04832, I would suggest staying away from LVCMOS output format, as this has a dramatically different tempco to the other output formats.
    • Divider settings mostly don't matter, since the output of the divider block is re-timed by the clock distribution path in most cases. There are a handful of exceptions, such as when the divider is bypassed or the duty cycle correction is enabled; in those cases, all outputs in all feedback paths should share the same bypass/DCC options.
    • Digital delay settings for the feedback path in zero-delay mode are not relevant, since the device will treat the delay as phase error and zero it out at the phase detector. The digital delay of other outputs is still relevant, and the digital delay of the feedback path in zero-delay mode can adjust the phase of the other outputs relative to the feedback path. I recommend keeping the digital delay settings constant across devices.
    • Analog delay on the clock outputs is not very consistent device-to-device, especially across temperature - the variation in step size across temperature is about three times larger than the step size. I would not recommend using analog delay on the clock outputs for precise input-to-output matching across multiple devices unless per-device calibration and temperature compensation can be afforded.
    • Arguably, keeping the SYNC procedure identical across devices is less important than designing a configuration in which the SYNC procedure is trivialized. For instance, zero-delay mode configurations bringing the SYSREF divider into the feedback loop can also re-time the SYNC event to the SYSREF divider edge, which can be correlated to the reference clock. A SYNC event could be issued to any device in the multi-device network with a frequency of GCD(input, output[0..13]), and the valid window is relaxed to almost a full SYSREF divider period (less some small amount of setup/hold time on the order of 200ps around the SYSREF divider rising edge).
    • Of course, if temperature and voltage can be held close to constant across systems, this eliminates drift due to these sources; while propagation delay variation as a function of voltage is small enough to be negligible, propagation delay variation as a function of temperature tends to be around 1ps/°C. All devices should have very similar tempco, and the sign never flips; so if temperature of the system increases or decreases, as long as the increase or decrease can be spread across the whole system, multiple devices should closely track each other.
    • VTUNE voltage variation at the charge pump output (and for PLL2, VCO calibration results) greatly impacts phase alignment. The parasitic capacitance across the FETs in the charge pump output result in an injected charge during FET turn-on and turn-off, which happens regularly at the phase detector frequency during pump-up and pump-down actions. The magnitude of this injected charge can become imbalanced between pump-up and pump-down cycles as VTUNE deviates from 50% of full-scale, which looks like an effective increase in up or down on-time. The effective increase in up or down on-time causes the R and N ports of the phase detector to skew away from each other to balance the total integrated up and down on-time while locked. On PLL2, the effect size is close to 2ns/V. On PLL1, the effect size is likely smaller due to lower charge pump currents requiring smaller FETs (I have not had a chance to measure it comprehensively), but the much-smaller overall VTUNE voltage variation required to keep the VCXO locked over temperature diminishes the impact.

      Making matters worse, the calibration routine for the LMK04828 PLL2 integrated VCO is not deterministic, which can lead to discrete ~200ps variations in input-to-output phase alignment between recalibrations at identical conditions. To briefly explain: the integrated VCOs in PLL2 include a programmable varactor capacitance to set the center frequency, allowing for higher Q and better phase noise at the locked frequency without sacrificing a wider overall VCO range. The digital capacitance codes (capcodes) which program the varactor have frequency regions that heavily overlap - each step only moves the center frequency by a few MHz - and each capcode gives the VCO enough overall frequency range to maintain lock across ±125°C temperature variation in L/C parameters for the VCO tank circuit. A calibration routine is performed each time the PLL2 N-divider LSBs are written, which iterates through capcodes and monitors the VTUNE voltage until a value is found which achieves approximately half of the full-scale VTUNE output. Because the capcodes overlap, sometimes more than one capcode can satisfy the calibration conditions, even at the same temperature. But since each capcode offsets the VCO center frequency by several MHz, the VTUNE voltage is offset by around 100mV - resulting in around 200ps difference in phase detector timing due to the previously described charge injection phenomenon. 

      Because the capcodes and the inductor in the VCO have some process variation, a single capcode can't be picked for all devices operating at the same frequency. Moreover, locked devices with different capcodes may still have a few tens of mV difference in VTUNE voltage, which can lead to ~100ps skew between devices even in identical circumstances. And since the calibration is in some cases nondeterministic even at identical conditions, that skew could be as much as 500ps, even if it was measured once as ~100ps initially. PLL2-only or cascaded zero-delay mode configurations are all subject to this substantial variation. There is a way to override the PLL2 capcode and skip the calibration procedure to make this variation constant between power cycles, which I can describe if it is of interest. 

      PLL1 (nested) zero-delay mode configurations are still subject to some variation from this charge injection mechanism, but because VCXO frequency stability across temperature is much better and the centering at the same temperature is generally much closer than with the integrated VCO of PLL2, the magnitude of the observed effect in PLL1 tends to be much smaller, and is mostly consistent between power cycles (with some allowance for VCXO aging, on the scale of years). Remember as well that the VTUNE voltage of the VCXO changes proportionally much less to achieve a greater frequency variation at PLL2 in nested zero-delay mode, and that the phase of the outputs with respect to the input is governed by the PLL1 phase detector relationship, independent of the PLL2 phase detector timing.

    So, if we are to synthesize some meaningful guidance for a real configuration:

    • Use a nested zero-delay mode configuration with a stable VCXO
    • Feed the SYSREF divider back if at all possible, to allow for re-timing the SYNC event to the SYSREF divider edge, which can then be correlated to a reference edge
    • Ensure that max(N, R) % min(N, R) == 0, so that only a single input-to-output phase relationship exists, and so that the SYNC event timing to align all other clock outputs is trivialized.

    I've included some flowcharts that describe the different ways that synchronization can be achieved for the LMK0482x and LMK04832 family of devices. If you can't satisfy the summary guidance above for whatever reason, the flowcharts are a good place to check for palatable alternatives.

    SYNC flowchart.pdf

    I will also caution against attempting to use the SYNC pin for any tight timing alignments, since the SYNC pin delay varies by as much as 2ns across temperature (a consequence of a CMOS pin implementation) vs CLKIN0 where the delay variation is less than 150ps across temperature (a CML implementation). Re-timing the SYNC event to the SYSREF divider is much simpler to satisfy given your constraints.

    I think I've given you something that answers or addresses the first three questions back to Michael, roughly summarized as "don't try to SYNC against the GHz VCO with the SYNC pin, use the SYSREF divider retimer". Let me also address the last two questions.

    How many SYNC pulses are needed for a deterministic reset in a one-shot configuration (SYNC_1SHOT_EN = 1) — a single edge, or is a multi-pulse sequence recommended for robustness across devices? And does the SYSREF divider need to be synchronized separately from the device-clock dividers, or does a single SYNC event handle both?

    Just a single edge is sufficient. Note that the one-shot is after the SYSREF_MUX, so if the SYNC event is to be re-timed by the SYSREF divider, the SYNC pin needs to be asserted for at least one VCO cycle after the SYSREF divider edge internally; if it's easier, you can just raise the pins high across all devices at some point for longer than one SYSREF divider cycle, and the one-shot will take care of the timing for the divider resets.

    Resetting the SYSREF divider can be done simultaneously, but if the SYSREF divider is in the feedback path, then it serves no purpose to reset the divider, since the phase of the divider will be restored to align with the reference input immediately afterward. Since the SYNC event can be re-timed to the SYSREF divider, it is still possible to establish a deterministic timing between the SYSREF divider edge and all other clock outputs without resetting the SYSREF divider.

    If the SYSREF divider is not in the feedback path, you could of course reset it with the other dividers simultaneously; but you'd have to align the SYNC event against the clock distribution path, which I've tried to caution is only really deterministic with CLKIN0 and very tight timing on the SYNC generator.

    Reconciling the two competing needs on the LMK04828B. I understand the LMK04832 offers an easier PLL1 R-divider SYNC (100 ns setup/hold), but it lacks analog delay on the DCLK outputs — which I need for per-output fine skew trim to meet the inter-board budget. The LMK04828B gives me the DCLK analog delay (25 ps step) but appears to have the harder ½-VCO-period SYNC requirement on the CMOS pin. Is there a way to get both on the LMK04828B — i.e., a SYNC/divider-reset approach that achieves deterministic multi-board reset without requiring the ~200 ps setup/hold on the CMOS SYNC pin, while retaining DCLK analog delay for skew trimming? For example, does using the SYSREF-request / continuous-SYSREF path, or a particular SYNC_MODE/SYSREF_MUX configuration, provide an easier timing path while keeping DCLK ADLY available?

    As I mentioned above, I think those analog delays are not going to be nearly as consistent as you want, unless you can afford per-device calibration and temperature compensation. It also increases the noise floor of the output clocks by several dB. If you are okay with the tradeoffs, I think re-timing the SYNC even to the SYSREF divider with the SYSREF divider fed back in zero-delay mode would mostly solve your problem.

  • Thank you — the SYSREF-divider retimer guidance and the flowcharts are exactly what we needed, and we are proceeding in that direction: nested zero-delay mode, SYSREF divider as the 0-delay feedback, SYNC re-timed to the SYSREF-divider edge. Working through the details against the flowchart raised five follow-up questions.

    Our configuration, for reference:

    • LMK04828B (on AMD CLK104), nested zero-delay mode
    • PLL1: CLKin0 = 10 MHz external reference; VCXO (OSCin) = 160 MHz; PLL1 PD = 10 MHz
    • PLL2: VCO = 2500 MHz (R2 = 8, N2 = 125, PD = 20 MHz)
    • Outputs: device clocks 156.25 MHz (VCO ÷16) on CLKout0/4/6/8/12; SYSREF 9.765625 MHz (VCO ÷256) on CLKout3/9/11; all HSDS8
    • 0-delay feedback currently CLKout8 (156.25 MHz); planning to move it to the SYSREF divider per your guidance

    Q1 — Reconciling the flowchart with the SYSREF-retimer guidance (our critical question).

    Walking our configuration through the LMK0482x flowchart (page 2): we satisfy "fOUT % fZDM == 0 for all fOUT" with fZDM = SYSREF = 9.765625 MHz (156.25/9.765625 = 16). But feeding the SYSREF divider back to PLL1 requires PLL1 dividers satisfying 10/R1 = 9.765625/N1, whose minimal solution is R1 = 128, N1 = 125 (PFD = 78.125 kHz) — and 125/128 are coprime, so max(N,R) % min(N,R) ≠ 0, giving (per the flowchart) min(N,R)/GCD(N,R) = 125 possible input-to-output phase relationships, with a non-resettable N-divider and no R-divider SYNC on the LMK0482x. This appears fundamental to our frequency set: SYSREF = VCO/256 lies in the power-of-2-submultiple family of the 2500 MHz VCO, while the 10 MHz reference (and any VCXO lockable to it) lies in the ×10 MHz family; the two share no usable common divider structure.

    Yet your written guidance says the SYSREF-divider feedback re-times the SYNC to the SYSREF-divider edge with a valid window of almost a full SYSREF period, issued at GCD(input, outputs) — which for us is 78.125 kHz. Can you reconcile these? Specifically:

    (a) Is the SYSREF-divider SYNC-retime mechanism valid despite the 125-phase ambiguity — e.g., because the ambiguity is granular in whole SYSREF periods (102.4 ns = exactly 256 sample periods at 2.5 GSPS) and therefore acceptable or correctable downstream — or does the ambiguity defeat multi-device determinism?

    (b) Which flowchart leaf/Type does our configuration actually land on?

    (c) If this frequency set cannot achieve a deterministic relaxed-window SYNC on the LMK04828B, one option we're considering is moving the distributed reference itself into the power-of-2 family: CLKin0 = 156.25 MHz (= VCO/16, generated centrally by our clock-distribution unit from its OCXO) with VCXO = 156.25 MHz. That gives, with SYSREF-divider feedback: PLL1 R1 = 16, N1 = 1 (PFD = 9.765625 MHz), max(N,R) % min(N,R) = 0, and — since the flowchart identifies the non-resettable N-divider as the source of phase randomization — N1 = 1 should eliminate the input-to-output ambiguity entirely. PLL2 becomes R2 = 1, N2 = 16 (PD = 156.25 MHz), also satisfying your condition. SYNC would be issued at GCD = 9.765625 MHz with the ~102.4 ns window. Does this configuration land on a deterministic leaf and achieve the relaxed window — and do you see any issue with the R1 = 16 (R > 1) side given the LMK0482x's lack of R-divider SYNC, or is that concern eliminated when N1 = 1? If you would recommend a different change instead (different SYSREF frequency, different feedback/SYNC structure), we're open to it.

    Q2 — Scope of the max(N,R) % min(N,R) == 0 condition, and a planned VCXO change.

    Our PLL2 pair is N2/R2 = 125/8, which is coprime, failing your condition on the PLL2 loop. Our PLL1 pair (10 MHz → 160 MHz, effectively 16/1) passes, and the composite 10 MHz → 2500 MHz ratio is exactly 250. Does the condition need to hold on the PLL2 pair specifically, on PLL1, or on the composite, for single-phase determinism in nested ZDM with the SYSREF divider fed back?

    If PLL2 binds, we plan to replace the CLK104 VCXO — with 100 MHz if we keep the 10 MHz reference (giving R2 = 1, N2 = 25, PLL2 PD = 100 MHz, and PLL1 N = 10 with PD unchanged at 10 MHz), or with 156.25 MHz under the Q1(c) configuration (R2 = 1, N2 = 16, PD = 156.25 MHz). Two loop-filter questions follow: (a) PLL1 uses external loop-filter components on CPout1, sized for the current 160 MHz VCXO's tuning gain — these presumably need re-sizing for a new VCXO with different Kv. What's the recommended method/tool for redesigning the PLL1 external filter (Clock Design Tool? PLLatinum Sim? a worksheet?), and can you advise a starting point given a target PLL1 bandwidth in the usual 10–300 Hz jitter-cleaning range? (b) PLL2's loop filter is internal/register-programmable — what internal filter settings (R3/R4/C3/C4) would you recommend at the new PD (100 or 156.25 MHz)? If it helps, we can pull the CLK104's populated PLL1 filter values from the board schematic as the baseline.

    Q3 — Constraints on SYSREF-sourced outputs with the SYSREF divider fed back.

    (Contingent on Q1 resolving favorably.) With the SYSREF divider selected via FB_MUX as the internal 0-delay feedback in nested mode — the CLK104 provides no route to FBCLKin, so the feedback must be internal — does that selection place any constraints on the SYSREF-sourced outputs (CLKout3/9/11), e.g., divider or delay settings that must be held common with the feedback path?

    Q4 — Phase-noise cost of the 9.765625 MHz feedback.

    You cautioned that a low SYSREF-divider frequency in the feedback path can limit performance. Moving the feedback from CLKout8 (156.25 MHz) to the SYSREF divider (9.765625 MHz) drops the feedback frequency ~16×. Can you characterize or bound the phase-noise/jitter penalty for our configuration (nested ZDM, 2500 MHz VCO, with the VCXO options above)? Is there a feedback frequency below which you would advise against this approach?

    Q5 — PLL2 capcode override.

    Yes, please describe the capcode override procedure you offered. Our inter-board calibration relies on each board's input-to-output phase being repeatable across power cycles (we calibrate per-board residual skew via downstream measurement of the RFSoC DAC output phase), so making the ~200 ps capcode variation static is valuable even in nested mode. Specifically: which registers set the manual capcode; how should the correct capcode be determined per device (per-unit characterization, or one value per frequency/temperature?); and does skipping calibration carry lock-reliability risk over a 25–45 °C operating range?

  • Yet your written guidance says the SYSREF-divider feedback re-times the SYNC to the SYSREF-divider edge with a valid window of almost a full SYSREF period, issued at GCD(input, outputs) — which for us is 78.125 kHz. Can you reconcile these?

    We can re-time the SYNC event with a large valid window. But if the SYSREF divider timing is still unknown with respect to the input, having a larger SYNC window doesn't help establish a single deterministic input-to-output phase.

    A bit further down the flowchart, we ask FZDM % FOSC == 0, and if not, SYNC is not possible because the R-divider randomizes input-to-output phase with R possible combinations. In your case, 9.765625MHz % 10MHz != 0, so you are unable to reliably synchronize because of the many possible phases of the R and N dividers. You'll notice that the same flowchart for the LMK04832 indicates that SYNC is possible with this configuration, because the R-divider can be reset.

    (a) Is the SYSREF-divider SYNC-retime mechanism valid despite the 125-phase ambiguity — e.g., because the ambiguity is granular in whole SYSREF periods (102.4 ns = exactly 256 sample periods at 2.5 GSPS) and therefore acceptable or correctable downstream — or does the ambiguity defeat multi-device determinism?

    The ambiguity defeats multi-device determinism. Again, this could be solved with R-divider SYNC on LMK04832, but with LMK04828 you are unable to synchronize given this frequency plan.

    (b) Which flowchart leaf/Type does our configuration actually land on?

    Type III.a, SYNC Not Possible.

    • Require Input-to-output determinism? YES
    • Use ZDM? YES
    • fOUT % fZDM == 0 for all fOUTYES
    • OSCIN doubler used? NO
    • fZDM % fOSC == 0? NO
    • Result: Type III.a, SYNC Not Possible
    (c) If this frequency set cannot achieve a deterministic relaxed-window SYNC on the LMK04828B, one option we're considering is moving the distributed reference itself into the power-of-2 family: CLKin0 = 156.25 MHz (= VCO/16, generated centrally by our clock-distribution unit from its OCXO) with VCXO = 156.25 MHz. That gives, with SYSREF-divider feedback: PLL1 R1 = 16, N1 = 1 (PFD = 9.765625 MHz), max(N,R) % min(N,R) = 0, and — since the flowchart identifies the non-resettable N-divider as the source of phase randomization — N1 = 1 should eliminate the input-to-output ambiguity entirely. PLL2 becomes R2 = 1, N2 = 16 (PD = 156.25 MHz), also satisfying your condition. SYNC would be issued at GCD = 9.765625 MHz with the ~102.4 ns window. Does this configuration land on a deterministic leaf and achieve the relaxed window — and do you see any issue with the R1 = 16 (R > 1) side given the LMK0482x's lack of R-divider SYNC, or is that concern eliminated when N1 = 1? If you would recommend a different change instead (different SYSREF frequency, different feedback/SYNC structure), we're open to it.

    This doesn't change the terminal leaf: you are still in Type III.a, SYNC Not Possible, because fZDM % fOSC != 0. Your SYSREF divider on each device would be locked to one reference edge of the 156.25MHz clock, but you have no control over which of the 16 possible valid edges. Again, this is a problem resolved with R-divider reset on LMK04832.

    As an alternative, you could pre-divide the 156.25MHz source and distribute 9.765625MHz to each LMK04828 CLKIN0. This satisfies fZDM % fOSC == 0, and you end up in Type I SYNC where SYNC timing has been trivialized (only one valid phase for the entire system).

    Synchronization aside, I think the switch to 156.25MHz VCXO is a good idea. Technically, 125MHz or 100MHz are also valid options - in nested ZDM, the VCXO is only fed forward into PLL2 reference, so any frequency which locks PLL2 is satisfactory. 156.25MHz is also the highest frequency I recommend providing to PLL2 phase detector (technically it's 1.25MHz over the limit in the datasheet, but this limit has some generous margin and 156.25MHz will present no problems), which is beneficial to performance - the higher phase detector frequency will reduce in-band noise by around 10log(156.25/20) = 8.9dB compared to the 20MHz suggested originally.

    Does the condition need to hold on the PLL2 pair specifically, on PLL1, or on the composite, for single-phase determinism in nested ZDM with the SYSREF divider fed back?

    The condition holds at whichever phase detector must produce repeatable input-to-output determinism. In nested ZDM this is the phase detector of PLL1. This is one of the unique benefits of nested ZDM - the conditions at PLL2 phase detector are largely irrelevant, as long as it locks, and the user is free to optimize for performance.

    By contrast, it is generally very challenging to achieve input-to-output determinism with a cascaded configuration because of the lack of control over the divider phase within PLL1. I'm not sure a cascaded use case can be reliably aligned. There might be some pathological edge cases (REF == VCXO, R1 == N1 == 1) combined with ZDM on PLL2 that could be supported, but generally PLL1 has either R or N > 1.

    For your use case, if we stick with nested ZDM, the conditions only needs to hold at PLL1 phase detector.

    (a) PLL1 uses external loop-filter components on CPout1, sized for the current 160 MHz VCXO's tuning gain — these presumably need re-sizing for a new VCXO with different Kv. What's the recommended method/tool for redesigning the PLL1 external filter (Clock Design Tool? PLLatinum Sim? a worksheet?), and can you advise a starting point given a target PLL1 bandwidth in the usual 10–300 Hz jitter-cleaning range?

    PLLatinum Sim works best. From what I can tell, a nested ZDM configuration behaves identically in simulation to a cascaded configuration, so the two PLLs can be designed as though they are independent and cascaded - design PLL1 first with the desired VCXO, then cascade this into PLL2.

    Normally, the PLL1 loop bandwidth is selected such that the reference noise is substituted above the loop bandwidth with the VCXO noise. This isn't necessarily 100Hz - sometimes the reference is really good, and we want as little of the VCXO noise as possible, hence we open the loop bandwidth a lot more, maybe to 1kHz. At some point, the 1/f noise of PLL1 will be worse than the slope of the VCXO. If you know the reference and VCXO performance and the gain of the VCXO, and your reference is better than your VCXO, you can load these parameters into PLLatinum Sim, and locate the point where the PLL noise and VCXO noise intersect, and target the loop bandwidth to that point or just a bit before it.

    In terms of general usage for PLLatinum Sim, I have the following advice:

    • Second order loop filter is probably sufficient for PLL1, unless your reference comes in with spurs that need to be filtered.
    • Since your phase detector frequency is constrained by ZDM, your available loop bandwidth pivots are the loop filter components and the charge pump gain. You can decrease charge pump gain to help reduce the size of loop filter components needed for smaller bandwidths, but be careful that the charge pump gain is still sufficient to drive the leakage impedance of the VCXO VTUNE node.
    • PLLatinum Sim will reset the VCO gain whenever a handful of parameters are changed (in most cases, this is to look up integrated VCO gain at a given frequency, but it is counterproductive when using external VCXO). Under the options menu, uncheck "Main Diagram Updates Performance Metrics" when working with PLL1 to make it save the Kvco value. This option should be reapplied for PLL2 (if it isn't automatically reapplied).
    • The default X-axis and Y-axis settings for PLL1 are not very helpful, I recommend adjusting the X-axis and Y-axis on the Phase Noise tab (uncheck Autoscale Axes) so you can see the whole phase noise curve.
    (b) PLL2's loop filter is internal/register-programmable — what internal filter settings (R3/R4/C3/C4) would you recommend at the new PD (100 or 156.25 MHz)? If it helps, we can pull the CLK104's populated PLL1 filter values from the board schematic as the baseline.

    I recommend leaving them at defaults unless there are specific VCXO spurs to be cleaned up. We don't see that much benefit to the third and fourth order components in PLL2 in general.

    With the SYSREF divider selected via FB_MUX as the internal 0-delay feedback in nested mode — the CLK104 provides no route to FBCLKin, so the feedback must be internal — does that selection place any constraints on the SYSREF-sourced outputs (CLKout3/9/11), e.g., divider or delay settings that must be held common with the feedback path?

    The feedback path comes directly from the divider output, before the SYSREF_MUX, and is a buffered copy. So there are no constraints on the SYSREF-sourced outputs.

    You cautioned that a low SYSREF-divider frequency in the feedback path can limit performance. Moving the feedback from CLKout8 (156.25 MHz) to the SYSREF divider (9.765625 MHz) drops the feedback frequency ~16×. Can you characterize or bound the phase-noise/jitter penalty for our configuration (nested ZDM, 2500 MHz VCO, with the VCXO options above)? Is there a feedback frequency below which you would advise against this approach?

    Generally, you will see a 10log(FPD_NEW/FPD_OLD) change in noise performance below the loop bandwidth from the PLL flat noise contribution, such that performance improves by 10dB/decade or 3dB/octave as phase detector frequency increases. By using nested ZDM, the impact of a reduced phase detector frequency is pushed back to PLL1 with a very low loop bandwidth, and a clean VCXO becomes the dominant noise source above the loop bandwidth of PLL1 but below the crossover frequency with PLL2 flicker/flat noise contribution. I previously mentioned that nested ZDM does not care about the precise alignment of the PLL2 phase detector since it aligns the output directly to the input, so we are free to select a PLL2 reference (VCXO) that gives the highest acceptable frequency to achieve the best in-band performance. 156.25MHz, as mentioned previously, is an excellent choice. 125MHz or 100MHz would also work, at a 1-2dB in-band noise penalty. 

    Yes, please describe the capcode override procedure you offered. Our inter-board calibration relies on each board's input-to-output phase being repeatable across power cycles (we calibrate per-board residual skew via downstream measurement of the RFSoC DAC output phase), so making the ~200 ps capcode variation static is valuable even in nested mode. Specifically: which registers set the manual capcode; how should the correct capcode be determined per device (per-unit characterization, or one value per frequency/temperature?); and does skipping calibration carry lock-reliability risk over a 25–45 °C operating range?

    I'll reiterate that nested ZDM doesn't care about the alignment at the PLL2 phase detector, so I don't feel this procedure is necessary. Nevertheless, the procedure is below. We break it up into two steps: a discovery phase, and a programming phase.

    For discovery:

    1. Lock PLL2 at the nominal operating temperature, with the desired register settings. Ensure that the LSBs of PLL2_N are written, so that the calibration of the internal VCO completes.
    2. Read back the value of the capcode from R396[6:0] (0x18D[6:0]). Record this value for each device (i.e. per-unit characterization).
    3. Optionally, this step could be repeated at several different temperatures to help bound the skew variation across temperature.

    For programming:

    1. Write all registers as described in the datasheet, including allowing initial calibration to happen. This is important for VCO amplitude calibration, which improves phase noise by increasing slew rate (without significantly impacting timing).
    2. Disable VCO calibration by writing R358[2] = 1 (0x166[2] = 1).
    3. Write the recorded capcode, bitwise OR'd with 0x80, to R376 (0x178). The bitwise OR is important, because it enables the override.

    In your case, over the 25°C to 45°C range, I don't think step 3 of discovery is required. Over a wider temperature range, the VTUNE variation might be large enough that you could bound the overall skew to a narrower window by selecting a new capcode. I'm less certain about how large this window needs to be. If you'd like to investigate on your own, you could monitor VTUNE voltage with a DMM across the operating temperature range with a forced capcode, and see if incrementing or decrementing the capcode at the boundary temperatures gives less input-to-output skew than sticking to one capcode over the whole operating range.

    I think there is no risk to lock reliability when skipping calibration over 25°C to 45°C range. Each capcode is designed to be usable across ±125°C.

  • Derek, thank you — the pre-divided reference suggestion resolves the problem completely. We are adopting it. Before closing this thread I'd like to confirm our final configuration and the multi-device determinism claims we're building on.

    Adopted configuration (each of the 8 LMK04828B devices, identical):

    • CLKin0 = 9.765625 MHz, pre-divided at the central distribution unit and distributed to all devices (per your suggestion)
    • Nested 0-delay mode; SYSREF divider fed back via FB_MUX (internal feedback)
    • PLL1: R1 = 1, N1 = 1, PD1 = 9.765625 MHz
    • VCXO = 156.25 MHz (per your recommendation); PLL2: R2 = 1, N2 = 16, PD2 = 156.25 MHz, VCO = 2500 MHz
    • Outputs: device clocks 156.25 MHz (VCO ÷16), SYSREF 9.765625 MHz (VCO ÷256)
    • SYNC: SYNC_1SHOT_EN = 1; a shared SYNC pulse to all devices' SYNC pins, held high longer than one SYSREF divider period (>102.4 ns), re-timed internally to the SYSREF divider edge

    Flowchart walk (LMK0482x page): input-to-output determinism required? YES → ZDM? YES, fZDM = 9.765625 → fOUT % fZDM == 0 for all outputs? YES (156.25/9.765625 = 16; 9.765625/9.765625 = 1) → OSCin doubler? NO → fZDM % fOSC == 0? YES (9.765625/9.765625 = 1, R1 = 1) → Type I, SYNC timing not critical, and with R1 = N1 = 1 there is no R- or N-divider phase ambiguity.

    Three confirmations requested:

    1. Multi-device determinism claim. With all 8 devices in this identical Type I configuration on the shared 9.765625 MHz reference, a single shared SYNC event places every device's output dividers on the same deterministic phase relative to the reference — one valid phase for the entire system, repeatable across power cycles. The residual device-to-device output phase variation is then bounded by: the nested-ZDM lock repeatability (~7.5 ps between power cycles, per Michael's earlier figure), propagation-delay tempco (~1 ps/°C, common-mode tracking across devices if temperature is spread evenly), and static per-device offsets (trimmed once, externally). Is that the correct residual model, or is there a term we're missing?

    2. SYNC pin usage. In this configuration the SYNC pin carries only the re-timed, non-critical event — so the ~2 ns CMOS pin delay variation over temperature you cautioned about is harmless inside the >102 ns valid window, and no tight timing is required of the SYNC generator or its distribution. Correct?

    3. SYNC frequency of issuance. Once locked and SYNC'd, the divider phases hold for as long as the PLLs remain locked — so SYNC needs issuing once at startup (and again only after a power cycle, loss of lock, or reconfiguration), not periodically. Correct?

    If all three are confirmed, we're done — and genuinely appreciative. This thread took us from "part-to-part skew isn't specified" to a configuration where the multi-device SYNC problem is eliminated by construction. The flowcharts in particular are an excellent resource; they'd make a good app note.

  • Adopted configuration (each of the 8 LMK04828B devices, identical):

    • CLKin0 = 9.765625 MHz, pre-divided at the central distribution unit and distributed to all devices (per your suggestion)
    • Nested 0-delay mode; SYSREF divider fed back via FB_MUX (internal feedback)
    • PLL1: R1 = 1, N1 = 1, PD1 = 9.765625 MHz
    • VCXO = 156.25 MHz (per your recommendation); PLL2: R2 = 1, N2 = 16, PD2 = 156.25 MHz, VCO = 2500 MHz
    • Outputs: device clocks 156.25 MHz (VCO ÷16), SYSREF 9.765625 MHz (VCO ÷256)
    • SYNC: SYNC_1SHOT_EN = 1; a shared SYNC pulse to all devices' SYNC pins, held high longer than one SYSREF divider period (>102.4 ns), re-timed internally to the SYSREF divider edge

    Flowchart walk (LMK0482x page): input-to-output determinism required? YES → ZDM? YES, fZDM = 9.765625 → fOUT % fZDM == 0 for all outputs? YES (156.25/9.765625 = 16; 9.765625/9.765625 = 1) → OSCin doubler? NO → fZDM % fOSC == 0? YES (9.765625/9.765625 = 1, R1 = 1) → Type I, SYNC timing not critical, and with R1 = N1 = 1 there is no R- or N-divider phase ambiguity.

    Looks good to me.

    1. Multi-device determinism claim. With all 8 devices in this identical Type I configuration on the shared 9.765625 MHz reference, a single shared SYNC event places every device's output dividers on the same deterministic phase relative to the reference — one valid phase for the entire system, repeatable across power cycles. The residual device-to-device output phase variation is then bounded by: the nested-ZDM lock repeatability (~7.5 ps between power cycles, per Michael's earlier figure), propagation-delay tempco (~1 ps/°C, common-mode tracking across devices if temperature is spread evenly), and static per-device offsets (trimmed once, externally). Is that the correct residual model, or is there a term we're missing?

    Mostly correct. I'll add two more considerations:

    • The 1ps/°C tempco is based on some combination of our own device's inherent propagation delay tempco, and the CVHD-950 VCXO on our EVM. Our tempco is relatively stable across devices. Depending on the VCXO selected, the tempco may go up or down as a function of the VCXO temperature stability. I have seen some, but not many, VCXOs that specify the nominal variation across temperature in the datasheet - this is a question for the VCXO vendors. They can usually give you an approximate PPM/°C characteristic curve, which can be combined with the gain to determine the expected VTUNE variation across temperature.
    • VCXO leakage from nonzero VTUNE resistance could add a current to the VTUNE node which would bias on/off time and thus input-to-output skew. I sometimes don't see DC resistance specified in the datasheets for the VCXOs (usually only modulation bandwidth), but the vendors can usually measure it or have the value on-hand. Lower leakage (higher DC input resistance) is better. I haven't seen a VCXO with leakage comparable to charge pump currents in over a decade, and I only saw it at high temperature (85°C), so in practice I doubt this will be a significant concern; generally most VCXOs are 10MΩ or better, but it's good to check beforehand if you have a chance. It's also possible that a lower but highly consistent DC input resistance becomes a fixed common skew term that can be calibrated out, but I'd look for high resistance VCXOs first since they are very common.
    2. SYNC pin usage. In this configuration the SYNC pin carries only the re-timed, non-critical event — so the ~2 ns CMOS pin delay variation over temperature you cautioned about is harmless inside the >102 ns valid window, and no tight timing is required of the SYNC generator or its distribution. Correct?

    Correct. As long as your SYNC pulse is long enough that a rising edge of the SYSREF divider re-times the signal and feeds it into the internal one-shot, timing variation at the SYNC pin is made irrelevant by the design of this nested ZDM configuration. You stated that you plan to provide a SYNC pulse at least this long, so you should have no additional timing constraints

    3. SYNC frequency of issuance. Once locked and SYNC'd, the divider phases hold for as long as the PLLs remain locked — so SYNC needs issuing once at startup (and again only after a power cycle, loss of lock, or reconfiguration), not periodically. Correct?

    In fact, it holds as long as the PLL was synchronized in lock. Loss of lock is recoverable, since the output phases are fixed with respect to the SYSREF divider phase, and the SYSREF divider phase is fixed with respect to the input reference phase. Power cycle or reconfiguration of the output dividers or SYSREF divider would disrupt the SYNC and require SYNC to be re-issued.

    For reference, any time the channel dividers or the SYSREF divider are programmed, the dividers are automatically reset with effectively randomized timing, because the SPI writes are passed from the SPI clock domain to the clock distribution path domain.

    The flowcharts in particular are an excellent resource; they'd make a good app note.

    That's the plan... it's just a matter of finding the time. I appreciate when I get asked these kind of questions, because it's a good justification to continue spending time on it!