DP83869HM: Sporadic CRC errors between two DP83869 ports on the same board (short cable) - error rate decided at link training

Part Number: DP83869HM
Other Parts Discussed in Thread: AM62P, DP83869

Hello,

Hardware setup

We have a custom carrier board with a Variscite VAR-SOM-AM62P (TI AM62P SoC, CPSW3G, both RGMII ports used). On the carrier there are two DP83869HM PHYs, both in RGMII-to-copper mode (ti,op-mode set in device tree, not by straps), each one going to its own RJ45 jack through integrated magnetics. Both PHYs receive their 25 MHz reference clock from the same clock generator chip (buffered crystal output, 8 mA drive, 3.3 V VDDIO). OS is Linux (kernel 6.12, TI/Variscite BSP).

The RGMII delay configuration is tuned and validated for this board: zero align/code errors on the MAC side in all tests, and a full delay sweep was done during bring-up. Port A runs phy-mode "rgmii" (MAC internal TX delay + board RX filter), port B runs "rgmii-rxid".

Test setup

We connect port A to port B with a 1.5 m patch cable. To force real traffic over the wire (and not let the Linux kernel loop packets internally), one port is moved into a separate network namespace. Then we run bidirectional UDP iperf3, 200 Mbit/s in each direction, 45 seconds per trial (~776,000 frames per direction). Before every trial we restart auto-negotiation with ethtool -r, so every trial gets a fresh 1000BASE-T link training. We log rx_crc_errors and rx_good_frames from the MAC, plus PHY registers, for every trial.

What we see

1. The error rate is decided at link training. Some trainings give 0 CRC errors in 776k frames. Other trainings, same cable, same boot, give hundreds to thousands of CRC errors. The rate is stable inside one training and changes only at the next retrain. Example from one session, consecutive trials: 0 / 938 / 3 / 84 / 7 / 0 errors.
2. The errors are invisible to the PHY line-level counters. In a trial with 907 CRC errors, RX_ERR_CNT (0x15, cleared before traffic, read after) stayed 0 on both PHYs and the 1000BASE-T idle error counter (0x0A low byte, read after traffic) stayed 0. Only in extreme trials (10,000+ CRC errors) a few counts appear (for example 69 counts for 10,102 CRC errors), always on the receiving side of the errors.
3. MSE looks excellent. All four pairs on both PHYs read 8-1nt threshold from SNLA443 Table 2-8, and the MSE value doesnot correlate with good or bad trainings.
4. Master/slave resolution does not correlate with the errord-role trials included).
5. At 100 Mbit/s (same loop, same cable, speed limited via advertisement): 12 of 12 trials completely clean, 0 errors in 8.4 million frames total.
6. Against an external link partner (a different board, same in more than 100 million frames at full rate, on both ports,both directions, even both ports at the same time. The problem exists only when our two ports talk to each other.
7. A second identical board reproduces the issue, even stron sometimes drops completely during traffic).

What we already tried

- The full SNLA443 section 2.3.3.1 short-cable script, execureset first, all 15 writes verified by readback, soft restart, values confirmed to survive retrains), A/B/A with 12+ trials per arm: no improvement — in two independent runs the script arm actually measured
worse than stock.
- DSPFFECFG 0x012C = 0x0E81 alone (the DP83867 short-cable FFE fix known from E2E): no change. Our DP83869 default in that register is 0x0C2D.
- The SNLA443 section 2.3.3.2 AGC script: no effect.
- EEE is disabled, auto-MDIX pinned manually made no difference, RGMII delay configuration swept and validated.

Questions

1. Is there a DP83869-specific register recipe for improving 1000BASE-T margin on short cables between two DP83869 PHYs, beyond section 2.3.3.1? For
the DP83867 there is a known single-register FFE fix for exaithm requires ISI to lock" per an earlier E2E answer) — whatis the DP83869 equivalent?
2. Which diagnostic registers would you recommend reading pehy a specific training converges badly, when MSE, RX_ERR_CNTand the idle error counter all look perfect?
3. Is training-to-training variance of this size expected beices on a 1.5 m cable?

Thank you!

  • Hello,

    Does this only occur with short cables? If you increase the cable length does the issue go away?

    1. The short cable script is the register recipe for improving short cable margin. 

    2. Its good that you've read these registers. It would be good to perform both an MII loopback and reverse loopback test in the error-case to ensure this is coming from the cable connection between the PHYs. MII loopback would involve sending data to the transmit PHY, where it will be looped back to the receive pins of the MAC. Reverse loopback would be set on the receiving PHY and will send packets back towards the transmitting PHY. Checking that returning packets are error-free on the MAC would tell us if the signal path is experiencing an issue.

    3. While link training parameters do impact performance, I would expect the short cable and AGC scripts you've applied to have some effect in preventing errors. Since these are not helping, I would like to see if loopback testing reveals any useful hints.

    Best,

    Shane

  • Thank you, we ran your suggestions. Results below.

    Cable length: we tested 1.5 m, 2 m and 5 m cables (the longer ones also on a second identical board): the error behavior is unchanged across all lengths. We have not yet tested above 10 m. Also relevant: on the same SoM, a loopback between a DP83867 and an ADIN1300 with the same 1.5 m cable and the same method shows 0 errors in 9.3 million frames, so this looks specific to the DP83869-to-DP83869 pairing rather than to cable length or the platform

    Loopback tests: MII and digital loopback could not source MAC traffic under Linux, phylib keeps the carrier down while the PHY is in near-end loopback, so the MAC never transmits (a software limitation, not a signal result). Reverse loopback worked and gave the clearest result we have so far. With the receiving PHY in reverse loopback (register 0x16 = 0x0020) and the other side sending UDP at about 590 Mbit/s, every frame crosses the 1.5 m cable twice and passes through the full receive path of the far PHY before returning:

    - 25.4 million frames, across 17 separate link trainings (we restarted auto-negotiation between rounds): 2 CRC errors total (~1e-7).
    - Normal-traffic trials interleaved on the same boot, same cable, same procedure: up to ~6e-4 (hundreds of errors per 500k frames on bad trainings).
    - We tested both directions (each PHY as the looping end): same clean result. The sending side's MAC TX and MAC RX are both inside this clean chain, so the MAC/RGMII paths are validated by this test as well

    So the signal path, cable, magnetics, both transmitters, both receivers, both MAC connections, carries 50+ million cable crossings essentially error-free. The one variable that changes everything: in reverse loopback the two directions of the cable carry correlated content (the far PHY re-transmits the near PHY's own stream), while in normal operation the two directions carry independent streams. Same wire, same PHY pair, same trainings, higher throughput, only the independence of the two data streams changed, and the error rate changed by roughly 5000x. To us this looks like the adaptive cancellation absorbing correlated interference but not independent far-end traffic, on trainings where the converged DSP has little margin.

    On point 3: agreed and this may explain why the short-cable and AGC scripts did not help: they target training time and gain convergence, while our failure appears only under simultaneous independent bidirectional traffic after an apparently successful training (MSE excellent, RX_ERR_CNT = 0, idle error counter = 0, no false carrier events — only the FCS sees it).

    Questions: does the correlated-vs-independent signature point to a specific block (echo canceller / NEXT canceller adaptation) in the DP83869? Is there any register visibility into canceller state or convergence quality that we could log per training? And is DP83869-to-DP83869 on short links with full-duplex independent traffic a known corner case?

    Thank you!

  • The correlated vs independent signature is a good observation, however I would not expect this in-itself to cause an issue. I am not aware of link dependent issues with DP83869 when it is connected to another DP83869 in other designs. Rather than narrowing in on internal PHY processing blocks I want to rule out a design issue.

    Are you able to share the schematic and/or layout of the design for review? If you would prefer to keep this off the public forum I can send you a message through E2E's direct message feature. This creates a private space only viewable to yourself and I.

    Can you share waveforms of the following RGMII signals in both the passing and failing cases? These should be measured on the receiving PHY connected to the MAC reporting received CRC errors:

    • RX CLK
    • RX CTRL
    • RX D0

    Additionally a waveform of the XI input clock to the PHY would help to see. 

    And is DP83869-to-DP83869 on short links with full-duplex independent traffic a known corner case?

    We have seen short cables produce link instability before, however this is addressed via the short cable script you've mentioned. IF this script is having no effect, and increasing the cable length also has no effect, this does not appear length-dependent. Certainly let me know if testing with a cable longer than 10m shows different results.

    Best,

    Shane

  • I have sent you the schematics privately - PHY A, PHY B and the clock section. Also in that thread is a scope capture of the XI reference clock at 25 MHz.

    On the rest of the waveforms you asked for, I have to be straight with you: the scope available to us does not have the bandwidth for RGMII. At 125 MHz with sub-nanosecond edges it cannot produce a trace anyone could draw a conclusion from, so rather than send you something misleading I am not sending it at all. What I can do instead is measure from inside the PHY itself, and that is what the rest of this message is.

    One thing I checked and closed on the running board: I read the Si5351 crystal load capacitance register directly over I2C and it is programmed correctly, so the 25 MHz reference is not sitting off frequency. That is measured, not assumed.

    The first measurement is the 0x55 skew FIFO idea, tested properly. I set the Sync FIFO Control register 0xE9 to 0xDF22 on both PHYs before each link as 7.3.2.1 requires, and I verified it read back as 0xDF22 on every single training. 32 trainings, bidirectional UDP at 200 Mb/s for 30 seconds each, CRC counted per direction.

    The result is negative. Grouping by the documented SFD variation field of the follower PHY, field = 0 gave 15 trainings with 3 completely clean and a median of 83 errors, while field = 1 gave 14 trainings with none clean and a median of 397 errors. Mann-Whitney p = 0.097, so not significant. More to the point, the overlap is total: the good group contains a training with 2115 errors and the bad group contains one with 5. So 0x55 does not predict what a given link will do. An earlier run without the 0xE9 initialisation trended the same way at p = 0.085, so there may be a weak shift, but it is not the state variable I was hoping for. The only documented per-link-persistent value in the part does not explain the lottery.

    The second measurement is the more interesting one. I looked at whether the two directions degrade together or independently. Of the 32 trainings, 14 had exactly one direction bad at 50 errors or more, only 3 had both directions bad, and 12 had one direction at least 20 times worse than the other. Spearman rho between the two directions is 0.355 with p = 0.037.

    So a training does not just pick an error rate, it picks a direction. That looks like a per-receiver effect, one of the two receivers converging badly at training time, rather than anything acting on both directions at once. Put next to the reverse loopback result I sent earlier, 25.4 million frames with 2 errors, and the fact that the fault needs two simultaneous streams, the picture I keep arriving at is the receive path's echo or NEXT canceller converging badly for one receiver, and only mattering when that receiver has to work while its own transmitter is active.

    For reference on the run: 4 of 32 trainings were completely clean, the median was 81.5 errors, the worst training had 6046 errors which is about 6e-3, and PHY A came up as leader in 29 of the 32.

    The question I still most want your view on is the RGMIICTL asymmetry, which is also a measured result. The two ports need different settings and each has exactly one that works. PHY A works at 0x32 = 0xD3 and its RX is dead at 0xD2. PHY B works at 0xD2 and is about 30 times worse at 0xD3. That is a 2 ns difference in required internal delay between two ports running the same silicon at the same speed. Is there any path by which the RGMII sync FIFO half-full phase in 0x32[6:3] and 0x33[1:0] can be established badly at training and stay bad until the next one?

    I can also run a 16 step 250 ps sweep of RGMIIDCTL 0x86 DLL_TX and DLL_RX independently on each port, scored by frame errors, as a functional eye margin map. That is the closest thing to a scope trace I can produce. Tell me if you want it and I will send tables.

    Best,
    Anastasios

  • Hi Anastasios,

    I see the private message. Please allow me time to look through the information you've provided. I will aim to reply early next week.

    Best,

    Shane

  • Hi Anastasios,

    Looking at the schematic:

    1. The insertion loss of the magnetics is slightly above the -1dB spec in the DP83869 datasheet table 9-3. Otherwise the magnetics look ok.

    2. RBIAS should have one 11k pulldown, not a capacitor. Is there a reason for C266 being here? I recommend depopulating this if it is not already:

    Otherwise the schematic looks ok. It would be good to see the layout as well since high speed signals are sensitive to trace routing.

    As for the theory that this is due to NEXT/reciever convergence, the results of your reverse loopback testing do not suggest this. If the link partner is in reverse loopback, and data is sent from the DUT, then there are two simultaneous directions active on the DUT PHY. This has data transmitting and receiving at the same time, yet the results always show a negligible error count. Please correct me if I'm mistaken in my analysis. If reverse loopback testing never shows the error case it suggests the PHY MDI path is stable. 

    • Was there ever a high-error case in your reverse loopback testing? It seems odd to me that this test is error free from either PHY's perspective, yet when the data passes through the whole signal chain there is no issue.

    I remember you mentioned that in extreme cases you noticed the idle error register incrementing. Its a long-shot, but can you test whether lowering the viterbi idle threshold in register 0x0053 has any effect when the link is showing crc errors?

    Best,

    Shane

  • Hi Shane,

    Thank you for the review, and sorry for the slow reply.

    C266 / RBIAS: it came from a TI application note we followed during the design — I don't remember exactly which one. Either way it does not matter: I depopulated C266 and retested, and there was no change in the error behavior, so it stays off the board. The 11k pulldown is fitted as required. The magnetics insertion-loss point is noted for the next board revision.

    Layout: agreed it matters. If you want, I can send you the routing views of those areas (MDI, RGMII, clocks) in a private message.

    Reverse loopback: no, there was never a high-error case: 17 trainings, 25.4M frames, 2 CRC errors in total. Normal traffic on the same link shows 0 to ~6000 errors per 60 s window, depending on the training. So I agree with you: the MDI path itself looks stable. One clear difference between the two tests: in reverse loopback the returned data is a copy of what the DUT transmits, while the failing traffic is two independent streams. An open question for you: after the side-stream scramblers, can the relationship between the two payloads matter to the DSP at all? If it cannot, this difference is a red herring.

    Clock offset: one more property of this failing pair. Both PHY XI inputs come from one Si5351A, and both outputs are driven straight from its crystal. So the two PHYs run at exactly 0 ppm frequency offset, permanently. Every pairing with an independent oscillator is clean: external link partners show >100M error-free frames at full rate, both ports, even fully bidirectional, and a DP83867<->ADIN1300 pair on the same SoM is clean too (9.3M frames). Our working theory: the failure needs both things at once — independent data streams AND the exact 0 ppm offset, so a bad DSP operating point picked at training is never averaged out by drift. Reverse loopback removes the first condition, an external partner removes the second, and both are clean. Two more facts fit this: at a fixed configuration the error count is decided at link training (spread 0 to 6046 per training), and within one training the errors usually pick one direction (14 of 32 trainings were fully one-sided). We are testing the theory right now by shifting one PHY reference by +50 ppm. I will post the result.

    Register 0x0053: on the bench list. I will lower the viterbi idle threshold during a bad training and report it together with the +50 ppm result.

    A question back: can you sanity-check our measurement method?
    - The two RJ45 ports of the same board, joined with a short patch cable.
    - One port moved to its own Linux network namespace, so traffic really crosses the wire.
    - Both MACs in promiscuous mode (needed by the CPSW ALE for same-board traffic), firewall rules flushed.
    - iperf3 UDP, 200 Mbit/s per direction, both directions at once, 60 s per run.
    - Each "training" = an ethtool -r renegotiation.
    - Errors counted with the MAC hardware FCS counters (ethtool -S rx_crc_errors) on both ports.

    What puzzles us: the PHY-level counters stay clean the whole time (RX_ERR_CNT = 0, no idle errors, MSE reads excellent) while the MAC counts CRC
    errors. Is that expected, and is there a PHY register you wonts at the line level?

    Device tree excerpts (clocks and PHYs):

    /* Si5351A: both PHY refclks straight from its 25 MHz crysta
    clkout@2 { reg = <2>; silabs,clock-source = <2>; /* xtal direct */
    silabs,drive-strength = <8>; clock-frequency = <2XI */
    clkout@3 { reg = <3>; silabs,clock-source = <2>; /* xtal direct */
    silabs,drive-strength = <8>; clock-frequency = <2XI */

    cpsw3g_phy_a: ethernet-phy@3 { reg = <3>; clocks = <&clk_ge
    ti,op-mode = <DP83869_RGMII_COPPER_ETHERNET>; enet-phy-lane-no-swap; };
    cpsw3g_phy_b: ethernet-phy@c { reg = <12>; clocks = <&clk_ge
    ti,op-mode = <DP83869_RGMII_COPPER_ETHERNET>; };

    &cpsw_port1 { phy-mode = "rgmii"; phy-handle = <&cpsw3g_phy_a>; }; /* RGMIICTL 0xD3 */
    &cpsw_port2 { phy-mode = "rgmii-rxid"; phy-handle = <&cpsw3gxD2;
    the SoM adds a passive RX delay on this port's RGMII bus, so the PHY RX delay is off */

    Best,
    Anastasios

  • Follow-up with the two promised tests.

    1. Reference offset test. I moved PHY A's XI from the shared crystal buffer to a spare fractional multisynth on the same Si5351 and verified the result on silicon. Three conditions, 15 link trainings each, 60 s of 200 Mbit/s bidirectional UDP per training, MAC FCS counters:

    - Both PHYs on the shared crystal, 0 ppm (the normal board): the usual lottery - median ~82 errors per training over our n=32 dataset, some trainings clean.
    - PHY A at +50 ppm (25.00125 MHz): PHY A's receive direction was bad in 15 of 15 trainings — median ~4,400 errors, minimum 375, maximum 13,851. Roughly 50x the baseline. PHY B's direction kept the usual lottery.
    - Control: the same multisynth path set to exactly 25.000000 MHz (same synthesizer jitter, 0 ppm): back to the normal lottery (median ~106, minimum 11, some near-clean trainings).

    So my 0 ppm theory was wrong — the lottery is still there with an offset, and also at 0 ppm through a different clock path. But the test found something new: a 50 ppm reference offset between the two PHYs (well inside the 802.3 tolerance) multiplies the error rate about 50x, concentrated on the offset PHY's receive direction, in every single training. The contrast is striking: against external link partners, which naturally run some ppm apart, both ports are clean at full rate. So the offset sensitivity also appears only in the DP83869<->DP83869 pairing. Is a 50x degradation at 50 ppm XI offset between two DP83869s expected behavior? This looks like the strongest lead we have so far.

    2. Register 0x0053 (viterbi idle threshold). Tested inside one bad training, no renegotiation, thresholds written to both PHYs, one 60 s measurement per setting (errors eth0/eth1): 0x2055 default: 2371/2580; 0x2054: 920/993; 0x2053: 697/2189; back to 0x2055: 967/2258. I would call this no clear effect - the changes are within the natural variance, because at fixed settings consecutive 60 s windows in the same training varied from 309 to 2371 on the same port. That variance is itself a data point: the error rate is decided at training, but it also wanders a lot within a training.

  • Hi Anastasios,

    1. The relationship of the output and input data should not affect the performance of DP83869. If the reverse loopback test is working, the PHYs are able to talk to one-another over the MDI

    2. DP83869 requires a +/-100ppm or less to work correctly. 50ppm is ok, and it seems this ppm works with other link partners. It is interesting that the ppm exacerbates the issue specifically on PHY A, yet even lowering the ppm to 0 does not fix this. 

    3. Are you able to share the layout in our private message? I would like to review this incase there are subtle layout issues contributing to the problem. Furthermore, are you performing any other register configuration to the PHY besides the RGMII RX/TX delays and the short cable script?

    Best,

    Shane 

  • Hi Shane,

    1. Understood, thanks.

    2. Agreed, 50 ppm is inside the +/-100 ppm spec; my point is only the sensitivity. One clarification: both PHYs on this board run from the same clock buffer, so all the normal tests were already at 0 ppm. The extra 0 ppm test only showed that the alternative clock path itself adds no errors.
    Summary: 0 ppm -> random per link-up, median ~82 CRC errors per 60 s at 200 Mbit/s bidirectional, some link-ups clean. +50 ppm on PHY A -> PHY A receive side bad in 15 of 15 link-ups, about 50x more errors.
    If useful, I can repeat the +50 ppm test and log which PHY is master/slave in each link-up (I did not record it the first time - see the master/slave note at the end, it is probably relevant).

    3. Layout: yes. I will send them to you in a private message.

    REGISTER CONFIGURATION - complete list, checked against the driver source.
    U-Boot does not touch the PHYs on this board (its DP83869 and CPSW drivers are not compiled in). Everything is done by the Linux dp83869 driver (kernel 6.12) from our device tree. Before probe, the MDIO bus hardware-resets both PHYs (RESET_N low 100 us, 12 ms wait). Then this sequence runs at probe and again at every interface-up (preceded by SW_RESET, reg 0x1F = 0x8000):

    - CFG2 (0x14): downshift enable, bits 9:8 set
    - OP_MODE_DECODE (MMD 0x1F, 0x01DF) = 0x0040, both PHYs - RGMII to copper (ti,op-mode)
    - BMCR (0x00) = 0x1140
    - PHY_CTRL (0x10) = 0x5048 - FIFO depth code 1 in both directions (driver default)
    - CTRL1000 (0x09) = 0x0B00, then phylib rewrites bits 9:8 from the advertised modes -> reads 0x0A00; no manual master/slave
    - LEDS_CFG1 (0x18) bits 7:0 = 0x10 - LED_0 link, LED_1 RX/TX activity; LEDS_CFG2 (0x19) bits 2 and 6 set (active high)
    - GEN_CFG3 (MMD 0x1F, 0x0031) bit 0 cleared - port mirroring off (enet-phy-lane-no-swap), strap not followed
    - RGMIIDCTL (MMD 0x1F, 0x0086) = 0x0077, both PHYs - 2.0 ns for TX and RX (driver default)
    - RGMIICTL (MMD 0x1F, 0x0032): read, set bits 1:0, clear the bit named by phy-mode, write back -> port A ("rgmii") = 0x00D3, port B ("rgmii-rxid") = 0x00D2; the upper bits are the reset value
    - CTRL (0x1F) = 0x4000 - SW_RESTART, then 1-2 ms wait
    - Not written on this board: IO_MUX_CFG (0x0170) - no impedance or clock-output property, so the impedance bits stay at factory trim and CLK_O_SEL stays at its strap/reset value (0xC = reference clock out; the CLK_OUT pin is unused). CFG4 INT_OE is not written (polling, no interrupt line).
    - Generic phylib on the same paths: MICR (0x12) = 0; power-down bit cleared; advertisement registers (0x04, 0x09 bits 9:8, EEE MMD 7 reg 0x3C) rewritten from the advertised modes; then BMCR autoneg enable + restart. Our per-link-up renegotiation (ethtool -r) is only a BMCR autoneg restart; it does not re-run the sequence above.

    Correction to my earlier post: I wrote that port B's PHY RX delay is off. Reading the driver again, it programs RGMIICTL following the datasheet definition "0 = clock shifted with respect to data, 1 = aligned" (upstream commit 2e1ec861a605, acked by TI). So with 0x00D2 port B has the PHY RX delay ON and TX delay off, and port A with 0x00D3 has both PHY delays OFF - port A's RX skew must come from the SoM/board, not from the PHY. Could you confirm that reading of bits 1:0? It decides which PHY carries the delay.

    Test-only settings: the short cable script was used only in the one comparison I reported earlier; it is not active in any later result, including the ppm tests. Forced master/slave was used only in the role tests. All test writes are cleared by the interface restart, and the board is power-cycled before each test series.

    FULL REGISTER DUMP - normal test configuration (loopback cable, both links up at 1000/full, no traffic). Format: register (address): PHY A (addr 3) / PHY B (addr 12)

    BMCR (0x00): 0x1140 / 0x1140
    ANAR (0x04): 0x05E1 / 0x05E1
    CTRL1000 (0x09): 0x0A00 / 0x0A00
    PHY_CTRL (0x10): 0x5048 / 0x5048
    PHYSTS (0x11): 0xBC02 / 0xBF02 (B resolved MDI-X)
    LEDS_CFG1 (0x18): 0x6110 / 0x6110
    LEDS_CFG2 (0x19): 0x4444 / 0x4044 (bit 10 = LED_2 polarity, strap)
    RGMIICTL (MMD 0x1F, 0x0032): 0x00D3 / 0x00D2
    RGMIIDCTL (MMD 0x1F, 0x0086): 0x0077 / 0x0077
    IO_MUX_CFG (MMD 0x1F, 0x0170): 0x0C0F / 0x0C10 (factory impedance trims)
    OP_MODE_DECODE (MMD 0x1F, 0x01DF): 0x0040 / 0x0040
    VTM_CFG (MMD 0x1F, 0x0053): 0x2055 / 0x2055
    MMD 0x1F, 0x0055: 0x0000 / 0x1001 (see note 1)
    MMD 0x1F, 0x00E9: 0x9F22 / 0x9F22
    DSPFFECFG (MMD 0x1F, 0x012C): 0x0C2D / 0x0C2D
    STRAP_STS1 (MMD 0x1F, 0x006E): 0x0034 / 0x02CC (see note 2)
    STRAP_STS2 (MMD 0x1F, 0x006F): 0x0000 / 0x0000

    Note 1 - 0x0055: we never write it. Over 12 link-ups it always read 0x0000 on the master and 0x1001 / 0x1101 / 0x1110 on the slave (the values swapped exactly in the one link-up where port A became slave). So it is a slave-side register whose low bits change per link-up. What is it?

    Note 2 - the strap registers differ between the two PHYs (MDIO address 3 vs 12 and the LED_2 polarity are expected). Software overrides op-mode, port mirroring and the RGMII delays anyway, but could you check the two STRAP_STS1 values for any other functional difference?

    Note 3 - master/slave: when I restart autoneg from port A (ethtool -r), port A comes up master in 11 of 12 link-ups, although neither PHY has a manual setting or port-type preference (CTRL1000 = 0x0A00 on both). Is that expected? It also means that in the +50 ppm test the affected receiver (PHY A) was most likely the master, i.e. the side the partner loop-times to.

    Best,
    Anastasios

  • Hi Anastasios,

    Shane is OoO and will be back on Tuesday following US holiday.

    Sincerely,

    Gerome

  • Hi Anastasios,

    Apologies for the delay and thank you for pinging this thread. To answer your questions:

    So with 0x00D2 port B has the PHY RX delay ON and TX delay off, and port A with 0x00D3 has both PHY delays OFF - port A's RX skew must come from the SoM/board, not from the PHY. Could you confirm that reading of bits 1:0? It decides which PHY carries the delay.

    Since RGMIICTL = 0x00D3 on PHYA, there is no RGMII clock delay on this PHY. PHYB allows shifting of the RGMII clk for the RX path, and the shift is 2ns (Set in RGMIIDCTL). 

    Note 1 - 0x0055: we never write it. Over 12 link-ups it always read 0x0000 on the master and 0x1001 / 0x1101 / 0x1110 on the slave (the values swapped exactly in the one link-up where port A became slave). So it is a slave-side register whose low bits change per link-up. What is it?

    0x0055 shows the MDI-facing channel delay of the PHY, where each group of 4 bits corresponds to a channel. A value of '0' indicates no delay is added to the corresponding channel, whereas a value of '1' indicates 8ns of delay is being added to that channel to align the symbols of that channel with the symbols from the other 3 channels. At least one hex value in this register should read '0' for no delay. And you are correct this is a slave-side register.

    Note 2 - the strap registers differ between the two PHYs (MDIO address 3 vs 12 and the LED_2 polarity are expected). Software overrides op-mode, port mirroring and the RGMII delays anyway, but could you check the two STRAP_STS1 values for any other functional difference?

    The only differences I see between the Strap_STS1 register of these two PHYs are the OPMODE, Address, and the autonegotiation mode (LED_2 polarity). Double check that these are addressed by your software correctly

    Note 3 - master/slave: when I restart autoneg from port A (ethtool -r), port A comes up master in 11 of 12 link-ups, although neither PHY has a manual setting or port-type preference (CTRL1000 = 0x0A00 on both). Is that expected?

    The master/slave selection should be randomized based on each PHY's seed value if no manual configuration is performed. Due to its random nature, its possible PHY A has generated the winning seed value in 11 out of the 12 link ups.

    Now for my side:

    1. I reviewed the layout images you sent over, and there are signals that exceed our 50mil length matching guideline (RX_D2 on PHY A and TX_CTRL on PHY B). We recommend having all signals matched within 50mils to their respective clock signal (RX_CLK or TX_CLK) to give margin for the timing requirements. That being said, the routing itself looks ok and if you haven't seen issues with other link partners, length matching may not be the root cause, though it can be improved. 

    2. You mentioned that the RGMII delays are optimized already, and that there were no MAC side align errors when the delay value was set. Can you elaborate what MAC testing was done when the delays were validated? I remember you mentioned the DP83869's MII loopback wouldn't source packets from the MAC, so it could not be tested, so I'm curious what MAC testing showed the delays are configured correctly.

    My thinking is that if the PHYs do not report RX errors on the MDI side and errors are being reported at the MAC, there could be something on the MAC side causing the link-dependent errors. The PHY will simply pass data from one side to the other, and so long as the RGMII delays are configured correctly, and the link is coming up, the PHY should not be introducing errors into the data. The fact that reverse loopback works from both sides supports this (Data can pass from one MAC through both PHYs and return ok, yet once both MACs are communicating there is a link up dependency). Have you reached out to the AM62 team for their input on this issue? I am not familiar with the processor side, however they may know other ways to narrow down link dependent issues.

    Probing the RGMII lines would be one way to see that the data is reaching the MAC ok, yet it seems this is not an option. If you already optimized the RGMII delays, we can tune the IO_MUX_CFG register to see if changing the RGMII impedance has any effect on the issue. In general, a lower impedance will give faster rise/fall times while a higher impedance gives slower Rise/fall times. Have you tried tuning this register? If not, perhaps see if raising or lowering this has any effect.

    Best,

    Shane