Part Number: DP83869HM
Other Parts Discussed in Thread: AM62P, DP83869
Hello,
Hardware setup
We have a custom carrier board with a Variscite VAR-SOM-AM62P (TI AM62P SoC, CPSW3G, both RGMII ports used). On the carrier there are two DP83869HM PHYs, both in RGMII-to-copper mode (ti,op-mode set in device tree, not by straps), each one going to its own RJ45 jack through integrated magnetics. Both PHYs receive their 25 MHz reference clock from the same clock generator chip (buffered crystal output, 8 mA drive, 3.3 V VDDIO). OS is Linux (kernel 6.12, TI/Variscite BSP).
The RGMII delay configuration is tuned and validated for this board: zero align/code errors on the MAC side in all tests, and a full delay sweep was done during bring-up. Port A runs phy-mode "rgmii" (MAC internal TX delay + board RX filter), port B runs "rgmii-rxid".
Test setup
We connect port A to port B with a 1.5 m patch cable. To force real traffic over the wire (and not let the Linux kernel loop packets internally), one port is moved into a separate network namespace. Then we run bidirectional UDP iperf3, 200 Mbit/s in each direction, 45 seconds per trial (~776,000 frames per direction). Before every trial we restart auto-negotiation with ethtool -r, so every trial gets a fresh 1000BASE-T link training. We log rx_crc_errors and rx_good_frames from the MAC, plus PHY registers, for every trial.
What we see
1. The error rate is decided at link training. Some trainings give 0 CRC errors in 776k frames. Other trainings, same cable, same boot, give hundreds to thousands of CRC errors. The rate is stable inside one training and changes only at the next retrain. Example from one session, consecutive trials: 0 / 938 / 3 / 84 / 7 / 0 errors.
2. The errors are invisible to the PHY line-level counters. In a trial with 907 CRC errors, RX_ERR_CNT (0x15, cleared before traffic, read after) stayed 0 on both PHYs and the 1000BASE-T idle error counter (0x0A low byte, read after traffic) stayed 0. Only in extreme trials (10,000+ CRC errors) a few counts appear (for example 69 counts for 10,102 CRC errors), always on the receiving side of the errors.
3. MSE looks excellent. All four pairs on both PHYs read 8-1nt threshold from SNLA443 Table 2-8, and the MSE value doesnot correlate with good or bad trainings.
4. Master/slave resolution does not correlate with the errord-role trials included).
5. At 100 Mbit/s (same loop, same cable, speed limited via advertisement): 12 of 12 trials completely clean, 0 errors in 8.4 million frames total.
6. Against an external link partner (a different board, same in more than 100 million frames at full rate, on both ports,both directions, even both ports at the same time. The problem exists only when our two ports talk to each other.
7. A second identical board reproduces the issue, even stron sometimes drops completely during traffic).
What we already tried
- The full SNLA443 section 2.3.3.1 short-cable script, execureset first, all 15 writes verified by readback, soft restart, values confirmed to survive retrains), A/B/A with 12+ trials per arm: no improvement — in two independent runs the script arm actually measured
worse than stock.
- DSPFFECFG 0x012C = 0x0E81 alone (the DP83867 short-cable FFE fix known from E2E): no change. Our DP83869 default in that register is 0x0C2D.
- The SNLA443 section 2.3.3.2 AGC script: no effect.
- EEE is disabled, auto-MDIX pinned manually made no difference, RGMII delay configuration swept and validated.
Questions
1. Is there a DP83869-specific register recipe for improving 1000BASE-T margin on short cables between two DP83869 PHYs, beyond section 2.3.3.1? For
the DP83867 there is a known single-register FFE fix for exaithm requires ISI to lock" per an earlier E2E answer) — whatis the DP83869 equivalent?
2. Which diagnostic registers would you recommend reading pehy a specific training converges badly, when MSE, RX_ERR_CNTand the idle error counter all look perfect?
3. Is training-to-training variance of this size expected beices on a 1.5 m cable?
Thank you!


