This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

J722SXH01EVM: CSI0 distortion during multiple vc id after FIFO overflow

Part Number: J722SXH01EVM

Hi,

We are using a GMSL chip which has multiple input ports to tunnel two CSI cameras via one CSI0 lane.

We assign VC ID 0 and 1 to each camera.

During the stream opening, we sometimes get some FIFO overflows. However, when this error occurs, the ti shim driver begins the DMA and it causes the stream to permanently be offset until closing and re-opening the streams.

When the FIFO overflows do not happen, the stream opens normally.

This also has implications for FPD link too, as any FIFO overflow can cause this permanent offset while the stream is active?

Is there some other way to make TI's DMA engine shim driver more tolerant of FIFO overflows?

Here's the csi log status when the errors occur:
v4l2-ctl --log-status -d /dev/v4l-subdev4

Status Log:

   cdns_csi2rx.30101000.csi-bridge: =================  START STATUS  =================
   cdns-csi2rx 30101000.csi-bridge: Overflow of the Stream 0 FIFO detected events: 543
   cdns-csi2rx 30101000.csi-bridge: Unrecoverable ECC error events: 1
   cdns-csi2rx 30101000.csi-bridge: CRC error events: 132
   cdns_csi2rx.30101000.csi-bridge: ==================  END STATUS  ==================

Here's what the bar looks like. It's actually very distracting because the bar will have partially corrupted data in it, causing the a lot of flickering in both video streams.
image.png

  • Hi Jared,

    Yes those patches are applied. I'm on SDK 11.01.02.01.

    I think I have a workaround.
    I found I can issue a soft link reset on both streams when the second channel opens to prevent the FIFO/CRC errors.
    Basically, the first stream always opens correctly, but the second stream opens with some delay and can cause corruption?

  • I also want to point out, this doesn't fix the underlying issue with the CSI driver. With multiple streams, FIFO errors could cause this problem. So the common case is fixed with my workaround, but there is still vulnerability.
    As a test, I changed the tuning on the GMSL chip to trigger FIFO/CRC failures and am still able to reproduce the corrupted vertical bar.
    This also has implications for BCI immunity testing, as active streams may not properly recover. Its fine to close the ticket for now, but I may open another one when we start testing BCI.

    It seems the csi_rx_if_csi_rx_if0_public registers for 'stream0' are affecting both vc id 0 and 1. I also found setting SOFT_RST can make the offset either significantly worse or fix it. However, I cannot get the start of vc id 1 to appear on my vc id 0 stream, so seems like something is almost correct. Maybe it's an artifact of how both frames end.

  • Hi ,

    I have not seen FIFO errors when using virtual channels with our FPD link serializers/deserializers. What are the parameters you are using for the cameras and CSI controller (data rate, pixel freq, resolution, data type, frame rate, etc)?

    Performing a soft reset to clear the FIFO appears to be the solution recommended within the TRM:

    12.6.1.4.8.5 Error Control With Soft Resets

    The CSI_RX_IF will perform control of the soft reset either for error event recovery or to clear a stream FIFO or
    internal state machine.

    • The FRONT block can be soft reset if the DPHY_RX becomes unresponsive and the controller wishes to maintain its configuration. In this case the DPHY_RX resets can be applied and the DPHY_RX enabled to begin the transfer again.
    • The Protocol block can be soft reset if the FRONT soft reset is required, and the protocol is not in the IDLE state.
    • The stream soft resets (CSI_RX_IF_VBUS2APB_STREAM0_CTRL - CSI_RX_IF_VBUS2APB_STREAM3_CTRL)[4] SOFT_RST can be used to clear the stream to the stop state and reset all the stream state machines and FIFO. If the system has a failure and wishes to clear the stream FIFO and return to a safe state on the pixel interface, the stream soft reset should be asserted.

    Although, it also states that your issue shouldn't occur:

    12.6.1.4.6.2 PSI_L DMA error handling due to FIFO overflow

    The DMA error handling is also called a PSI_L protocol enforcer. It is intended to prevent hang of the PSILSS0. Unnatural packet size should not cause hang so only SOP/SOL/EOL/EOP framing is enforced. The following list highlights the error handling mechanism:

    • For context cleanup the protocol enforcer:
      • cycle through each context index checking if context is in MOL or MOP
      •  close out MOP with EOP and MOL&MOP with EOL&EOP
    • dropOnFloor SOP if currently in MOPstate
    • dropOnFloor EOP if not currently in MOPstate
    • dropOnFloor FIFO data
    • EOL context if EOP and MOLstate
    • After closing out all open contexts the PSILSS0 logic will then wait till end of frame per virtual channel. Once a new frame starts it will then start sending out data from that new frame

    Best,
    Jared

  • Hi Jared,

    Thanks for this. I'm also getting this failure with an isl79987 chip, which is entirely on the PCB.
    The failure isn't immediate, it's only after running a few hours.


    From the device tree, we are using:

    data-lanes = <1 2>;
    clock-lanes = <0>;
    link-frequencies = /bits/ 64 <800000000>;

    The resolution of the analog cameras is 720x480.
    We are clocking it at 60 frames a second.

    Here's the failure on this chip too:

    v4l2-ctl --log-status -d /dev/v4l-subdev1

    Status Log:

    cdns_csi2rx.30121000.csi-bridge: ================= START STATUS =================
    cdns-csi2rx 30121000.csi-bridge: A reserved or invalid short packet has been received events: 33
    cdns-csi2rx 30121000.csi-bridge: Data ID error in the header packet events: 33
    cdns-csi2rx 30121000.csi-bridge: ECC error detected and corrected events: 579
    cdns-csi2rx 30121000.csi-bridge: Unrecoverable ECC error events: 608
    cdns-csi2rx 30121000.csi-bridge: CRC error events: 22
    cdns_csi2rx.30121000.csi-bridge: ================== END STATUS ==================

    It also makes a bar of corruption in the same way. So maybe it happens with any CRC or header packet error? This only happens with multiple streams going at once.

    media-ctl -p -d /dev/media1
    Media controller API version 6.12.43

    Media device information
    ------------------------
    driver j721e-csi2rx
    model TI-CSI2RX
    serial
    bus info platform:30122000.ticsi2rx
    hw revision 0x1
    driver version 6.12.43

    Device topology
    - entity 1: 30122000.ticsi2rx (5 pads, 5 links, 4 routes)
    type V4L2 subdev subtype Unknown flags 0
    device node name /dev/v4l-subdev0
    routes:
    0/0 -> 1/0 [ACTIVE]
    0/1 -> 2/0 [ACTIVE]
    0/2 -> 3/0 [ACTIVE]
    0/3 -> 4/0 [ACTIVE]
    pad0: SINK
    [stream:0 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:1 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:2 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:3 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    <- "cdns_csi2rx.30121000.csi-bridge":1 [ENABLED,IMMUTABLE]
    pad1: SOURCE
    [stream:0 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    -> "30122000.ticsi2rx context 1":0 [ENABLED,IMMUTABLE]
    pad2: SOURCE
    [stream:0 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    -> "30122000.ticsi2rx context 2":0 [ENABLED,IMMUTABLE]
    pad3: SOURCE
    [stream:0 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    -> "30122000.ticsi2rx context 3":0 [ENABLED,IMMUTABLE]
    pad4: SOURCE
    [stream:0 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    -> "30122000.ticsi2rx context 4":0 [ENABLED,IMMUTABLE]

    - entity 7: cdns_csi2rx.30121000.csi-bridge (5 pads, 2 links, 4 routes)
    type V4L2 subdev subtype Unknown flags 0
    device node name /dev/v4l-subdev1
    routes:
    0/0 -> 1/0 [ACTIVE]
    0/1 -> 1/1 [ACTIVE]
    0/2 -> 1/2 [ACTIVE]
    0/3 -> 1/3 [ACTIVE]
    pad0: SINK
    [stream:0 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:1 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:2 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:3 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    <- "isl7998x 4-003c":0 [ENABLED,IMMUTABLE]
    pad1: SOURCE
    [stream:0 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:1 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:2 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    [stream:3 fmt:UYVY8_1X16/720x480 field:none colorspace:srgb xfer:srgb ycbcr:601 quantization:lim-range]
    -> "30122000.ticsi2rx":0 [ENABLED,IMMUTABLE]
    pad2: SOURCE
    pad3: SOURCE
    pad4: SOURCE

    - entity 13: isl7998x 4-003c (5 pads, 1 link, 4 routes)
    type V4L2 subdev subtype Unknown flags 0
    device node name /dev/v4l-subdev2
    routes:
    1/0 -> 0/0 [ACTIVE]
    2/0 -> 0/1 [ACTIVE]
    3/0 -> 0/2 [ACTIVE]
    4/0 -> 0/3 [ACTIVE]
    pad0: SOURCE
    [stream:0 fmt:UYVY8_1X16/720x480 field:seq-bt]
    [stream:1 fmt:UYVY8_1X16/720x480 field:seq-bt]
    [stream:2 fmt:UYVY8_1X16/720x480 field:seq-bt]
    [stream:3 fmt:UYVY8_1X16/720x480 field:seq-bt]
    -> "cdns_csi2rx.30121000.csi-bridge":0 [ENABLED,IMMUTABLE]
    pad1: SINK
    [stream:0 fmt:UYVY8_1X16/720x480 field:seq-bt]
    pad2: SINK
    [stream:0 fmt:UYVY8_1X16/720x480 field:seq-bt]
    pad3: SINK
    [stream:0 fmt:UYVY8_1X16/720x480 field:seq-bt]
    pad4: SINK
    [stream:0 fmt:UYVY8_1X16/720x480 field:seq-bt]

    - entity 23: 30122000.ticsi2rx context 1 (1 pad, 1 link)
    type Node subtype V4L flags 0
    device node name /dev/video2
    pad0: SINK
    <- "30122000.ticsi2rx":1 [ENABLED,IMMUTABLE]

    - entity 29: 30122000.ticsi2rx context 2 (1 pad, 1 link)
    type Node subtype V4L flags 0
    device node name /dev/video3
    pad0: SINK
    <- "30122000.ticsi2rx":2 [ENABLED,IMMUTABLE]

    - entity 35: 30122000.ticsi2rx context 3 (1 pad, 1 link)
    type Node subtype V4L flags 0
    device node name /dev/video4
    pad0: SINK
    <- "30122000.ticsi2rx":3 [ENABLED,IMMUTABLE]

    - entity 41: 30122000.ticsi2rx context 4 (1 pad, 1 link)
    type Node subtype V4L flags 0
    device node name /dev/video5
    pad0: SINK
    <- "30122000.ticsi2rx":4 [ENABLED,IMMUTABLE]

  • Hi ,

    Looking through the datasheet, I see: 

    • MIPI Output
    • MIPI CSI-2 version 1.1 compliant unidirectional output
    • Standard virtual identification channel support
    • Non-standard pseudo virtual channel support
    • One or two data lanes
    • YUV422 or RGB565 output format

    Perhaps it's using the "Non-standard pseudo virtual channel support"?

    Best,
    Jared

  • Hi Jared,

    Thanks for looking at this.
    I have the non-standard bit cleared, so it operates under normal standard virtual id channel.
    For example, if I set the bit, then only stream 0 clocks data, and it has both cameras ghosted on top of each other, one camera per frame.
     i2cset -y -f 4 0x3c 0x06 0x01

    I found if I set the 8BHDR bit, then clear it, both video feeds become shifted the same way like when I see the CRC errors.
    This causes the feed to be temporarily skewed by about 45 degrees until I clear the bit.

    "1 = Add the 8-byte header in the MIPI output(1448 bytes in the long-packet payload)"

    So:
    i2cset -y -f 4 0x3c 0x06 0x20
    i2cset -y -f 4 0x3c 0x06 0x00

    I'd add an image, but that's not working for me right now.

  • Hi ,

    I should say, I was looking at the "data short". I don't have access to the entire datasheet.

    There are a lot of header, ECC, and CRC errors occurring. I do not observe those when testing CSI cameras and FPD-link cameras/serdes.

    Are there any other registers that change the data to something non-standard? 

    Best,
    Jared

  • Hi Jared,

    I noticed it's accumulating CRC errors on the csi1 peripheral in bursts over time even without the right shift.
    I imagine eventually a corrupted packet is shifting the video feed 4 pixels to the right.
    I'm not sure if there is a pattern yet.

    Perhaps there are CSI settings I need to tune or a hardware issue?
    I think I'm seeing errors occur in 2 second bursts maybe every ~500 seconds?

    Errors occurred at:
    238 seconds
    239
    280
    281
    729
    730
    771
    1219
    1220
    1262
    1263
    1710
    1711
    1752
    1753
    2200
    2201

    So it's every 500 seconds and then 40 seconds later?

    After it's shifted the data over 4 pixels, I can actually see the corrupted data 'ticking' Every second it shifts by a line up, so this probably matches about the number of lines, about 480? So maybe some kind of clock skew between the streams?

  • Hi ,

    At data rates above 1.5Gbps (which yours appears to be) requires a skew calibration to be done.

    Within our DS90UB960, we have the following register:

    There may be a similar one for your part.

    Best,
    Jared

  • Hi Jared,

    The resolution is about 720x480 pixels x 2 cameras x 8 bits  x 29.97 frames a second.
    It should be far below 1.5Gb/s.

    When I use two identical analog cameras, the 1 second 'ticking' in the corrupted section goes away, and the CRC errors too.

    In my initial setup, maybe one of my cameras is 30fps, and the other is 29.7fps?
    I think it's causing some kind of slight offset in the CSI peripheral that eventually overflows a buffer.

    I think there's still a vulnerability in the CSI peripheral or driver though, where it is unable to properly snap to the left after CRC errors during multiple virtual channel streams.
    I'll drop this for now.