This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

TDA4VH-Q1: Questions Regarding Maximum Encoder Instances and Parallel Compression Latency on TDA4VH (SDK 11)

Part Number: TDA4VH-Q1
Other Parts Discussed in Thread: TDA4VH

When testing the video compression functionality on the TDA4 VH, we found that during parallel testing with 64 channels, the maximum number of instances that can run is 32. When the number exceeds 32, errors occur. In the SDK 11 source code, we found that the defined MAX_NUM_INSTANCE is 32. If we want to support 64 instances, do we need to modify the SDK source code?
In addition, we have two questions regarding the use of gst-launch-1.0 to test compression performance:

1.When testing parallel compression, should we launch multiple gst-launch-1.0 processes simultaneously, with each process running one compression instance, or should we launch a single gst-launch-1.0 process and run multiple compression instances within that process?


2.When testing compression latency, we observed that the latency for a single instance is 7.67989e+06, while running four instances simultaneously results in a latency of 3.27993e+07. Our question is: in a parallel compression scenario, should the latency of multiple instances be the same as that of a single instance, or is an increase in latency expected?

BC15FF7F27C9B6B7EE5C337E89794579.PNGC12A09FF41DC058989F3ED6650AD2250.JPG

  • Hello, 

    If we want to support 64 instances, do we need to modify the SDK source code?

    The MAX number of stream instances that the Wave5 supports is 32 Instances per VPU core. On the TDA4VH there are 2 VPU cores so as long as you split the number of instances between the two cores this is supported without modifying the source.

    1.When testing parallel compression, should we launch multiple gst-launch-1.0 processes simultaneously, with each process running one compression instance, or should we launch a single gst-launch-1.0 process and run multiple compression instances within that process?

    You would run a single GStreamer process with multiple compression instances. An example would look like this below: 

    • GST_DEBUG_FILE=gst_trace.log GST_DEBUG_NO_COLOR=1 GST_DEBUG="GST_TRACER:7" GST_TRACERS="latency(flags=element):v4l2" gst-launch-1.0 filesrc location=airshow_p720x480_nv12.264 ! h264parse ! tee name=t \
      t. ! queue ! v4l2h264dec capture-io-mode=4 ! fakevideosink \
      t. ! queue ! v4l2h264dec capture-io-mode=4 ! fakevideosink \
      t. ! queue ! v4l2video2h264dec capture-io-mode=4 ! fakevideosink \
      t. ! queue ! v4l2video2h264dec capture-io-mode=4 ! fakevideosink
    2.When testing compression latency, we observed that the latency for a single instance is 7.67989e+06, while running four instances simultaneously results in a latency of 3.27993e+07. Our question is: in a parallel compression scenario, should the latency of multiple instances be the same as that of a single instance, or is an increase in latency expected?

    I believe the latency should be around the same. In fact I have an FAQ here, [FAQ] TDA4VH-Q1: Analysis on Latency Wave521CL CODEC IP , that you can reference on latency information. If you run the python script on your gstreamer log you can see the isolated latency measurements of the v4l2h264/5 encoder/decoder.

    Thank you,
    Sarabesh S.

  • Hello, Sarabesh :

    1、How can video streams be distributed across the two VPU cores for compression? When we use gst-launch-1.0 to test 64 parallel streams, the maximum number of streams that can be compressed is still limited to 32. Errors occur when the number exceeds 32. The following image shows the error messages from dmesg.

    2、Based on the parallel compression command you provided, we wrote a test script to measure compression latency under different levels of parallelism. We found that the latency for a single stream is not the same as for multiple streams; instead, it increases proportionally with the number of parallel compression instances. For example, the latency is 4.85 for a single stream and increases to 15.92 when running four streams in parallel, as shown in the test results below. Could you please help confirm whether our testing method is correct?

    and this is our test script.Thanks

    #!/bin/bash

    NUM_STREAMS=$1

    PIPELINE="GST_TRACERS='latency(flags=pipeline+element)' GST_DEBUG=GST_TRACER:7 GST_DEBUG_FILE='latency_server.txt' gst-launch-1.0 -e "

    PIPELINE+=" filesrc location=/dev/shm/test_1080p.yuv ! \
    rawvideoparse width=1920 height=1080 format=nv12 framerate=30/1 colorimetry=bt709 ! tee name=t "

    for i in $(seq 1 $NUM_STREAMS); do
    PIPELINE+="t. ! queue ! v4l2h265enc output-io-mode=dmabuf ! filesink location=output_$i.265 "
    done

    echo "$PIPELINE"
    eval "$PIPELINE"

  • Hi j G, 

    Wasn't available to get back to you on this today. I will take a look and let you know.

    Thanks,
    Sarabesh S.

  • Hi Sarabesh,

    Any updates on this?

    Thanks,
    J. G.

  • Hi J.G.

    Apologies for the delay, got occupied on some priority tasks.

    How can video streams be distributed across the two VPU cores for compression? When we use gst-launch-1.0 to test 64 parallel streams, the maximum number of streams that can be compressed is still limited to 32. Errors occur when the number exceeds 32. The following image shows the error messages from dmesg.

    I should've explained this on my last response but as you can see Gstreamer command I ran, there are two v4l2 decoders that I call (i.e. v4l2h264dec and v4l2video2h264dec). You can see the list of encoder and decoders by running gst-inspect-1.0. Could you try running it with 32 streams across each of the two cores.

    2、Based on the parallel compression command you provided, we wrote a test script to measure compression latency under different levels of parallelism. We found that the latency for a single stream is not the same as for multiple streams; instead, it increases proportionally with the number of parallel compression instances. For example, the latency is 4.85 for a single stream and increases to 15.92 when running four streams in parallel, as shown in the test results below. Could you please help confirm whether our testing method is correct?

    I will check on your test method and get back to you. 

    Thanks,
    Sarabesh S.

  • Hi Sarabesh,


    Any updates on the second question?


    Thanks,
    J. G.

  • No updates as of now. B/w has been limited, I will get back to you.

    Thanks,
    Sarabesh S.

  • Hello Unlocking this thread.

    Apologies for the delay. Your test method does look correct. However, I do want to run this on my end and get back to you on why that is the case. 

    Regards,
    Sarabesh S.