This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

TDA4VH: OpenVX remapping is a lot slower than OpenCV

Part Number: TDA4VM

I was using opencv to create a remap table, and do the following exp:

1) use opencv remap

cv remapping: 8.36 ms

2) use openvx remap

159.625 ms
GRAPH: graph_70 (#nodes = 7, #executions = 1)
NODE: DSP_C7-2: node_85: avg = 45236 usecs, min/max = 45236 / 45236 usecs, #executions = 1
NODE: DSP_C7-2: node_91: avg = 7893 usecs, min/max = 7893 / 7893 usecs, #executions = 1
NODE: DSP_C7-2: node_83: avg = 45147 usecs, min/max = 45147 / 45147 usecs, #executions = 1
NODE: DSP_C7-2: node_89: avg = 7794 usecs, min/max = 7794 / 7794 usecs, #executions = 1
NODE: DSP_C7-2: node_81: avg = 45081 usecs, min/max = 45081 / 45081 usecs, #executions = 1
NODE: DSP_C7-2: node_87: avg = 7772 usecs, min/max = 7772 / 7772 usecs, #executions = 1
NODE: DSP_C7-2: node_92: avg = 330 usecs, min/max = 330 / 330 usecs, #executions = 1

The way I implemented on openvx:

    vxChannelExtractNode(graph, vxImageIn, VX_CHANNEL_R, virtChanR);
    vxChannelExtractNode(graph, vxImageIn, VX_CHANNEL_G, virtChanG);
    vxChannelExtractNode(graph, vxImageIn, VX_CHANNEL_B, virtChanB);

    vxRemapNode(graph, virtChanR, remapTable, VX_INTERPOLATION_BILINEAR, virtChanRemappedR);
    vxRemapNode(graph, virtChanG, remapTable, VX_INTERPOLATION_BILINEAR, virtChanRemappedG);
    vxRemapNode(graph, virtChanB, remapTable, VX_INTERPOLATION_BILINEAR, virtChanRemappedB);

    vxChannelCombineNode(graph, virtChanRemappedR, virtChanRemappedG, virtChanRemappedB, nullptr, vxImageOut);


for opencv, i also use `cv::INTER_LINEAR`.
In my understanding, the openvx is running on c7x, while opencv is running on mpu?
so I was expecting openVX will be faster, but it turns out its is not.

any explanation or suggestion to speed up the remapping process using openVX?


  • Hi ST,

    Due to a holiday in India, half of our team is currently out of office. Please expect a 1~2 day delay in responses.

    Apologies for the delay, and thank you for you patience.

    Regards,
    Takuma

  • Hi,

    In the description, I see that you have mentioned TDA4VM.

    In this SOC, these nodes should have run on C66 DSP core.

    But the logs, I see that you are running this on C7x_2 core. 
    Could you please let me know which SOC are you running this application?

    Regards,

    Nikhil

  • Actually, I use j784s4 evm for now.

  • Hi,


    cv remapping: 8.36 ms

    Does this also split the RGB channel, then remap each channel and combine back again on A72?

    GRAPH: graph_70 (#nodes = 7, #executions = 1)
    NODE: DSP_C7-2: node_85: avg = 45236 usecs, min/max = 45236 / 45236 usecs, #executions = 1
    NODE: DSP_C7-2: node_91: avg = 7893 usecs, min/max = 7893 / 7893 usecs, #executions = 1
    NODE: DSP_C7-2: node_83: avg = 45147 usecs, min/max = 45147 / 45147 usecs, #executions = 1
    NODE: DSP_C7-2: node_89: avg = 7794 usecs, min/max = 7794 / 7794 usecs, #executions = 1
    NODE: DSP_C7-2: node_81: avg = 45081 usecs, min/max = 45081 / 45081 usecs, #executions = 1
    NODE: DSP_C7-2: node_87: avg = 7772 usecs, min/max = 7772 / 7772 usecs, #executions = 1
    NODE: DSP_C7-2: node_92: avg = 330 usecs, min/max = 330 / 330 usecs, #executions = 1

    May I know what does each node here correspond to ? could you give the names to these nodes?

    May I also know the resolution of the RGB image?

    Regards,

    Nikhil

  • the image is from 3840x2160 -> remap to 320x180

    This is the same code with some additional reference names:

    GRAPH: graph_70 (#nodes = 7, #executions = 1)
    NODE: DSP_C7-2: extractB: avg = 48223 usecs, min/max = 48223 / 48223 usecs, #executions = 1
    NODE: DSP_C7-2: remapB: avg = 7524 usecs, min/max = 7524 / 7524 usecs, #executions = 1
    NODE: DSP_C7-2: extractG: avg = 48211 usecs, min/max = 48211 / 48211 usecs, #executions = 1
    NODE: DSP_C7-2: remapG: avg = 7490 usecs, min/max = 7490 / 7490 usecs, #executions = 1
    NODE: DSP_C7-2: extractR: avg = 48218 usecs, min/max = 48218 / 48218 usecs, #executions = 1
    NODE: DSP_C7-2: remapR: avg = 7482 usecs, min/max = 7482 / 7482 usecs, #executions = 1
    NODE: DSP_C7-2: combine: avg = 337 usecs, min/max = 337 / 337 usecs, #executions = 1

  • Hi,

    The numbers mentioned seems to be in alignment with the testing done at TI's end. Please refer the datasheet below 
    TIOVX User Guide: J784S4 Linux Performance Report

    Does this also split the RGB channel, then remap each channel and combine back again on A72?

    If only remap node is considered, then we have around 7.5 msec. So only remap happens on OpenCV?

    Regards,

    Nikhil

  • Hi Nikhil,

    In OpenCV, I only need to provide two single channel map mapX and mapY

    And the OpenCV can do the rest for me, which means I can use remap directly on rgb image.

    So there's no need to separate channel and remap then combine in OpenCV.

    But in OpenVX, it seems that the remap operation only support single channel input, so I have to do what I did in the aforementioned code.

  • Hi,


    The numbers mentioned seems to be in alignment with the testing done at TI's end. Please refer the datasheet below 
    TIOVX User Guide: J784S4 Linux Performance Report

    Understood. As mentioned above, the numbers seems to aligned with the performance report of the node and in the sdk, remap operation only support single channel input.

    Also please note that on J784s4, these kernels are rebuilt on C7 but not optimized for the same..

    They are the same C66 kernels that is rebuilt for C7x.

    Regards,

    Nikhil

  • Hi Niki,

    Do you have any suggestion to speed up?

    It sounds crazy to me that mpu is faster than dsp in this case haha.

    Or this is the best I can get, so I have to stick to the OpenCV implementation now?

  • Hi,

    Currently the suggestion from TI would be to stick with the OpenCV implementation as the kernels are not optimized for C7x DSP.

    Regards,

    Nikhil