This thread has been locked.

If you have a related question, please click the "Ask a related question" button in the top right corner. The newly created question will be automatically linked to this question.

[参考译文] SK-AM62A-LP:edgeai 应用存在问题

Guru**** 2952510 points
请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

https://e2e.ti.com/support/processors-group/processors/f/processors-forum/1654114/sk-am62a-lp-issue-with-edgeai-application

器件型号: SK-AM62A-LP

工程师、您好:

我正在 OD 应用领域工作。

    input0:
        source: /opt/video/sample.mp4
        format: h264
        width: 1280
        height: 720
        framerate: 20
        loop: True

在视频结束后、循环仍为真、它会卡住。 下面是 EOS 之后的 GST_DEBUG 级别 3 日志。

0:03:07.665746600  2387 0xffff54136c60 WARN              bufferpool gstbufferpool.c:1246:default_reset_buffer:<v4l2h264dec0:pool0:src> Buffer 0xffff54350dd0 without the memory tag has maxsize (0) that is smaller than the configured buffer pool size (1382400). The buffer will be not be reused. This is most likely a bug in this GstBufferPool subclass
0:03:07.665859455  2387 0xffff54136c60 WARN              bufferpool gstbufferpool.c:1246:default_reset_buffer:<v4l2h264dec0:pool0:src> Buffer 0xffff5800e020 without the memory tag has maxsize (0) that is smaller than the configured buffer pool size (1382400). The buffer will be not be reused. This is most likely a bug in this GstBufferPool subclass
0:03:07.665927894  2387 0xffff54136c60 WARN              bufferpool gstbufferpool.c:1246:default_reset_buffer:<v4l2h264dec0:pool0:src> Buffer 0xffff5800e3f0 without the memory tag has maxsize (0) that is smaller than the configured buffer pool size (1382400). The buffer will be not be reused. This is most likely a bug in this GstBufferPool subclass
0:03:07.665984936  2387 0xffff54136c60 WARN              bufferpool gstbufferpool.c:1246:default_reset_buffer:<v4l2h264dec0:pool0:src> Buffer 0xffff5800e7c0 without the memory tag has maxsize (0) that is smaller than the configured buffer pool size (1382400). The buffer will be not be reused. This is most likely a bug in this GstBufferPool subclass


此外,应用程序有时崩溃,我了解到 CMA 内存变得饥饿,但整个系统当时使用约 600MB 的内存。 终端日志是 [error]从 GST 流水线拉取帧时出错。  

此致、
Sajan

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Sajan、

    您能否分享您尝试运行的完整配置? 我假设这个问题来自 SDK 中的 apps_python? 如果不是、请提供您尝试运行的应用程序、以及 apps_python 或 apps_cpp 是否也存在问题

    此致、
    Jay

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    尊敬的 Jay:

    请分享您尝试运行的完整配置吗? [/报价]


    title: "Preview"
    
    inputs:
        input0:
            source: /opt/video/sample.mp4
            format: h264
            width: 1280
            height: 720
            framerate: 20
            loop: True
    
    models:
        model0:
            model_path: models/ONR-OD-8220-yolox-s-lite-mmdet-coco-640x640
            viz_threshold: 0.5
    
    outputs:
        output0:
            sink: kmssink
            width: 1280
            height: 720
            framerate: 20
    
    flows:
        flow0: [input0, model0, output0]
    
    

    、我假设此问题来自我们 SDK 中的 apps_python?

    不、我创建了应用程序 apps_python、然后仅通过该代码创建了 OD。  

    视频循环现在可以正常工作。

    diff --git a/apps_python/gst_wrapper.py b/apps_python/gst_wrapper.py
    index bed3734..44f222c 100644
    --- a/apps_python/gst_wrapper.py
    +++ b/apps_python/gst_wrapper.py
    @@ -78,58 +78,72 @@ class GstPipe:
             sink.set_caps(caps)
             return sink
     
    -    def pull_frame(self, src, loop):
    +    def _restart_src_pipe(self, pipe_index):
             """
    -        Pull a frame from gst pipeline
    -        Args:
    -            src: gst src element from which the frame is pulled
    -            loop: If src need to be looped after eos
    +        Destroy and recreate src pipeline for looping
             """
    +        pipe = self.src_pipe[pipe_index]
    +        pipe.set_state(Gst.State.NULL)
    +        # Wait for NULL state to complete fully
    +        pipe.get_state(Gst.CLOCK_TIME_NONE)
    +        pipe.set_state(Gst.State.PLAYING)
    +        ret = pipe.get_state(5 * Gst.SECOND)
    +        if ret[0] == Gst.StateChangeReturn.FAILURE:
    +            print("[ERROR] Failed to restart pipeline for loop")
    +            return False
    +        return True
    +
    +    def pull_frame(self, src, loop):
             sample = src.try_pull_sample(5000000000)
             if type(sample) != Gst.Sample:
                 if src.is_eos():
                     if loop:
    -                    # Seek can be called from various sources hence putting lock
                         with self.mutex:
    -                        src.seek_simple(Gst.Format.TIME, Gst.SeekFlags.FLUSH, 0)
    +                        restarted = False
    +                        for i, pipe in enumerate(self.src_pipe):
    +                            if pipe.get_by_name(src.get_name()):
    +                                restarted = self._restart_src_pipe(i)
    +                                break
    +                        if not restarted:
    +                            return None
                             sample = src.try_pull_sample(5000000000)
    +                        if type(sample) != Gst.Sample:
    +                            print("[ERROR] Error pulling frame after restart")
    +                            return None
                     else:
                         return None
                 else:
    -                print("[ERROR] Error pulling frame from GST Pipeline")
    +                print(
    +                    f"[ERROR] pull timeout: eos={src.is_eos()} "
    +                )
                     return None
             caps = sample.get_caps()
    -
             struct = caps.get_structure(0)
             width = struct.get_value("width")
             height = struct.get_value("height")
    -
             buffer = sample.get_buffer()
             _, map_info = buffer.map(Gst.MapFlags.READ)
             frame = copy.deepcopy(np.ndarray((height, width, 3), np.uint8, map_info.data))
             buffer.unmap(map_info)
    -
             return frame
     
         def pull_tensor(self, src, loop, width, height, layout, data_type):
    -        """
    -        Pull a frame from gst pipeline
    -        Args:
    -            src: gst src element from which the frame is pulled
    -            loop: If src need to be looped after eos
    -            width: width of the tensor
    -            height: height of the tensor
    -            layout: data layout (NHWC or NCHW)
    -            data_type: data type of the tensor
    -        """
             sample = src.try_pull_sample(5000000000)
             if type(sample) != Gst.Sample:
                 if src.is_eos():
                     if loop:
    -                    # Seek can be called from various sources hence putting lock
                         with self.mutex:
    -                        src.seek_simple(Gst.Format.TIME, Gst.SeekFlags.FLUSH, 0)
    +                        restarted = False
    +                        for i, pipe in enumerate(self.src_pipe):
    +                            if pipe.get_by_name(src.get_name()):
    +                                restarted = self._restart_src_pipe(i)
    +                                break
    +                        if not restarted:
    +                            return None
                             sample = src.try_pull_sample(5000000000)
    +                        if type(sample) != Gst.Sample:
    +                            print("[ERROR] Error pulling tensor after restart")
    +                            return None
                     else:
                         return None
                 else:
    @@ -141,8 +155,8 @@ class GstPipe:
                 frame = np.ndarray((1, height, width, 3), data_type, map_info.data)
             elif layout == "NCHW":
                 frame = np.ndarray((1, 3, height, width), data_type, map_info.data)
    +        frame = copy.deepcopy(frame)
             buffer.unmap(map_info)
    -
             return frame
     
         def push_frame(self, frame, sink):
    @@ -1675,3 +1689,4 @@ def get_gst_pipe(flows, outputs):
                         sink_player = add_and_link(o.gst_disp_elements, player=sink_player)
                         link_elements(s.gst_post_proc_elements[-1], o.gst_disp_elements[0])
         return src_players, sink_player
    +

    问题:
    现在、实际问题是在随机时间运行视频超过 5 分钟后、内存从~580MB 呈指数增长、系统因没有内存而卡住。

    请提供有关调试内存泄漏的说明。

    此致、
    Sajan

    [/quote]
  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Sajan、

    Gstreamer 是否已升级到 1.24 版(可解决内存泄漏问题)?

    此致

    Suren

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Suren、

    这是否将 GStreamer 升级到版本 1.24(可解决内存泄漏问题)?

    编号


    我现在无法升级到 SDK 11、因为在当前设置中已经投入了大量精力、包括模型训练和集成。 此外、SDK 11 重新训练/重新编译不能通过 Model Composer 获得、我也尝试了 EdgeAI TensorLab CLI 方法、但训练无法正常工作。

    我使用两个开箱即用的应用程序进行了测试、流切换不起作用。 您提供的流水线只会重新打开现有的摄像头流、该流按预期工作、 但切换到其他摄像头流长时间没有工作。


    此致、
    Sajan

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Sajan、

    那么、如果我理解正确、您将两个异构摄像头连接到 AM62A 板、并且您是否尝试在这些摄像头之间切换并对其进行编码? 是这样的。 这将有助于我们在最后重现这个问题

    此外、您尝试循环对 sample.mp4 进行解码并发现内存泄漏?

    这两个问题是单独的吗?

    此致

    Suren

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Suren、

    通过 arducam v3 链路板将两个 imx219 摄像头连接到 am62a 中。 尝试使用 od 和 cl 应用程序将馈送到 rtsp 中的数据流进行切换。

    这两个问题是不是分开的?

    是的。

    另外,我的 od 应用程序的长期运行似乎:
    [error]从 GST 流水线拉取张量时出错
    我使用了 https://github.com/TexasInstruments/edgeai-gst-apps/blob/main/apps_python/gst_wrapper.py#L81 中的 pull_frame () 和 pull_tensor ()

    但在循环运行视频 5 小时或更长时间时、似乎拉帧或拉张量错误。

    此致、
    Sajan

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    这是我运行应用程序 2 小时 28 分钟后出现的确切错误。

    [ERROR] Error pulling tensor from GST Pipeline

    根据我的理解、我认为这不是视频 EOS 或跳转来开始相关错误。 因为错误发生在视频的任何时间

    这不是与视频流相关的问题。 我尝试了不同的视频。
    下面是错误发生时的系统日志。

        PID     ELAPSED %CPU %MEM   RSS    VSZ
       3770    02:28:30  111  9.1 311344 1789244
    

    此致、
    Sajan

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Sajan、

    我想根据流开关问题进行响应。 您能否准确地分享您尝试做的事情?

    是否正在尝试动态切换输入传感器? 还是你想做别的事?

    此致、
    Jay

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Suren、

    是的、CAM1 提供 CL 应用程序输出帧、CAM1 提供 OD。 我尝试通过更改端口号来切换流。

    此致、
    Sajan

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Sajan、

    才能正确理解您的问题。 您正在将 IMX219 cam0 与分类功能配合使用、并将 IMX219 CAM1 与物体检测工作流程配合使用。 这是有效的。 当您在切换摄像机的情况下启动应用程序时、应用程序将失败。 这是正确的吗?

    如果是、请共享 media-ctl -p 的输出、以及两种情况下的输入和输出流水线。

    此致、
    Jay

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Jay、

    您使用的是具有分类功能的 IMX219 cam0 以及具有对象检测工作流的 IMX219 CAM1。 [/报价]

    此陈述是正确的。

    outputs:
        output0:
            sink: remote
            host: 127.0.0.1
            port: 8081
            encoding: h264
            width: 1280
            height: 720
            bitrate: 4000000
            gop-size: 15
    

        output0:
            sink: remote
            host: 127.0.0.1
            port: 8082
            encoding: h264
            width: 1280
            height: 720
            bitrate: 4000000
            gop-size: 15
    


    与此 udpsink 一样、该帧会发送到 8081 (CL) 和 8082 (OD)。

    然后启动了一个脚本、用于从该端口号切换流。

    接收器端使用的流水线为:

                f"udpsrc port={cfg['port']} address=127.0.0.1 "
                "caps=\"application/x-rtp,media=video,encoding-name=H264,payload=96\" ! "
                "rtph264depay ! h264parse config-interval=-1 ! "
                "v4l2h264dec capture-io-mode=dmabuf ! "
                "kmssink sync=false force-modesetting=true"
    

    此致、
    Sajan

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Sajan、

    当您在切换摄像机的情况下启动应用程序时、它将失败。 这是正确的吗?

    您能否确认此器件? 固有问题是什么?

    我仍然不清楚切换流意味着什么。 您是在更改动态传输哪个型号的摄像机、还是在执行其他操作。 如果您可以丢弃处理此问题的代码库部分、则可能会有所帮助。 请记住、这是一个公共论坛。

    此致、
    Jay

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    尊敬的 Jay:

    您是在动态更改哪些摄像机馈送型号、还是在执行其他操作。

    编号 模型 OD 和 CL 应用程序没有任何相关内容。 我说的是 OD 和 CL 应用程序只能将帧放入分配的端口中。 这是关于显示访问端口的流的问题。


    用于切换端口的代码是 https://docs.google.com/document/d/1fGB0UB5NguTpstUQpELoFXumcoMxu2BUg2zMCF5azEc/edit?usp=sharing

    这将在切换后的某个时间崩溃。

    此致、
    Sajan

  • 请注意,本文内容源自机器翻译,可能存在语法或其它翻译错误,仅供参考。如需获取准确内容,请参阅链接中的英语原文或自行翻译。

    您好、Sajan、

    这在理论上应该是可行的、但它需要相当数量的样板、才能使其与以下现有的 edgeai-gst-apps 配合使用。 如果这是正确的,你看到错误的拉张量从管道,下面是我认为发生的情况:

    1.切换管道时,将触发从管道中拉取帧的任务,因为它在不同的线程上运行,并且由于无法拉取帧而崩溃。
    2.我不确定整个架构、但底层的 GStreamer 流水线及相关节点也会发生变化。 因此、您需要将这些变量与任务同步。

    如果您遇到其他问题、请告诉我。

    此致、
    Jay