نوع مقاله : مقاله های برگرفته از رساله و پایان نامه
عنوان مقاله English
نویسنده English
Road semantic segmentation plays a key role in scene understanding for autonomous and assisted driving systems, but accurate pixel-level labeling of each frame of road videos is expensive and time-consuming. In this paper, a semi-supervised learning framework for semantic segmentation of road videos is introduced, which relies on exploiting motion constraints extracted from point tracking. In this approach, only the first frame of each 15-frame sequence is manually pixel-labeled, and for subsequent frames, pseudo-label masks are generated using point tracking by the CoTracker3 model. These pseudo-labels are then used to semi-supervisedly train a lightweight LR-ASPP-based segmentation network with a MobileNetV3 backbone. The combined error function consists of a supervised component on labeled frames and a pseudo-supervised component considering tracked pixels. The proposed framework is evaluated on the KITTI-STEP dataset, which is an extended version of KITTI for segmentation and tracking of each pixel in video. The results show that the proposed semi-supervised model SSL-Tracker achieves an mIoU of 27.67% and a pixel accuracy of 72.58%, recording a 2.54% improvement in mIoU over the fully supervised base model trained only with labeled frames (25.13%). A more detailed analysis of the classes shows that the use of point tracking provides a significant advantage, especially for dynamic classes such as “car” and “bicycle”. These results indicate that a significant improvement in semantic road segmentation can be achieved by exploiting the motion information embedded in the video without increasing the number of manual labels.
کلیدواژهها English