CNN-Based Shot Boundary Detection and Video Annotation

Abstract

With the explosive growth of video data, content-based video analysis and management technologies such as indexing, browsing and retrieval have drawn much attention. Video shot boundary detection (SBD) is usually the first and important step for those technologies. Great efforts have been made to improve the accuracy of SBD algorithms. However, most works are based on signal rather than interpretable features of frames. In this paper, we propose a novel video shot boundary detection framework based on interpretable TAGs learned by Convolutional Neural Networks (CNNs). Firstly, we adopt a candidate segment selection to predict the positions of shot boundaries and discard most non-boundary frames. This preprocessing method can help to improve both accuracy and speed of the SBD algorithm. Then, cut transition and gradual transition detections which are based on the interpretable TAGs are conducted to identify the shot boundaries in the candidate segments. Afterwards, we synthesize the features of frames in a shot and get semantic labels for the shot. Experiments on TRECVID 2001 test data show that the proposed scheme can achieve a better performance compared with the state-of-the-art schemes. Besides, the semantic labels obtained by the framework can be used to depict the content of a shot.

Publication
2015 IEEE International Symposium on Broadband Multimedia Systems and Broadcasting
Wenjing Tong
Wenjing Tong
Master Student

I am interested in reading and running.Also I often go to gym for health building.

Li Song
Li Song
Professor, IEEE Senior Member
Hui Qu
Hui Qu
Master Student

I enjoy running and riding bycicle. I have paticipated in Marathon races several times since 2010. Sometimes I also travel by bicycle for short journeys.

Related