
6/23/2022 · Xinyang Guo, Baoyu Han
What this post added
Details the implementation of video deduplication using Milvus, including the workflow of extracting feature vectors from video frames, indexing these vectors in Milvus, and performing similarity searches to identify duplicate videos. The system architecture involves Kafka for data ingestion, deep learning models for feature extraction, Milvus for vector indexing and search, Ceph for storage, and TiDB/Pika for video ID mapping. The similarity search process is described in three steps: batch recall, refined search based on video IDs, and final similarity scoring.