Watch and track your favorite playlist.
Curated by: 김성범[ 교수 / 산업경영공학부 ] (327 videos)
이미지 및 비디오 이상탐지에는 vision 기반 모델들이 널리 활용되어 왔다. 그러나 이러한 모델들은 시각적 픽셀 정보에만 의존하기 때문에, 이미지 내 semantic한 특성을 충분히 반영하기 어렵다는 한계가 있다. 이를 보완하기 위해, 최근에는 언어 정보를 함께 활용하여 다양한 특성을 추가적으로 고려할 수 있는 vision-language model (VLM) 기반 이상탐지 연구가 활발히 수행되고 있다. 특히, OpenAI의 CLIP이나 Alibaba의 Qwen 등 빅테크 기업에서 개발한 foundation VLM들이 등장하면서, 이들의 풍부한 사전 지식을 기반으로 이상탐지 성능을 크게 향상시키는 연구들이 활발히 등장하고 있다. 본 세미나에서는 이러한 foundation VLM을 활용하여 이미지 및 비디오 이상탐지를 수행한 최신 연구 사례들을 살펴보고자 한다. 참고자료: [1] Abdalla, M., Javed, S., Al Radi, M., Ulhaq, A., & Werghi, N. (2025). Video anomaly detection in 10 years: A survey and outlook. Neural Computing and Applications. [2] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., ... & Sutskever, I. (2021, July). Learning transferable visual models from natural language supervision. In ICML. [3] Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual instruction tuning. Advances in neural information processing systems, 36, 34892-34916. [4] Wei-Lin, C., Zhuohan, L., Lin, Z., Ying, S., Wu, Z., Hao, Z., ... & Ion, S. (2023). Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. LMSYS. [5] Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., & Yuan, L. (2024, November). Video-llava: Learning united visual representation by alignment before projection. In EMNLP. [6] Jiang, X., Li, J., Deng, H., Liu, Y., Gao, B. B., Zhou, Y., ... & Zheng, F. (2025) MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection. In ICLR. [7] Li, Y., Yuan, S., Wang, H., Li, Q., Liu, M., Xu, C., ... & Zuo, W. (2025). Triad: Empowering LMM-based Anomaly Detection with Expert-guided Region-of-Interest Tokenizer and Manufacturing Process. In ICCV. [8] Chen, Z., & Imani, F. (2026). A multi-expert framework for enhancing multimodal large language models in industrial anomaly detection. Pattern Recognition, 112752. [9] Ye, M., Liu, W., & He, P. (2025). Vera: Explainable video anomaly detection via verbalized learning of vision-language models. In CVPR.