Watch and track your favorite playlist.
Curated by: 김성범[ 교수 / 산업경영공학부 ] (327 videos)
멀티모달 대규모 언어모델(MLLM)은 이미지와 텍스트를 함께 이해하며 다양한 과제를 수행할 수 있지만, 입력 이미지와 일치하지 않는 내용을 생성하는 hallucination 문제가 주요 한계로 지적되고 있다. 이를 완화하기 위해 모델의 생성 행동을 직접 교정하는 학습 기반 방법, 추론 과정에서 시각적 근거를 강화하는 decoding 기반 방법, 그리고 최근에는 모델 내부의 시각 정보 처리 과정을 분석하고 제어하는 접근까지 다양한 연구가 제안되고 있다. 본 세미나에서는 MLLM hallucination의 기본 개념과 주요 원인을 살펴보고, 대표적인 완화 방법들이 어떻게 발전해왔는지를 소개하고자 한다. 참고자료: [1] Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H. T., Sun, M., & Chua, T. S. (2024). RLHF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional human feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [2] Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., & Bing, L. (2024). Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [3] Jiang, Z., Chen, J., Zhu, B., Luo, T., Shen, Y., & Yang, X. (2025). Devils in middle layers of large vision-language models: Interpreting, detecting and mitigating object hallucinations via attention lens. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.