多模態 AI 如何重塑工廠工安與自動光學檢測的新未來 How Multimodal AI is Reshaping the Future of Factory Safety and Automated Optical Inspection
摘要 Abstract
多模態模型的高泛化能力已是AI產業化關注的重點。 本演講將探討多模態大型語言模型 (VLM) 在智慧工廠中的應用,包括透過 VLM 實現工安檢查(如著裝檢查、煙火偵測、門禁管控及禁制區域監測等)。接著將介紹如何利用自監督學習 (SSL) 微調大型視覺模型,快速構建針對特定領域的檢測任務。最後,強調大型模型的泛化能力及其在零樣本與少樣本場景中的卓越表現,助力智慧工廠實現安全高效的運營目標。
The strong generalizability of multimodal models has become a key focus in AI industrialization. This talk will explore the applications of multimodal Vision-Language Models (VLMs) in smart factories, including their use in safety inspections such as dress code checks, fire/smoke detection, access control, and restricted area monitoring. It will also discuss leveraging self-supervised learning (SSL) to fine-tune large vision models, enabling rapid development of domain-specific detection tasks. The talk will emphasize the generalizability of large models and their outstanding performance in zero-shot and few-shot scenarios, driving smart factories toward safer and more efficient operations.







