The WWW 2025 Multimodal Dialogue Intent Recognition competition is jointly organised by Taobao and Tmall Group (淘天集团), the WWW 2025 conference, and the Tianchi platform. It is drawn directly from real e-commerce customer-service operations, where shoppers reach out both by chatting with an agent and by sending screenshots of what they are looking at.
Given either a multi-turn Chinese conversation between a customer and a service agent, or a screenshot taken inside a shopping app, the model has to work out what the user actually wants and return a single, fine-grained category label. The text and image inputs are judged together by overall accuracy on a held-out test set, so the real difficulty is labelling consistently and precisely across dozens of closely related categories.
用户:<image>
客服:您好!您对这款商品有什么想了解的吗?
用户:这款风扇灯安装时对电线有没有特殊要求?
客服:需要连接零线、火线和地线;地线没预留可不接,零线火线必须接上。
用户:如何操控风扇
客服:选好款式颜色后点击图片可看功能参数,支持遥控器操作,对应图片上会有遥控器图标。
candidates: 用法用量 · 控制方式 · 能否调光 · 功效功能 · …