Hyeryeon Park, Bohyun Choi, Yerim Choi

2026.6.11Fashion and Textiles

DOI: 10.1186/s40691-026-00475-w

Abstract

The rapid proliferation of AI-generated images in the fashion industry has created a need for systematic evaluation methods to assess garment consistency between original and AI-generated images. Conventional quantitative metrics often fail to capture fine-grained garment attributes, while human evaluation, though accurate, is costly and difficult to scale. This study proposes an automated evaluation method leveraging Vision–Language Models (VLMs) to assess garment consistency in AI-generated images. To enable systematic evaluation, we developed a garment-specific evaluation framework by operationalizing DeLong’s visual definers into 20 attributes, which were embedded into the prompt to guide the VLMs. To validate the proposed method, experiments were conducted using a real-world dataset of original fashion images paired with AI-generated ghost mannequin photography. While traditional quantitative metrics failed to effectively capture garment consistency, the proposed method demonstrated substantial alignment with human evaluation on overall garment consistency. Compared to human evaluation, the proposed method successfully identified inconsistencies across the defined attributes; however, at the attribute level, it showed higher sensitivity to color and texture but lower sensitivity to shape and line dimensions. These findings suggest that VLM-based evaluation can effectively complement existing evaluation methods by providing scalable and theoretically grounded insights.

Citation format

PARK, Hyeryeon; CHOI, Bohyun; CHOI, Yerim. A VLM-based framework for evaluating garment consistency in AI-generated images based on delong’s theory. Fashion and Textiles, 2026, 13.