Transactions of the Association for Computational Linguistics · 2026
CROC Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
Why this publication matters
A score for text-to-image generation should notice when an image misses a crucial part of a prompt. CROC tests that ability using closely contrasted examples, including human-checked cases. It also uses these examples to improve an evaluation model, connecting the discovery of weaknesses to a concrete training approach.
Abstract
The assessment of evaluation metrics (metaevaluation) is crucial for determining the suitability of existing metrics in text-toimage (T2I) generation tasks. Humanbased meta-evaluation is costly and timeintensive, and automated alternatives are scarce. We address this gap and propose CROC: a scalable framework for automated Contrastive Robustness Checks that systematically probes and quantifies metric robustness by synthesizing contrastive test cases across a comprehensive taxonomy of image properties. With CROC, we generate a pseudo-labeled dataset (CROCsyn) of over 1 million contrastive prompt–image pairs to enable a fine-grained comparison of evaluation metrics. We also use this dataset to train CROCScore, a new metric that achieves state-of-the-art performance among open-source methods, demonstrating an additional key application of our framework. To complement this dataset, we introduce a human-supervised benchmark (CROChum) targeting especially challenging categories. Our results highlight robustness issues in existing metrics: for example, many fail on prompts involving negation, and all tested open-source metrics fail on at least 24% of cases involving correct identification of body parts.
Figures
Cite this paper
@article{leiter2026crocevaluatingand92,
title = {{CROC Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks}},
author = {Christoph Leiter and Yuki M. Asano and Margret Keuper and Steffen Eger},
journal = {Transactions of the Association for Computational Linguistics},
year = {2026},
url = {http://hdl.handle.net/21.11116/0000-0013-3F33-C}
}
Figures and abstract are reproduced from the linked research sources. Credit remains with the authors and publishers.