Appearance descriptions still carry graded gender associations; 16 LLMs compress them
The GAPA dataset shows 316 appearance attributes still carry structured gender associations among 304 US annotators, while 16 LLMs only partially reproduce them.
ImportanceLocalEvidenceE2 unreplicated
Readers can now verify that seemingly "objective" appearance descriptions still carry structured, graded gender associations in human interpretation: the GAPA dataset built by Yingjia Wan, Lin Lin and Elisa Kreiss covers 316 common appearance attributes, with 304 US annotators providing 14,706 gender-association ratings.
Previously, substituting descriptions for gender labels was treated as neutral communication, at the cost of hiding the strength and graded structure of the associations these descriptions themselves carry.
The authors report that 16 LLMs only partially reproduce the human ratings: the distributions are compressed, alignment with male-associated items is weaker, and abstentions cluster asymmetrically on the non-binary category; the dataset, code and proxy prediction model have been released by the authors.
The result does not cover non-US annotator populations and has not yet been independently reproduced; the preprint is arXiv 2609.16366, with the arXiv page noting publication at COLM 2026.