Researchers use the women-are-wonderful effect to name a relative evaluation pattern: on some attitude and trait measures, the broad category women receives more favorable responses than the broad category men. It does not mean women are objectively better, that every respondent shows the pattern, or that favorable ratings produce favorable treatment.

Those limits are central to the definition. A person can rate a group as warm yet doubt its members' authority. An average response can change with the traits on the questionnaire, the social role attached to the target, the respondent, the culture, and the historical moment.

A fictional profile check

The community workshop, volunteer profiles, reviewers, category labels, scores, and outcome in this section are invented. This is an illustration, not study data or a real organization.

An imaginary workshop is testing a new form for selecting volunteer coordinators. Two fictional profiles list the same experience and the same record of completed projects. One profile identifies the applicant as a woman; the other identifies the applicant as a man.

Before anyone compares overall impressions, the reviewers score four separate questions: How much evidence shows warmth? Competence? Reliability? Ability to perform the coordinator's actual tasks? For each score, they must point to a line in the profile.

The exercise does not assume that the reviewers will produce a gender difference. If one appears, it would not prove a hidden motive or establish a general truth about women and men. It would show where the form, the evidence, and the category cue deserve a closer look.

What the classic studies measured

The label grew from research on attitudes and gender stereotypes, not from a single test of behavior.

In a 1989 study, Alice Eagly and Antonio Mladinic asked respondents to evaluate women, men, and two political categories. The same respondents rated each category, so the design allowed direct comparison but also made the comparison conspicuous. Attitudes toward women and the evaluative content of the female stereotype were more favorable than the corresponding responses about men.[^eagly1989]

A 1991 experiment changed the design. Its 324 Purdue psychology students, evenly divided between men and women, evaluated one assigned social category rather than rating both women and men side by side. The researchers used five kinds of measures: a general attitude scale, free-response beliefs, a belief list, free-response emotions, and an emotion list.[^eagly1991]

Women received more favorable scores on the attitude measure and both belief measures. The two emotion measures did not show a significant women-versus-men difference, and respondents did not show more ambivalence toward women than men. The result was therefore specific: a difference in broad attitudes and attributed qualities in one student sample, not a uniform difference across every kind of response.

A later review reached an equally important conclusion. Favorable attitudes toward women as a general category could coexist with prejudice in masculine or male-dominated domains.[^eagly1994] “Liking” a broad category is not the same judgment as deciding who fits a particular job, who seems authoritative, or whose work counts as competent.

What implicit gender-attitude tasks add

An Implicit Association Test, or IAT, compares performance across sorting conditions. In a gender-attitude version, names categorized as female or male may share response keys with positive or negative words. Faster performance in one combined condition is interpreted as a difference in relative association under that task.

John Skowronski and Melissa Lawrence compared explicit ratings with IAT variants among fifth-graders and college students. Their tasks used general gender names or names of male and female soldiers. Explicit ratings favored women overall. Response latencies indicated pro-female attitudes in the soldier condition and among women and college students, but the error results included a pro-male pattern among male respondents in the general-gender condition.[^skowronski]

That mixture supports a bounded conclusion: some female-name IAT studies have found faster positive associations. It also blocks a stronger claim. Age group, respondent gender, target framing, latency, and error rate did not tell one perfectly consistent story.

Laurie Rudman and Stephanie Goodwin reported four experiments in which women's automatic own-group preference was stronger than men's. The studies examined links with gender identity, self-esteem, parental associations, associations between men and violence, and, among sexually experienced men, attitudes toward sex.[^rudman]

These are findings about task performance and proposed correlates. They do not show that women possess an innate tendency to favor women, that men lack positive regard for men, or that a particular score predicts how someone will treat another person.

Why “positive” depends on the item

An overall favorable score compresses several judgments into one number. That can hide the content doing the work.

The classic research linked women's favorable evaluation largely to communal qualities, such as being helpful, gentle, kind, or understanding.[^eagly1991] These are positive descriptors. They are not synonyms for competence, agency, status, or power.

The stereotype content model formalizes part of this problem by treating warmth and competence as separate dimensions. Across nine varied samples, Susan Fiske and colleagues found that social stereotypes often combined high warmth with low competence or high competence with low warmth.[^fiske] A group can therefore receive praise on one dimension while being discounted on another.

Cross-national evidence also shows this split. A 16-nation study involving 8,360 participants found that traits associated with men were less positively valenced but more associated with power than traits associated with women.[^glick2004] A single “Who is viewed more positively?” question would miss that tradeoff.

This is why the effect cannot be used to infer that women are seen as better leaders, stronger candidates, more credible witnesses, safer people, or more capable decision-makers. Those are separate claims requiring evidence from the relevant task and setting.

Culture and time change the comparison

The effect has appeared beyond one U.S. campus, but its size is not fixed.

Kuba Krys and colleagues reanalyzed face ratings from 4,519 post-secondary students across 44 cultures. Participants rated photos of women and men on eight items related to honesty and intelligence without being explicitly asked for gender stereotypes. Across the full sample, the female-target advantage was small, d = .16.[^krys]

The difference varied with national gender egalitarianism. Its correlation with the researchers' composite measure was r(42) = -.50. In the ten least egalitarian samples, the target difference was d = .32. In the ten most egalitarian samples, it was d = .05 and marginal rather than conventionally significant (p = .055).

The study broadens the evidence and supplies counterevidence to a universal reading. Its participants were students, the targets were decontextualized photographs, only smiling and neutral expressions were used, and attractiveness was not fully controlled. “Observed across the combined sample” is not the same as “present in every culture or decision.”

Historical evidence makes a related point. Across the period from 1946 through 2018, a cross-temporal meta-analysis synthesized responses from 30,093 adults in 16 U.S. polls designed to represent the national population. Respondents considered whether traits related to communion, agency, and competence were more true of women, more true of men, or equally true of both.[^eagly2020]

Women's relative advantage in communion increased over time. Men's relative advantage in agency did not show a temporal change. Competence judgments moved toward equality, with some female advantage among respondents who perceived a difference. The stereotype was not one stable line called “favorability.” Different trait domains followed different paths.

The effect is not the same as benevolent sexism

Benevolent sexism refers to a specific cluster of subjectively positive but paternalistic and role-restricting attitudes. In the original Ambivalent Sexism Inventory research, it was measured separately from hostile sexism, even though the two components were positively related.[^glick1996]

The concepts can overlap, but they are not interchangeable. A warm evaluation of women does not by itself establish protectionism, paternalism, heterosexual intimacy beliefs, or support for restricted roles. Conversely, a favorable-sounding belief can still carry a condition about how women should behave.

A useful follow-up to a compliment is therefore not “Is this secretly sexist?” It is more concrete: What quality is being praised, what evidence supports it, and does the praise come with a narrower expectation for the person's role?

What the finding cannot establish

The research does not justify statements that women are more moral, men are less worthy, or one gender is naturally disposed toward particular traits. It does not estimate how many people “have” the effect. It does not diagnose an individual from one rating or response-time score.

It also cannot settle whether women receive fair treatment. Positive stereotypes, negative stereotypes, mixed stereotypes, and unequal outcomes can coexist. Evidence about a global category rating cannot erase evidence from hiring, promotion, health care, education, law, or relationships. Each domain needs its own measures.

The binary categories used in much of the literature are another boundary. They do not represent every identity or intersection of gender with age, race, ethnicity, class, sexuality, occupation, and culture.

Sources

[^eagly1989]: Eagly, Alice H., and Antonio Mladinic. “Gender Stereotypes and Attitudes Toward Women and Men.” Personality and Social Psychology Bulletin 15, no. 4 (1989): 543-558. https://doi.org/10.1177/0146167289154008

[^eagly1991]: Eagly, Alice H., Antonio Mladinic, and Stacey Otto. “Are Women Evaluated More Favorably Than Men? An Analysis of Attitudes, Beliefs, and Emotions.” Psychology of Women Quarterly 15, no. 2 (1991): 203-216. https://doi.org/10.1111/j.1471-6402.1991.tb00792.x

[^eagly1994]: Eagly, Alice H., and Antonio Mladinic. “Are People Prejudiced Against Women? Some Answers From Research on Attitudes, Gender Stereotypes, and Judgments of Competence.” European Review of Social Psychology 5, no. 1 (1994): 1-35. https://doi.org/10.1080/14792779543000002

[^skowronski]: Skowronski, John J., and Melissa A. Lawrence. “A Comparative Study of the Implicit and Explicit Gender Attitudes of Children and College Students.” Psychology of Women Quarterly 25, no. 2 (2001): 155-165. https://doi.org/10.1111/1471-6402.00017

[^rudman]: Rudman, Laurie A., and Stephanie A. Goodwin. “Gender Differences in Automatic In-Group Bias: Why Do Women Like Women More Than Men Like Men?” Journal of Personality and Social Psychology 87, no. 4 (2004): 494-509. https://doi.org/10.1037/0022-3514.87.4.494

[^glick1996]: Glick, Peter, and Susan T. Fiske. “The Ambivalent Sexism Inventory: Differentiating Hostile and Benevolent Sexism.” Journal of Personality and Social Psychology 70, no. 3 (1996): 491-512. https://doi.org/10.1037/0022-3514.70.3.491

[^fiske]: Fiske, Susan T., Amy J. C. Cuddy, Peter Glick, and Jun Xu. “A Model of (Often Mixed) Stereotype Content: Competence and Warmth Respectively Follow From Perceived Status and Competition.” Journal of Personality and Social Psychology 82, no. 6 (2002): 878-902. https://doi.org/10.1037/0022-3514.82.6.878

[^glick2004]: Glick, Peter, et al. “Bad but Bold: Ambivalent Attitudes Toward Men Predict Gender Inequality in 16 Nations.” Journal of Personality and Social Psychology 86, no. 5 (2004): 713-728. https://doi.org/10.1037/0022-3514.86.5.713

[^krys]: Krys, Kuba, et al. “Catching Up With Wonderful Women: The Women-Are-Wonderful Effect Is Smaller in More Gender Egalitarian Societies.” International Journal of Psychology 53, suppl. 1 (2018): 21-26. https://doi.org/10.1002/ijop.12420

[^eagly2020]: Eagly, Alice H., Christa Nater, David I. Miller, Michèle Kaufmann, and Sabine Sczesny. “Gender Stereotypes Have Changed: A Cross-Temporal Meta-Analysis of U.S. Public Opinion Polls From 1946 to 2018.” American Psychologist 75, no. 3 (2020): 301-315. https://doi.org/10.1037/amp0000494