Project Activities
The work comprised two major components. First, the research team developed new tools for measuring the messages contained in images and in text. These included creating a way to measure the representation of skin color, race, gender, and age of pictured characters in both photographs and illustrations, and to measure the messages contained in text when people of different identities were mentioned (spanning ethnicity, race, gender, sexuality, and other dimensions previously beyond the reach of existing tools). Second, the project applied these tools, alongside other state-of-the-art tools, to measure gender- and race-based messages in multiple sets of materials often used in education. One corpus included a century of award-winning children’s books commonly found in libraries, classrooms, and homes. A second was the corpus of textbooks used in core subjects in elementary schools in the state of Texas between 1985–2011. A third corpus was a random sample of 250,000 articles from two major newspapers, the Wall Street Journal and the New York Times, also commonly used in educational settings. The team also paired these data with administrative data on the demographics of consumers to assess the relationship between exposure to these messages and the traits of their consumers.
Structured Abstract
Setting
The study focused the educational materials available to children across the United States, including a century of award-winning children’s books commonly found in homes, libraries, and classrooms. In addition, the study examined 25 years of books used by the Texas Department of Education, which instituted a standardized, required corpus of textbooks across the entire state over the study period and offers rich public-access data on student outcomes.
Sample
The sample of children’s books came from 19 awards recognized by the American Library Association; the sample of textbooks came from the administrative records of the Texas Department of Education; the data on child outcomes came from publicly available data at the cohort-by-subgroup (gender and race) level acquired from the Texas Student Data System. According to these data, over the period of this study, the number of students in Texas public schools ranged from 3.2 to 4.8 million; the proportion of these students who were Black ranged from 12–14%, and the percentage who were Latinx ranged from 35–53%. An additional sample of 500,000 news articles (250,000 each from the New York Times and Wall Street Journal) from 1/1/2000 to 6/30/2022 were analyzed.
The factors that were measured in this project were: 1) the levels and manner of gender and race representation contained in children’s books and school textbooks (as well as other materials), as measured by different machine-led methods for estimating them, and 2) the relationship between these messages and the demographic (race and gender) and political characteristics (political views and voting behavior, as collected by the Cooperative Election Study) of the people who consume them across different levels of exposure to these messages.
Research design and methods
The project’s research design had three main parts:
- Methodological: The project generated new tools to measure several features of image and text data which were previously beyond the reach of frontier methods.
- Application: The project applied these tools, along with other frontier tools, to characterize the messages about race and gender contained in a large body of print media, including a century of children’s books, three decades of textbooks from the states of Texas and California and from homeschooling settings, and half a million newspaper articles from 2000 to 2022.
- Analytic: The project analyzed the traits of consumers of these books over time, spanning demographics, purchasing behavior, and political views, to understand patterns over time and across space in the consumption of these books
Control condition
Not applicable.
Key measures
The study included two key sets of measures: (1) the levels of the targeted messages in the images and text of each set of educational materials; and (2) the relationship between these messages and the demographic (race and gender) and political views (political beliefs and voting behavior) of the consumers of these materials.
Data analytic strategy
The first step in the project was to collect the data, converting printed materials into digital data on images and text. The second step was to use new methods from machine learning, computer vision, and natural language processing to quantify the implicit and explicit messages about ability by gender and race contained in the images and text of the books. Finally, the project conducted quantitative analysis of the correlation between these measurements and the political and demographic traits of the consumers, and how this varied by geography
Key outcomes
The main outcomes of this project are as follows:
- The project developed and applied a series of new computational tools to measure the way skin color, race, gender, and age are represented in images that can be applied to a wide range of educational and other settings (Adukia et al., 2023; Szasz et al., 2022).
- The project developed and applied a series of new computational tools to measure the way that groups (including, but not limited to, those defined by race, gender, age, and sexuality, as well as the intersection of these identities) are measured in text. (Adukia et al., 2023; Adukia, Chiril, et al. 2022; Adukia, Christ, et al., 2022).
- Results showed that: 1) people are more likely to consume books that are centered on their own identities than others; 2) the types of children’s books purchased in a given locality correlate with local political beliefs; 3) books that center nondominant social identities cost more in major retailers than books which center dominant social identities; and 4) there are fewer copies of books that center nondominant social identities in libraries that serve predominantly White communities than there are in other communities. These findings advanced our understanding of the persistence of inequality in representation in books commonly used in education (Adukia et al., 2023).
- When examining a century of award-winning children’s books found in a large portion of American homes, libraries, and schools, it was found that the way these books portray gender reproduces traditional gender norms in society. Specifically, results showed that: 1) relative to males, females are more likely to be represented in relation to their appearance than in relation to their competence; 2) females are more likely to be represented in relation to their role in the family than their role in business; and 3) it finds that non-binary or gender-fluid individuals are rarely mentioned (Adukia, Chiril, et al. 2022).
People and institutions involved
IES program contact(s)
Project contributors
Products and publications
Find available citations in ERIC for this award here.
Project website:
Publications:
Inside IES Research Blog
Chhin, C. S. (2021, February 11). Representation Matters: Exploring the Role of Gender and Race on Educational Outcomes. Inside IES Research.
Journal articles
Adukia, A., Eble, A., Harrison, E., Runesha, H. B., & Szasz, T. (2023). What We Teach About Race andGender: Representation in Images and Text of Children’s Books. The Quarterly Journal of Economics, 138(4), 2225-2285.
Proceedings
Adukia, A., Chiril, P., Christ, C., Das, A., Eble, A., Harrison, E., & Runesha, H. B. (2022, October). Tales and Tropes: Gender Roles from Word Embeddings in a Century of Children’s Books. In Proceedings of the 29th International Conference on Computational Linguistics (pp. 3086-3097).
Adukia, A., Christ, C., Das, A., & Raj, A. (2022, June). Portrayals of Race and Gender: Sentiment in 100 Years of Children’s Literature. In Proceedings of the 5th ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies (pp. 20-28).
Questions about this project?
To answer additional questions about this project or provide feedback, please contact the program officer.