About The Fair Feedback Project
Why The Fair Feedback Project exists
Student evaluations of teaching (SETs) are among the most consequential and controversial documents in higher education. At more than 16,000 institutions worldwide, these instruments shape decisions about hiring, reappointment, tenure, promotion, and compensation. In many departments, they represent the only systematic metric of teaching performance and quality.
Yet decades of peer-reviewed research — spanning multiple disciplines, methodologies, and national contexts — demonstrate that SETs do not reliably measure teaching effectiveness. What they often measure instead is the degree to which an instructor conforms to students' expectations, expectations shaped by the instructor's gender, race, ethnicity, language background, age, disability status, and sexual orientation. The consequences are not abstract. Faculty who are women, people of color, non-native English speakers, and members of other marginalized groups receive systematically lower evaluation scores and disproportionately more abusive comments — even when teaching identical courses with identical materials. These biased scores then follow them into personnel files, where small numerical differences can determine who is promoted and who is not.
The Fair Feedback Project was created by Remi Kalir, PhD, in collaboration with Claude (Claude Code/Opus 4.6 & 4.8), as a practical resource that simultaneously demonstrates how generative AI may be responsibly used to advance pedagogical innovation. It was designed to address this problem at two levels. For individual instructors, it provides evidence-based, course-specific strategies and materials that can be implemented immediately. For centers for teaching and learning, department chairs, and administrators, it provides resources to support institutional reform of how student feedback is collected, interpreted, and used.
It is not a solution. It is a starting point.
The research foundation
The design of The Fair Feedback Project is grounded in a substantial body of peer-reviewed scholarship. Rather than summarizing every study, we highlight the key findings that informed our approach.
The problem: bias in student evaluations
Research consistently finds that SETs are influenced by factors unrelated to teaching quality. Among the most well-documented:
Gender bias. Studies using controlled and quasi-experimental designs — in which students evaluate identical courses taught by instructors of different genders, or in which the perceived gender of the instructor is manipulated — find that women receive significantly lower scores than men (Boring, Ottoboni, & Stark, 2016; Chávez & Mitchell, 2020; MacNell, Driscoll, & Hunt, 2015; Mitchell & Martin, 2018). Qualitative analysis of evaluation comments reveals that students evaluate women on different criteria than men, focusing more on personality and appearance and less on expertise, and that women are penalized for failing to conform to gendered expectations of warmth and nurturing (Adams et al., 2022). A review of more than 100 articles confirmed that this bias is robust across data sources and methodologies, though its magnitude is conditional on other factors including discipline, course level, and student demographics (Kreitzer & Sweet-Cushman, 2022).
Racial and ethnic bias. Faculty of color receive lower evaluation scores than white faculty, even with course content and format held constant (Chávez & Mitchell, 2020; Reid, 2010; Smith & Hawkins, 2011). Faculty from non-English-speaking backgrounds face similar penalties (Fan et al., 2019). These effects compound with gender: women of color face the intersection of both biases (Daskalopoulou, 2024; Merritt, 2008).
Structural and contextual effects. Evaluations are influenced by class size, course level, whether the course is required or elective, discipline, time of day, and grading leniency — none of which reflect teaching quality (Kogan, 2014; Kreitzer & Sweet-Cushman, 2022). A department's gender composition creates role-congruity expectations that bias evaluations of faculty in the gender minority (Aragón, Pietri, & Powell, 2023).
Abusive comments. Anonymous evaluation comments directed at women and marginalized faculty include personal attacks on appearance, intelligence, and identity. A survey of 674 academics found that previous estimates of the rate and severity of abuse had underestimated the problem, and that many institutions are aware of the issue but prioritize data collection over staff protection (Heffernan, 2023). The mental health consequences of biased and abusive evaluations are significant and underexamined (Daskalopoulou, 2024; Heffernan, 2022).
Weak relationship to learning. Meta-analyses have found little to no correlation between SET scores and student learning outcomes (Boring, Ottoboni, & Stark, 2016; Uttl, White, & Gonzalez, 2017). SETs appear to reward practices associated with student satisfaction — such as lenient grading — more than practices associated with deep learning.
What works: evidence on mitigation strategies
A smaller but growing body of research has tested interventions designed to reduce bias in SETs. The Fair Feedback Project draws on those findings while being transparent about their limitations.
Informational anti-bias statements show promise. In a randomized experiment at Iowa State University, students who received language informing them of unconscious bias and asking them to focus on course content rated female instructors significantly higher — as much as half a point on a five-point scale — with no effect on ratings of male instructors (Peterson, Biederman, Andersen, Ditonto, & Roe, 2019). A larger replication at Ohio State University across 800+ classes found that a "high-stakes" treatment emphasizing the career consequences of evaluations led to higher scores for racial and ethnic minority instructors (Genetin, Chen, Kogan, & Kalish, 2022). A field experiment at a French university found that an informational message sharing research data on bias was effective, while a normative "don't discriminate" message was not (Boring & Philippe, 2021).
Framing matters — and can backfire. A study at the University of Girona tested debiasing videos and found that a video about implicit bias generally reduced the gender gap, but a video explicitly focused on gender bias caused male students to rate female instructors lower than the control group (Ayllón & Zamora, 2025). An Australian replication found that male and female students responded to bias messaging in opposite directions (Kim, Williams, Johnston, & Fan, 2024). These findings demonstrate that how a message is framed is as important as whether a message is delivered.
Self-affirmation reduces bias through a different mechanism. Belgian researchers found that having students complete a self-affirmation exercise before evaluations eliminated gender bias — but by reducing inflated ratings for male professors rather than raising ratings for female professors (Hoorens, Dekkers, & Deschrijver, 2021).
Some interventions have shown null effects. A study at selective liberal arts colleges found that neither modified evaluation questions nor delayed timing of evaluations reduced gender disparities in qualitative comments (Owen, De Bruin, & Wu, 2025). A biology department replication of Peterson et al. yielded variable results across subfields (Mitchem et al., 2025). These null findings are important: they indicate that messaging interventions are not universally effective and should not be treated as a standalone solution.
An invitation to institutional action
If you are an instructor, staff at a center for teaching and learning or similar office, or an administrator, we want to be direct: the most impactful thing you can do is not to share The Fair Feedback Project with your colleagues, though we hope you will. The most impactful thing you can do is to change — within your respective sphere of influence — how your institution collects, interprets, and uses student evaluation data.
The evidence supports several institutional reforms. Rename evaluation instruments to emphasize student feedback rather than teaching ratings (as Augsburg University and UNC Asheville have done). Redesign evaluation questions to focus on specific student experiences and learning rather than global judgments of instructor quality. Report distributions, medians, and trends rather than means, and never use small numerical differences between instructors as evidence of differential teaching quality. Restrict or redesign open-ended comment prompts, where bias and abuse are most concentrated. Adopt a holistic approach to teaching assessment that includes peer observation, teaching portfolios, and evidence of student learning alongside student feedback. And make explicit, in tenure and promotion guidelines, that raw SET scores should not be used as a primary or standalone metric of teaching effectiveness.
Several institutions have already taken meaningful steps. In 2018, the University of Southern California stopped using student evaluations of teaching in tenure and promotion decisions, a change prompted by the provost's concern about documented bias against women and faculty of color. USC moved to a peer-review model of teaching assessment, in which peer evaluation is based on classroom observation and review of course materials, design, and assignments; faculty submit teaching reflection statements that describe how they use student feedback to improve instruction; and student evaluations were redesigned to focus on student engagement and responsibility rather than instructor ratings. Students are now asked about factors like hours dedicated to study, engagement outside class, and their own approaches to learning course material. As USC's UCAPT manual notes, student ratings and comments may be considered as indicators of student engagement, but the well-known limitations of those evaluations should be recognized. Hamilton College now requires departments to use multiple types of evidence in tenure and promotion decisions, and has explicitly prohibited using student feedback to compare faculty to one another. As Hamilton economist Ann Owen, who led the faculty committee on evaluation reform, has explained, the college can no longer conclude that one person is a better teacher than another based on more positive student feedback. Hamilton's Department Chair Handbook instructs chairs not to allow student evaluations to define good teaching exclusively, to corroborate numerical evaluations with written comments, to consider contextual issues such as class size and course level, and to discuss evidence of possible bias. At the departmental level, Hamilton's Sociology department has gone further, stating that student evaluations will be assessed within the broader gendered, racialized, and heteronormative context of the college.
Beyond these individual examples, a broader movement to transform teaching evaluation is underway. In a comprehensive survey of over 20 institutions, Michael McCreary (2026) documents a growing consensus that traditional reliance on student evaluations is inadequate and that modern approaches must incorporate multiple dimensions of teaching effectiveness assessed through multiple sources of evidence — not just the student voice, but also the instructor's self-reflection and peer review. Institutions as varied as UCLA, the University of Oregon, the University of Colorado Boulder, Clemson, Boise State, and the University of Georgia have adopted or are piloting teaching quality frameworks that define effective teaching across four to seven dimensions, evaluated through portfolios of evidence rather than a single numerical score. McCreary identifies six stages of institutional change, from recognizing the problem to reaching a sustainable steady state, and notes that even minimal policy changes — such as requiring that peer observation be available to faculty who want it — can catalyze broader cultural shifts. For institutions just beginning this work, the TEval approach, the DeLTA project, and the 2025 publication Transforming College Teaching Evaluation (Austin et al.) provide particularly useful models and resources.
The 2019 American Sociological Association Statement on Student Evaluations of Teaching, now endorsed by nearly two dozen scholarly organizations, provides an authoritative summary of the case for reform and concrete recommendations for institutions. We urge institutions to consult it.
The Fair Feedback Project recognizes that reform is slow and uneven; nonetheless, instructors need support now. The project will have succeeded most fully if it contributes, even modestly, to making itself unnecessary.
Accessibility
The Fair Feedback Project is designed for use by everyone, including instructors who rely on assistive technology. This project site has been audited against WCAG 2.1 Level AA — the international standard for digital accessibility — and revised to meet it (audit completed in June, 2026). This accessibility audit was conducted in collaboration with Claude (Opus 4.8), consistent with how the overall project was built. Practically, that means every interactive element can be reached and operated by keyboard; the question flows in the Instructor Track are structured so screen readers announce them clearly; focus moves sensibly as you navigate; generated materials and status messages are announced rather than appearing silently; and text meets contrast and readability guidelines.
In keeping with this project's commitment to transparency, we describe this as an aim we work toward, not a guarantee. Accessibility standards evolve, assistive technologies vary, and no audit catches everything. If you encounter a barrier — anything that makes the site harder to use than it should be — please tell us through the feedback form (link below in the footer), and we will work to fix it.
Selected references
Adams, S., Bekker, S., Fan, Y., Gordon, T., Shepherd, L. J., Slavich, E., & Waters, D. (2022). Gender bias in student evaluations of teaching: 'Punish[ing] those who fail to do their gender right.' Higher Education, 83, 787–807.
Aragón, O. R., Pietri, E. S., & Powell, B. A. (2023). Gender bias in teaching evaluations: The causal role of department gender composition. Proceedings of the National Academy of Sciences, 120(4), e2118466120.
Austin, A. E., et al. (2025). Transforming college teaching evaluation: A framework for advancing instructional excellence. Harvard Education Press.
Ayllón, S., & Zamora, C. (2025). Gender bias in student evaluations of teaching: Do debiasing campaigns work? IZA Discussion Paper No. 17632.
Boring, A., Ottoboni, K., & Stark, P. B. (2016). Student evaluations of teaching (mostly) do not measure teaching effectiveness. ScienceOpen Research.
Boring, A., & Philippe, A. (2021). Reducing discrimination in the field: Evidence from an awareness raising intervention targeting gender biases in student evaluations of teaching. Journal of Public Economics, 193, 104323.
Chávez, K., & Mitchell, K. M. W. (2020). Exploring bias in student evaluations: Gender, race, and ethnicity. PS: Political Science & Politics, 53(2), 270–274.
Daskalopoulou, A. (2024). Understanding the impact of biased student evaluations: An intersectional analysis of academics' experiences in the UK higher education context. Studies in Higher Education, 49(12), 2411–2422.
Fan, Y., Shepherd, L. J., Slavich, E., Waters, D., Stone, M., Abel, R., & Johnston, E. L. (2019). Gender and cultural bias in student evaluations: Why representation matters. PLoS ONE, 14(2), e0209749.
Genetin, B., Chen, J., Kogan, V., & Kalish, A. (2022). Mitigating implicit bias in student evaluations: A randomized intervention. Applied Economic Perspectives and Policy, 44, 110–128.
Heffernan, T. (2022). Sexism, racism, prejudice, and bias: A literature review and synthesis of research surrounding student evaluations of courses and teaching. Assessment & Evaluation in Higher Education, 47(1), 144–154.
Heffernan, T. (2023). Abusive comments in student evaluations of courses and teaching: The attacks women and marginalised academics endure. Higher Education, 85, 225–239.
Hoorens, V., Dekkers, G., & Deschrijver, E. (2021). Gender bias in student evaluations of teaching: Students' self-affirmation reduces the bias by lowering evaluations of male professors. Sex Roles, 84, 34–48.
Kim, F., Williams, L. A., Johnston, E. L., & Fan, Y. (2024). Bias intervention messaging in student evaluations of teaching: The role of gendered perceptions of bias. Heliyon, 10(17).
Kogan, J. (2014). Student course evaluation: Class size, class level, discipline and gender bias. In Proceedings of the 6th International Conference on Computer Supported Education (CSEDU-2014), 221–225.
Kreitzer, R. J., & Sweet-Cushman, J. (2022). Evaluating student evaluations of teaching: A review of measurement and equity bias in SETs and recommendations for ethical reform. Journal of Academic Ethics, 20, 73–84.
MacNell, L., Driscoll, A., & Hunt, A. N. (2015). What's in a name: Exposing gender bias in student ratings of teaching. Innovative Higher Education, 40, 291–303.
McCreary, M. (2026, February 10). A practical guide to modern teaching evaluation. Engaged Learning Collective (Substack).
Merritt, D. J. (2008). Bias, the brain, and student evaluations of teaching. St. John's Law Review, 82, 235–287.
Mitchem, L. D., et al. (2025). Replication of an intervention to mitigate gender bias in student evaluations of teaching yields variable results across a biology department. CBE—Life Sciences Education, 24(3), ar35.
Mitchell, K. M. W., & Martin, J. (2018). Gender bias in student evaluations. PS: Political Science & Politics, 51(3), 648–652.
Owen, A. L., De Bruin, E., & Wu, S. (2025). Can you mitigate gender bias in student evaluations of teaching? Evaluating alternative methods of soliciting feedback. Assessment & Evaluation in Higher Education, 50(3), 442–457.
Peterson, D. A. M., Biederman, L. A., Andersen, D., Ditonto, T. M., & Roe, K. (2019). Mitigating gender bias in student evaluations of teaching. PLoS ONE, 14(5), e0216241.
Reinsch, R. W., Goltz, S. M., & Hietapelto, A. B. (2020). Student evaluations and the problem of implicit bias. Journal of College and University Law, 45(1), 114–140.
Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42.