How AI Can Shape Trauma-Informed Evaluation

Image of the author

October 2026

About the author: Eyerusalem Tessera is a Project Lead at Three Hive Consulting and a Credentialed Evaluator with a Master of Public Health and Master of Business Administration. Eyerusalem’s background is grounded in evaluation, healthcare, and consulting. She brings expertise in interest holder engagement, facilitation, and evaluation design, with a focus on helping organizations use evidence to support learning, improvement, and decision-making.


This article is rated as:

Evaluation is my main role
 

AI is increasingly being used to support evaluation activities, including developing interview guides, generating survey questions, organizing qualitative data, identifying themes, summarizing findings, and drafting reports. While evaluators remain responsible for data collection, interpretation, and decision-making, AI can still shape how participant experiences are understood and represented.

From a trauma-informed perspective, AI may influence how the principles of trauma-informed evaluation are applied: Safety; Trustworthiness and Transparency; Peer Support; Collaboration and Mutuality; Empowerment, Voice and Choice; and Cultural, Historical, and Gender Responsiveness.

Three challenges are particularly important: bias, difficulty interpreting intersectional experiences, and the tendency to simplify complex experiences into coherent narratives.


1. AI May Reproduce Existing Biases 

Research demonstrates that AI systems can produce different outcomes for different groups, even when they appear objective. For example, a widely used health-care algorithm was found to systematically underestimate the needs of Black patients in the US because it used health-care spending as a proxy for health need, thereby reproducing existing inequities in access to care[1]. Other studies have found that language-processing systems are more likely to classify African American English as abusive or problematic and that large language models continue to reproduce gender stereotypes, associating women more frequently with domestic roles and men with leadership and professional occupations[2].

In evaluation, these biases may affect much more than analysis. AI-generated interview questions and probes may unknowingly embed assumptions about participants, communities, or social issues. There might be questions that are inappropriate and undermine the safety of participants despite fitting nicely within the scope of the evaluation. During coding and synthesis, AI may assign different meanings to similar experiences depending on how they are expressed. AI-generated reports may also reinforce dominant cultural perspectives while under-representing the experiences of equity-deserving groups.

[1] https://www.science.org/doi/10.1126/science.aax2342

[2] https://arxiv.org/pdf/1905.12516


2. AI Often Struggles with Intersectionality

Trauma-informed evaluation requires evaluators to understand how identities, histories, and systems of power interact. Experiences are rarely shaped by race, gender, disability, culture, income, sexuality, or age in isolation. Rather, people experience these factors simultaneously.

Current research suggests that AI is not consistently effective at interpreting intersectional experiences. Studies examining intersectional bias in language models have found that model performance often varies most significantly at the intersection of identities rather than within any single demographic group. Research has found substantial differences in how models respond to combinations of race and gender, disability and gender, or other intersecting identities[1].

One challenge is that AI is trained to identify patterns across large datasets. Smaller groups often appear less frequently in training data and may therefore be interpreted less accurately. AI may recognize individual identities while missing the way systemic factors such as racism, ableism, sexism, poverty, colonialism, or discrimination interact to shape lived experience.

For evaluators, this creates particular challenges when working with communities, racialized participants, people with disabilities, gender-diverse individuals, Indigenous people, newcomers, or participants whose experiences sit at multiple intersections of marginalization. AI may identify demographic categories without recognizing the structural forces that give those experiences meaning.

[1] https://arxiv.org/html/2508.07111v1


3. AI Tends to Smooth Over Complexity

Lastly, AI’s tendency to create coherent narratives undermines participants’ voices. Large language models are designed to identify patterns, summarize information, and produce concise explanations. While these abilities can be useful, trauma-informed evaluation often requires evaluators to sit with contradictions, ambiguity, competing truths, and diverse experiences. Not all participant experiences fit neatly into common themes.

AI-assisted analysis may prioritize dominant patterns and overlook less common yet highly important findings. In qualitative data, a single participant's experience may reveal a critical issue. Because AI frequently prioritizes recurring responses, these less common experiences may receive less attention or disappear entirely. We have also found AI to perform well at identifying major themes but less consistently at identifying less frequent findings that may still be highly significant, as discussed in our AI Qualitative Analysis article.

This tendency to simplify complexity can also affect reporting. Participant stories may be transformed into broad categories. Along the way, nuance, disagreement, uncertainty, and context may be lost. The resulting report may be coherent and well-written while failing to fully represent the complexity of participant experiences.

From a trauma-informed perspective, this matters because participant stories are not merely data points. How people understand and communicate their experiences is often central to the meaning of those experiences.

Ultimately, AI can support efficiency and pattern recognition, but trauma-informed evaluation requires human interpretation, contextual understanding, and relationship-based sensemaking. Evaluators should therefore be intentional when using AI to overcome its shortcomings. To reduce the risks, consider the following strategies:

  • Review AI-generated questions and probes for safety and appropriateness. Consider whether questions and probes are safe and how they might affect individuals with trauma.

  • Be transparent about AI use and protect participant confidentiality. Clearly communicate how AI is used within the evaluation process and ensure appropriate safeguards are in place for sensitive information.

  • Use AI-generated outputs as a starting point, not the final analysis. All themes, interpretations, and recommendations should be reviewed and validated by the evaluator.

    • Instead of asking AI for a single summary, evaluators conduct multiple analyses. For example, starting with major themes, followed by contradictory experiences, equity and intersectionality review, and strengths and resilience. Conduct a review at each stage for omissions, hallucinations, and oversimplification. This prevents AI from collapsing everything into one dominant story.

  • Actively look beyond dominant themes. Review less common responses, contradictory findings, and outlier experiences, recognizing that frequency does not determine importance.

    • One overlooked limitation of AI is that it often identifies barriers, risks, and problems because these are easier to detect as patterns.

    • To prevent findings from becoming overly deficit-focused, intentionally ask AI to identify resilience, coping strategies, participant strengths, community assets, examples of healing and growth, and participant-defined success.

  • Use AI to surface complexity rather than summarize it. Most evaluators use AI to reduce information. A trauma-informed approach may instead use AI to identify complexity.

    • For example, instead of asking: "What are the key findings?" ask: "What tensions, contradictions, or competing experiences are present in the data?" or "Which participants appear to have experienced the program differently and why?" This shifts AI from finding consensus to exploring variation.

  • Preserve participant voice and context. Use participant language and direct quotes where appropriate and ensure findings reflect the complexity and nuance of lived experiences rather than only AI-generated summaries.

  • Engage participants, peer advisors, Elders, and community partners in interpreting findings. Collaborative sensemaking can help identify biases, validate interpretations, and ensure findings resonate with lived experience.


Check out our other articles: Using AI To Do An Environmental Scan and Trauma informed evaluation articles and resources.

Trauma informed evaluation practice checklist

Building trauma informed evaluation tools

Previous
Previous

Trauma Informed Evaluation: Red Flags to Watch Out For

Next
Next

Designing Reports People Actually Want to Read - Previously Recorded Webinar