Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,550 words · 1 segments analyzed
In the very first paper session at CHI 2026, I found myself thinking, “I wish I wrote this paper!” And then I found myself thinking that a few times more. Since I’m on the hunt for new research questions, it seems worth digging into why I had these reactions. I’ll do so here. For each paper, I’ll address the following questions: What is the paper about? Why do I wish I wrote it? Could I, in fact, have written it? What next steps does the paper inspire? Monday, April 13, 11:51 am Liu, Alicia T. H., Mina Lee, and Xuechunzi Bai. 2026. “Writing with AI Can Reduce Gender Bias in Hiring Evaluations.” Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (New York, NY, USA), CHI ’26, April 13, 1–30. https://doi.org/10.1145/3772318.3791136. What is the paper about? The authors contribute a large-N, between-subjects experiment in which participants evaluate résumés from “male” and “female” job applicants (“John” and “Jennifer”). Participants compose their evaluations using a writing tool with LLM-based autocomplete suggestions. For “John” these suggestions are neutral with respect to gender stereotypes; for “Jennifer,” these suggestions may be neutral, or they may enforce or counter gender stereotypes. (As an aside: I question whether these AI suggestions are or can be truly neutral.) The authors report that counter-stereotypical suggestions increased Jennifer’s perceived competence and likelihood of being selected as a trusted leader. Such suggestions also brought Jennifer’s salary offers up to parity with John’s. However, counter-stereotypical suggestions activated gender backlash, with Jennifer seen as less likable than in the other two conditions. John remained the preferred candidate across all experimental conditions. Participants were largely unaware of the intent to manipulate gender stereotype activation, with open-ended feedback focused mainly on the utility of the autocompletion-based writing support. Why do I wish I wrote it? It’s a natural post-LLM follow-up to the currently-abandoned project Reading for Gender Bias, to which I contributed in summer 2020. In a nutshell, Reading for Gender Bias provides academic recommendation letter writers with advice for reducing gender bias in their letters. The project predates the current LLM craze, but when I asked for advice in 2022 and 2023 about how to move it forward, it was universally recommended to integrate LLMs as a source of revisions and advice. I always had a feeling that those reading recommendation letters as part of making hiring decisions might be a more valuable audience than those who were writing them. I often thought about impact beyond academia – and not only because it’s so hard to obtain a corpus of academic recommendation letters. And even before joining Reading for Gender Bias in 2020 – even before joining Whitman in 2015 – I was curious about the potential for persuasive technology focused on language use to affect people’s attitudes and biases. I appreciate how this paper uses suggestions to integrate debiasing into task performance rather than operating at the reflective or metacognitive level. I’m further attracted to this study in part because its potential applications are ethically murky – the kind of problem that brought me to persuasive technology in 2006 when I was on the cusp of completing my PhD and embarking on a new research program. I find myself wondering if it would be more ethical and more effective to use tools that obscure candidates’ gender, as in the gender-blind orchestra auditions pioneered in the 1970s and 80s. Hmm, I wonder if that experiment has already been done. Finally, it received a CHI Best Paper Award. With its social implications, I feel like it deserves some visibility beyond CHI. Could I, in fact, have written it? Probably not. I’ve hesitated to conduct formal experiments as part of my research. I’ve even tried, without success, to seek collaborators with expertise in experimental methods. At the same time, this kind of experiment seems within my reach – particularly with an experimental collaborator and with student support for software development. It’s probably no accident this paper was written by two psychologists in collaboration with a computer scientist. I still wish I’d thought of it first. What next steps does the paper inspire? Regarding the gender-blinded résumé experiment, I’ve thought before about developing “gender-neutralizing” text manipulation tools in the context of Degender the Web and in the context of developing training data for Reading for Gender Bias. I might have a potential collaborator in Cambridge behavioral economist Konstantinos Ioannidis (who is the spouse of a Cybercrime Centre PhD student). But first, I should probably find out if such an experiment has already been done. Another direction would be to return to debiasing academic recommendation letters, à la Reading for Gender Bias. I don’t think that’s useless – a number of colleagues have told me they use Thomas Forth‘s Gender Bias Calculator to get feedback on potential bias in their recommendation letters. How would I approach that problem differently after reading this paper? First, I’m curious about using LLMs to generate autocomplete suggestions, rather than the hand-coded keyword-search approach taken in that earlier work. Or perhaps suggestions could be integrated with the spellchecker-style feedback approach taken by the current prototype. Second, I think my focus would be on building a system good enough to evaluate experimentally, before building a system good enough to deploy. How would I approach the recommendation letter problem differently than the candidate evaluation problem? First, I think the tool and study would need to be transparent about the intention to manipulate gender stereotypes. Second, given the goal of the manipulation, it seems like there should be some kind of external assessment of gender bias in the resulting letters. And finally, study participants would need richer scenarios to write from, going beyond the fictional applicant’s résumé. Tuesday, April 14, 9:36 am Chanenson, Jake, Tara Matthews, Sunny Consolvo, et al. 2026. “‘It Didn’t Feel Right but I Needed a Job so Desperately’: Understanding People’s Emotions and Help Needs During Scams.” Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (New York, NY, USA), CHI ’26, April 13, 1–22. https://doi.org/10.1145/3772318.3790556. What is the paper about? The authors examine 405 Reddit posts “seeking help for a range of known and emerging scams.” The authors aim to understand why people engage with scammers, how they feel at different stages of the scam, and what kinds of help they seek. They look for the tactics scammers use to elicit emotional responses along with the contextual factors that increase susceptibility. The end goal is to inform interventions that reduce potential targets’ vulnerability to online financial scams. Across 12 different scam types (Beals, et al., 2015), the authors found five fundamental emotional motivations for target engagement: fear, hope, trust, guilt, and belonging. Across the stages of the User States Framework (Matthews, et al., 2025), the authors identified four types of help needs: sensemaking, guidance, emotional support, and action. Factors that elevate risk include financial, legal, and employment precarity, as well as neurodiversity and mental health conditions. Proposed interventions focus on “just-in-time” messaging (Intille, 2004) rather than the preventative education being explored elsewhere (e.g., Deng, et al., 2026, also on my reading list from CHI). This approach is more likely to be effective in high-stakes events, but also more difficult to get right. Why do I wish I wrote it? As a member of the Cambridge Cybercrime Centre, I felt obliged to come to this session on scams. This paper showed me a potential connection between cybercrime and persuasion, or more specifically, emotional manipulation. A question I have is whether scammers use classic influence strategies à la Robert Cialdini to manipulate their targets. The work also points the way towards persuasive technologies that might help potential victims steer clear of scams. This is a kind of connection I might have hoped to make myself. I also found the paper personally relevant. One of the strangest phone calls of my life came on a summer night from a former student who was having a tough time on the job market. As in the title of the paper, they had received a job offer that “didn’t feel right.” They wanted my help figuring out if the job offer was legitimate or a scam – the Diagnostic substate in the expanded User States Framework developed in this paper. Like the scam targets whose help requests are analyzed in this paper, my student was made vulnerable by their hope. They called me seeking help with making sense of the offer and guidance about whether to accept it, as well as emotional support. Finally, I’ve always admired the work of Sunny Consolvo, who I met at UW in Seattle and had long conversations with at PERSUASIVE 2009 in Claremont, CA (my college stomping grounds). Although she wasn’t the presenter, I see her fingerprints on the paper. I was sitting behind her at the presentation but didn’t get a chance to say hello. Could I, in fact, have written it? Probably not. First, I don’t have access a large team of dedicated researchers. Second, my interest in cybercrime is far too new. On the other hand, the data is publicly available. The analysis methods strike me as similar to those used in the Cybercrime Centre to learn about cybercrime methods and motivations – albeit with a focus on criminals rather than their targets.