Skip to content
HN On Hacker News ↗

No Easy Fix for Bogus Respondents in Online Opt-In Polls

▲ 61 points • 37 comments • by luu • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,507
PEAK AI % 0% · §1
Analyzed
Sep 23
backend: pangram/v3.3
Segments scanned
1 windows
avg 1507 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,507 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

About this research This study was designed to measure the impact of bogus respondents on opt-in surveys. It also compares three methods for identifying and removing bogus respondents: 1) the use of trap questions, sometimes called “attention checks,” 2) the Sentry prescreening system proprietary to CloudResearch, and 3) matching respondents to a national voter file. Why did we do this? Pew Research Center does high-quality research to help the public, the media and decision-makers understand important topics. Our methodological research investigates the current challenges facing the polling industry, including past work on the impact of bogus respondents in opt-in surveys. Learn more about Pew Research Center and our methodological research. How did we do this? We fielded a large online opt-in survey Nov. 14-19, 2024, among 11,114 U.S. adult respondents. Respondents who agreed to provide their name and contact information were matched to a registered voter file after the survey concluded. We evaluated how each approach for removing bogus cases performed on three data quality metrics: “yea-saying,” or agreeing regardless of what is asked, 2) quality of open-end text responses, and 3) response order effects. We also examined how screening methods substantively affected estimates of voter turnout and vote choice in the 2024 presidential election. Here are the questions used for this report the survey methodology. One of the most urgent problems in online opt-in polling is bogus (or fraudulent) respondents. These are survey-takers who make no effort to answer questions truthfully and instead are just looking to finish surveys quickly and collect rewards. To combat this threat, researchers have developed various ways to identify and purge bogus cases from survey samples. Approaches include 1) trap questions, sometimes called “attention checks,” that genuine respondents should always answer correctly, 2) automated prescreening services, and 3) matching respondents to a registered voter file. But how well do they work? A new Pew Research Center study finds that: Overall, purging bogus cases lowers error on most metrics, but a surefire solution remains elusive. Matching an opt-in sample to voter files slightly increased error by removing mostly good respondents (e.g., those who simply declined to give their name or address). Trap questions and an automated prescreening service performed similarly, improving data quality somewhat. All three approaches modestly increased the overestimation of Democratic support in the 2024 election. This appears to be due not to a systematic partisan bias but to bogus respondents’ tendency to say they voted for the winning candidate – in this case, Donald Trump, the Republican. Matching opt-in samples to voter files may serve other purposes, such as providing data on respondents’ voting history. This study speaks only to whether this is an effective tool for purging bogus cases. Do all polls have bogus respondents? No. Bogus respondents are primarily a threat to online opt-in polls, which are recruited through methods like online advertising, self-enrollment and email lists. Polls recruited offline using random sampling (e.g., Pew Research Center’s American Trends Panel) are generally immune to this threat, though they face other challenges. Related: Do AI and bogus respondents threaten polling’s future? Methods for removing bogus respondents Studies conducted by Center researchers and others have found that standard data quality checks such as looking for “speeders” who complete surveys too quickly or “straightliners” who consistently select the same answer choice (e.g., always pick the first response option) fail to detect most bogus respondents. The same goes for many trap questions or attention checks, which not only fail to detect bogus respondents but can also confuse legitimate respondents, leading to false positives. Faced with these challenges, pollsters have developed a variety of new approaches in hopes of better identifying bogus respondents. In this study, we evaluate three such approaches: Trap questions about unlikely behaviors Automated prescreening by a leading fraud-detection company Matching respondents to a commercial voter file Trap questions A difficulty in identifying bogus respondents lies in the fact that, for most questions, we have no way of knowing if a respondent has answered truthfully. To get around this problem, we asked respondents if they had ever engaged in four activities where we can be virtually certain that the true answer is “No.” Under this screening procedure, respondents were coded as bogus if they answered “Yes” to one or more of these questions. The questions were designed so that diligent respondents would not be confused about how to answer, while bogus respondents trying to appear eligible for surveys targeting specific groups or answering randomly would be more likely to answer “Yes.” Specifically: We asked respondents if they used any of six different social media platforms, one of which was a made-up platform called Fizzypress. We also asked if they had received any of six different government benefits in 2023, including payments from the United States Railroad Administration (USRA) – a real government agency, but one that ceased to exist in 1920. Both of these were asked early on in the survey and were intended to resemble screening questions looking for users of specific social media platforms or recipients of certain kinds of benefits. Of the full sample, 6% said they used Fizzypress and 9% claimed to have received USRA payments. Later in the survey, we asked respondents if they had ever visited the International Space Station; 10% said they had. Finally, we asked if they had ever served on a Polar-class icebreaker ship, only two of which were ever built and only one of which remains in active use by the U.S. Coast Guard; 8% answered in the affirmative. A total of 1,963 cases (18%) answered “Yes” to one or more of these trap questions and were flagged as bogus. Automated prescreening CloudResearch’s proprietary Sentry prescreening system is designed to automatically identify problematic respondents before they begin a survey. When a potential respondent first clicks the survey invitation link, they are routed to the Sentry platform and asked a short series of questions designed to elicit problematic survey-taking behaviors like yea-saying and inattentive responding. Open-end answers are checked to confirm that they are responsive to the question asked and not pasted from another source via an automated process. The system also performs passive checks using metadata about the respondent’s device, location and browser for other signs that they are misrepresenting themselves through technical means. Respondents who pass these checks are then routed to take the survey. Typically, those who fail are not forwarded to the main survey, but for this study, no respondents were terminated for failing the preescreening checks. We used metadata about which cases failed and why to simulate what would have happened to the survey results if they were excluded at the outset. Under this procedure, 5,369 cases – nearly half of all completes – failed at least one prescreening check, including 1,879 (35% of failed cases) that failed multiple. 70% of failed cases had a problematic open-end, making this the most frequently failed check by far. This was followed by yea-saying (47%), inattentiveness (12%) and duplicate IP addresses (10%). 7% of these cases failed checks for fraudulent behavior, and only 1% failed other passive technical checks. Voter file matching The third approach we tested involves asking respondents to provide their name and address and then looking for a matching record in a commercial voter file, a national database of nearly all registered voters in the United States. If a matching record can be found, the case is considered valid. If no matching record is found, the case is thrown out. This method assumes that respondents who can be matched are most likely being honest about their identity, and that someone willing to provide detailed contact information that can be validated against official voting records will likely be diligent about answering other survey questions. This approach has primarily been used by pollsters who work for political campaigns. For cases that are successfully matched, voter files can provide a great deal of information beyond what was asked in the survey, such as a respondent’s registration status and voting history. But while many voter file vendors attempt to include the unregistered population, Center research has found that a sizable share of this group is not covered by these databases. This can lead to throwing away data for many otherwise-valid respondents who simply aren’t registered to vote. This is not a large drawback for political pollsters, who typically focus on surveying registered voters. But if certain kinds of people who are more likely to be unregistered are underrepresented in a sample or missing altogether, it may lead to biased results if applied to a general population survey. In this study, we asked respondents if they were willing to provide their name and home address so that they could be matched to the voter file. Out of all respondents, 49% agreed to provide their contact information; of these, 73% were successfully matched to the TargetSmart voter file based on name, address, age and sex. Altogether, 3,977 cases (36% of the full sample) were successfully matched while 7,137 (64%) were screened out under this procedure. Distinct demographic patterns in flagged cases