Sign UpLogin With Facebook
Sign UpLogin With Google

Nonresponse bias: when the people who skip your poll change the answer

A low response rate is a warning light, not a verdict. Bias appears only when the reason people stay silent is connected to the answer they would have given.

  • Low rate is not bias
  • Worked example
  • Detect it yourself
  • Informal polls too

Make a poll here, free. No account needed, and you get a link you can send anywhere.

Polling 1 million people every day since 2003

Amazon, Men's Health, Weebly, Tumblr, YouTube and eHow have run polls with Poll Maker

What nonresponse bias actually is

Nonresponse bias is the gap between the answer your poll reports and the answer you would have got if everyone you invited had replied. It is a different thing from a low response rate, and the difference is the most useful idea on this page.

Methodologists break it into two parts. Nonresponse bias is the nonresponse rate multiplied by the difference between respondents and nonrespondents on the thing you are measuring. The National Academies review states the consequence plainly: high nonresponse rates could yield low nonresponse errors if the difference between respondents and nonrespondents is quite small.

Multiply anything by zero and you get zero. If the people who ignored your poll would have answered the same way as the people who replied, you can lose most of your sample and still land on the right number. If the silence is connected to the answer, a fifth missing can wreck the result.

A low response rate raises the ceiling on how wrong you could be. It does not tell you how wrong you actually are.

Bias needs non-participation to correlate with the answer. No correlation, no bias, however few people replied.

How many did not replyThe nonresponse rate on its own. A number most people fixate on.
x
How differently they would have answeredThe gap between the people who replied and the people who did not.
=
Nonresponse biasSet either side to zero and the bias is zero. It is only large when both are.

One consequence catches people out. Bias belongs to a question, not to a survey. The National Academies review notes that the response rate is not directly linked to bias and is not variable specific. The same poll, to the same list, on the same day, can be near perfect on one question and badly out on the next.

The evidence is stronger than most people expect. Groves and Peytcheva assembled 59 methodological studies that had each measured nonresponse bias directly, against rich sampling frames, matched administrative records and follow-ups of nonrespondents. Across that body of work they found very little correlation between the nonresponse rate and the measured bias.

Pew Research Center has run the same test on itself for three decades. Its telephone polls fell from a 37 percent response rate in 1996 to 9 percent in 2014, then 7 percent in 2017 and 6 percent in 2018. Across 13 benchmarked lifestyle, health and demographic questions the average absolute error was 2.7 percentage points in 2016 against 2.8 points in 2012. The rate collapsed and the accuracy did not move. Yet the same study found a question on which that 9 percent poll was 38 points out.

One telephone survey. One response rate of about 9 percent. Eight very different amounts of bias, decided only by what was asked.
Worked with neighbors to fix a problem38pts
Joined a sports or recreation group16pts
Contacted a public official15pts
Joined a civic or service association9pts
Rated own health as excellent or very good8pts
Registered to vote7pts
Received food stamps4pts
Typical demographic question, average of 143pts

Pew Research Center, What Low Response Rates Mean for Telephone Surveys (2017). Gap between the Pew phone estimate and a high response rate government benchmark. The health item is an understatement by the poll, the rest are overstatements. The final bar is the average absolute gap across 14 demographic and personal measures.

The pattern is not random. Answering a survey is itself a small act of civic-mindedness, so the people who answer are more civic-minded than the people who do not. Ask about volunteering and the bias is enormous. Ask household size and it vanishes, because that has nothing to do with willingness to talk to a pollster.

There is no response rate threshold. Figures get quoted, usually between 50 and 80 percent, below which a poll is supposedly invalid. No such threshold exists in the survey methodology literature, and the chart above is why.

If someone names a cutoff, ask them which question they mean.

A worked example: same response rate, opposite reliability

The numbers below are invented to illustrate the arithmetic. They are not data. Everything else on this page is sourced.

A company with 1,000 employees runs two internal polls in the same week, to the same list. Both get exactly 200 replies, a 20 percent response rate. A response rate audit would grade them identically.

Poll A asks which of two logo colors people prefer. Whether you bother to answer a company poll has nothing to do with your taste in logos, so repliers and non-repliers hold near identical views.

Poll B asks whether people are satisfied with their manager. Here the two are tightly connected. Employees unhappy with their manager are more likely to doubt the survey is really confidential, more likely to be disengaged, and less likely to reply at all. The silence is not noise. The silence is the finding.

Illustration only, not data. Two polls in the same 1,000 person company, both with a 20 percent response rate.
Poll A: logo colorPoll B: manager satisfaction
People invited1,0001,000
People who replied200 (20 percent)200 (20 percent)
What the poll reports61 percent prefer blue78 percent satisfied
Truth among the 800 who stayed silent60 percent prefer blue45 percent satisfied
Gap between repliers and non-repliers1 point33 points
Bias: 0.80 nonresponse rate times the gap0.8 points26.4 points
True figure across all 1,00060.2 percent51.6 percent

Poll A is fine. Reporting 61 percent when the truth is 60.2 percent is an error smaller than the rounding most people apply anyway. Poll B is a disaster. It reports 78 percent satisfaction when barely half the workforce is satisfied, off the same 200 responses and the same 20 percent rate that made Poll A trustworthy.

No amount of statistical confidence rescues Poll B. Sent to 2,000 employees for 400 replies at the same rate, it still reports 78 percent, just with a tighter margin of error. Sample size shrinks random noise and does nothing to bias. Read that alongside how accurate poll results actually are.

How to detect nonresponse bias in your own poll

You cannot measure bias directly, because measuring it would need the answers you did not get. You can triangulate. These are the methods national survey teams use, and every one shrinks to a poll of 300 customers.

  1. Compare respondents to something you already know

    If you polled your own staff, customers or members, you already hold the true profile of the whole list: department, tenure, plan tier, region, purchase recency. Cross-tabulate the people who replied against the list they came from. Any group that is short in your responses is a group whose views are under-weighted in the result. Pew does exactly this at national scale, against government surveys with response rates of 60 percent or more.

  2. Run a wave analysis

    Sort responses by arrival time, or by how many reminders it took. Late responders are the best available proxy for people who never responded, so if your headline number drifts steadily from the first wave to the third, project the drift and see how far it could still travel. Treat it as a hint, not proof: the National Academies review warns that pushing hard on reluctant people can degrade data quality rather than improve it.

  3. Chase a random sample of the silent

    Take 30 people at random from those who did not reply and pursue them properly: a phone call, a different channel, a real incentive, a three question version. Twenty answers from the silent majority beat another thousand from the willing minority, because they give you a direct estimate of the gap term.

  4. Match the silent group to a record

    Where a record exists, use it. Pew matched the phone numbers it called against a national voter file, which let it describe people who never picked up. Most organisations have an equivalent: order history, login frequency, support tickets. You cannot know what a nonrespondent thinks, but you usually know what they did.

  5. Ask the correlation question, item by item

    For each question, ask plainly whether the people who skipped it would have answered differently. For logo color the honest answer is no, and a low response rate is survivable. For satisfaction, engagement or intent to leave, the honest answer is usually yes, and no sample size will rescue you.

The fourth step gives the most quotable evidence, because it puts both groups in one table.

Pew matched the phone numbers in its 2016 sample against a national voter file, so the people who answered could be compared directly with the people who did not.
MeasureAnswered the surveyDid not answer
Registered to vote85 percent81 percent
Voted in the 2012 elections62 percent52 percent
Voted in the 2014 elections49 percent33 percent
Mean 2016 turnout propensity score, 0 to 1007769

Read that the way you should read your own. The gaps are modest on registration and much larger on turnout in a midterm year. Nonresponse pulled the sample toward habitual voters, so anything correlated with the voting habit is biased here and anything unrelated to it is not.

To make the first step possible, add profile questions whose true distribution you can look up in your own records.

How long have you been a customer?

  • Less than 6 months
  • 6 to 24 months
  • More than 2 years
  • I am not a customer

You already know the true tenure split of your list, so this doubles as a bias check rather than just a segmentation field.

How often do you use the product?

  • Daily
  • Weekly
  • Monthly or less
  • I have stopped using it

Usage frequency is the most common source of nonresponse bias in product surveys. Heavy users reply, lapsed users vanish, and lapsed users hold the criticism.

Which team are you in?

  • Sales
  • Engineering
  • Operations
  • Support
  • Something else

For internal polls, compare against headcount by team. A department that is 20 percent of staff and 6 percent of your responses is a hole in the result, not a rounding issue.

None of these are opinion questions, which is the point: they exist so you can hold your responses against a known truth. More on writing clean answer options in survey question examples, and on picking a response format in levels of measurement.

3 questions in this set. Copy them all, or one at a time.

How to reduce it, and what only feels like it works

Almost all advice about nonresponse is really advice about response rates. They are not the same target, and confusing them is how a budget gets spent without an estimate improving.

Prepaid incentives

The best-evidenced way to lift a response rate. A meta-analysis of 21 years of experiments across mail, telephone and in-person surveys found a strong, non-linear effect, with prepaid cash in mail producing the largest response per dollar.

Buys response rate reliably. Buys accuracy only if the extra people differ from the ones you already had.

Target the groups you are short of

Spend the effort where the hole is. If your responses are 6 percent Operations and your company is 20 percent Operations, another all-staff reminder is worthless while one conversation with an Operations lead is not.

The only lever that attacks the gap term rather than the rate term.

Make the poll shorter

Length is one of the very few things you control completely, and the people who abandon a long poll are rarely a random subset of your list. Cut every question you cannot name a decision for.

Reduces both whole-poll and per-question dropout, and costs nothing.

Switch channel for the hard to reach

A different channel reaches a genuinely different kind of person. Email plus SMS plus an in-product prompt pulls a different mix than any one alone, and the mix is what you are trying to change.

Changes who is in the sample, not just how many.

And the things that feel productive but move the wrong number.

  • More reminders to the same willing list. It lifts the response rate and leaves the composition of the sample unchanged, which is the definition of improving the metric that does not matter.
  • Treating the response rate as the goal. There is no proof that raising it automatically reduces bias, and the National Academies review notes that extraordinary efforts to secure responses from a reluctant population may even increase bias on some estimates.
  • A bigger sample. It narrows the interval around a wrong number. A biased poll of 50,000 is a more confident wrong answer than a biased poll of 500, not a better one.
  • Quoting a response rate as though it settles the matter. Report it, then report what you did to check the composition of your sample.

And one that is easy to miss: reduce the reason for the silence, not just the friction. Anonymity people actually believe in does more for an employee poll than any incentive, because the correlation you are fighting is between being unhappy and fearing identification. Remove it and you have shrunk the gap term, which is the only term that ever mattered. Mechanics in how to make a poll.

What weighting can and cannot fix

Weighting takes the people who did respond and stretches them to match a population you already know. It works, inside a boundary that two US election cycles marked out with unusual clarity.

2016 shows weighting working. AAPOR's evaluation found that adjusting for over-representation of college graduates was critical, but that many polls did not do it. Education correlated strongly with vote choice in the key states, people with more formal education are significantly more likely to answer surveys, and many state polls never corrected for it. The result was an over-estimate of support for Clinton: a real bias from nonresponse, and weighting was the fix.

2020 shows weighting running out of road. By then the correction was near universal: of the 317 state-level presidential polls in the final two weeks that disclosed their adjustments, 92 percent accounted for education. The polls were still wrong, and by more. AAPOR found the average signed error too favorable for Biden by 3.9 points nationally and 4.3 points in statewide presidential polls, the largest national error in 40 years.

The task force then ruled the demographic explanation out. Polling error was not primarily caused by incorrect assumptions about the composition of the electorate in terms of age, race, ethnicity, gender or education level, and reweighting the data to match the actual outcome produced only minor changes to the demographic weights.

That is the boundary. Weighting corrects a difference in who is in your sample. It cannot correct a difference inside a weighting cell. If the Republicans who answered polls differed from the Republicans who did not, matching the party split to a target repairs nothing, because both groups sit in the same cell and get the same weight. AAPOR named that as the leading remaining hypothesis, then added something more useful than a conclusion: reliable information is lacking on the demographics, opinions and vote choice of those not included in polls.

Weighting is not free, and it can improve one estimate while damaging another.

Pew's phone polls overstated working with neighbors by 38 points. Weighting the volunteers down would have corrected the civic numbers, but in 2016 the same adjustment would likely have made the political estimates worse, because the over-represented volunteers leaned Republican. One scheme, two questions, opposite effects.

Three rules follow. Weight on variables where you know the population truth, not variables you wish you knew. Check what a new scheme does to your other estimates, not just the one you built it for. And never let weighting persuade you a hole has been filled: it has been covered over with the opinions of the people who did turn up.

Informal polls: who is missing from a Twitter, WhatsApp or website poll

Nobody weights a Twitter poll. There is no frame, no denominator, and no record of who saw it and scrolled past. Everything above still applies, so it pays to know which people each format removes for you.

Two filters run before anyone votes on a poll posted to X or Twitter. The first is the platform. Pew found 22 percent of US adults used it, with a median age of 40 against 47 for all US adults, 42 percent holding at least a bachelor's degree against 31 percent of the public, and 60 percent identifying as or leaning Democratic against 52 percent of US adults.

The second filter is sharper. The median Twitter user posted twice a month, while the most prolific 10 percent produced 80 percent of all tweets from US adults, and 69 percent of that heavy group had tweeted about politics against 39 percent of users generally. A social poll is answered by the active fraction of an already unrepresentative population, and that fraction is its most opinionated part.

Everyone whose view you care aboutThe population you would like the answer to describe.
People on that platform at allRemoves anyone not using the app, which already skews by age and country.
People who follow the accountRemoves everyone who has not already opted into your point of view.
People the feed actually showed it toThe ranking algorithm decides this, and it favours recent and engaged users.
People who bother to voteA small, unusually motivated minority. This is your sample.

A WhatsApp group poll fails differently. The group is small enough to name every member, which makes the silence feel meaningless. It is not. The non-voters are the ones who muted the group, the ones in another timezone, and the ones who can already see their preference losing. Ask twelve people whether to move the standing meeting to 7am and those who would suffer most are precisely those least likely to be awake and reading. A scheduling poll or meeting availability poll handles that better, because it takes a full grid from each person and the empty rows are visible.

A website poll is the most reliably biased of the three, and it leans the same way every time. Only people who reached the page, stayed long enough to scroll, and cared enough to click can answer. Ask whether the page was helpful and the visitor who left after five seconds because it was useless is structurally prevented from saying so. That is a satisfaction score among people already satisfied enough to stay. See website feedback and feedback question examples.

Opt-in online panels add a problem the others do not have: some respondents are not answering sincerely. Pew replicated an opt-in poll reporting that 20 percent of US adults under 30 agreed the Holocaust is a myth, and using a probability-based panel found 3 percent. The damage concentrates in exactly the small subgroups people most want to quote.

  • Do not quote an informal poll as a population percentage. Say what people said, not what share of a country believes it.
  • State the audience. My followers, 340 votes is honest. Americans think is not.
  • For anything with a decision attached, collect it somewhere you control rather than in a feed. A form builder gives you the respondent list that makes every check above possible.

None of this makes informal polls worthless. A poll among people who care about a topic tells you what people who care about that topic think, which is often the question you actually had. The failure is only ever in the generalization. And if the poll exists to entertain rather than to decide, fun poll questions are there because not every poll needs to survive scrutiny.

Common questions

Does a low response rate mean my poll results are wrong?

No. It means the potential for error is larger, not that error is present. Nonresponse bias is the nonresponse rate multiplied by how differently the non-repliers would have answered, so if that second term is near zero a low response rate produces almost no bias.

The strongest evidence is a meta-analysis by Groves and Peytcheva covering 59 studies that each measured nonresponse bias directly. It found very little correlation between the nonresponse rate and the measured bias.

Is there a minimum acceptable response rate, like 50 or 70 percent?

No. There is no threshold in the survey methodology literature below which results become invalid, and a confident claim of one should make you suspicious of the source rather than of the poll.

Pew telephone polls ran at 9 percent, then 7 percent, then 6 percent, and stayed within about 3 percentage points of high response rate government benchmarks on demographic measures. The same polls were 38 points out on a question about helping neighbors. The rate predicted neither outcome. The subject of the question did.

What is the difference between nonresponse bias and sampling error?

Sampling error is random. It comes from measuring a sample rather than everyone, it is what a margin of error describes, and it shrinks predictably as the sample grows.

Nonresponse bias is systematic. It does not shrink with sample size and the margin of error does not cover it. A poll can report plus or minus 2 points and still be 26 points wrong, because those two numbers measure different things.

Will a bigger sample fix nonresponse bias?

No, and this is the most expensive misunderstanding in the field. Doubling the sample halves nothing about bias, it only tightens the interval around the biased estimate.

If your poll systematically misses unhappy customers, polling ten times as many people means missing ten times as many unhappy customers.

Does weighting remove nonresponse bias?

It removes the part you can see. If you undersampled a group whose views differ, and you know the true size of that group, weighting corrects the imbalance. AAPOR found this was decisive in 2016, when many state polls failed to adjust for the over-representation of college graduates and overstated support for Clinton.

It cannot remove a difference that sits inside a weighting cell. In 2020, 92 percent of the state presidential polls that disclosed their adjustments did weight on education, and the polls were still 4.3 points too favorable for Biden on the state-level margin. AAPOR ruled out demographic composition as the primary cause.

How do I estimate bias if I know nothing about the people who did not reply?

You almost always know more than you think. If you sent the poll to a list, you hold the list, so you can compare the profile of your responses against the profile of everyone invited on tenure, region, team, plan or purchase history.

If you genuinely have nothing, run a wave analysis: sort responses by arrival order and watch whether the headline number drifts as later, more reluctant responses come in. Then chase a small random sample of non-repliers directly.

Do incentives reduce nonresponse bias?

Incentives reliably raise response rates. A meta-analysis of 21 years of household survey experiments found a strong, non-linear effect, with prepaid incentives in mail surveys producing the largest response per dollar.

Whether they reduce bias is a separate question, and it depends entirely on who the money brings in. An incentive that recruits more of the people already answering raises the rate and leaves bias untouched. The National Academies review is blunt that there is no proof raising response rates automatically reduces bias, and that heavy pressure on reluctant respondents can make data quality worse.

Was the 2016 US polling miss caused by rural voters refusing to answer?

Mostly no, and this is a common misreading. AAPOR tested that theory directly by comparing where respondents lived against Census benchmarks, and reported that if anything, people living in the most pro-Trump parts of the country were slightly over-represented in the polls.

The explanations with the most evidence were real late movement in vote preference in the final week, and the failure of many state polls to weight for the over-representation of college graduates. Differential nonresponse was not ruled out, but the simple geographic version of the story was not supported by the data.

What happened in 2020, then?

2020 was worse and less explicable. The error was the highest in 40 years for the national popular vote, with polls in the final two weeks too favorable for Biden by 3.9 points nationally and 4.3 points in statewide presidential races.

AAPOR ruled out late deciders, education weighting, demographic composition, respondents hiding their support, and turnout misestimation. The leading remaining hypothesis is differential nonresponse within party, meaning the Republicans who answered polls differed from the Republicans who did not. The task force could not confirm it, because confirming it would require data on the people who never took part.

Is item nonresponse the same problem?

It is the same mechanism at a smaller scale. Whole-poll nonresponse is a person skipping the entire survey. Item nonresponse is a person skipping one question inside it. Both bias the estimate whenever the skipping correlates with the answer.

Item nonresponse is easier to spot, because you can see which questions bleed respondents. A salary question that 30 percent of people leave blank is not giving you a salary distribution for your population, it is giving you one for the people comfortable disclosing.

Are Twitter and WhatsApp polls useless?

Not useless, just narrow. They tell you what an engaged, self-selected, platform-specific slice of people thinks, and sometimes that slice is genuinely the group you cared about.

They stop being useful the moment the result is generalized. Pew found US adult Twitter users skew younger, more educated, higher income and more Democratic than the public, and that the most active 10 percent produce 80 percent of tweets while being far more politically engaged than the rest.

Sources

  1. Groves, R. M., and Peytcheva, E. (2008). The Impact of Nonresponse Rates on Nonresponse Bias: A Meta-Analysis. Public Opinion Quarterly, 72(2), 167-189.
  2. National Research Council (2013). Nonresponse in Social Science Surveys: A Research Agenda, chapter on Nonresponse Bias. National Academies Press.
  3. Pew Research Center (2017). What Low Response Rates Mean for Telephone Surveys.
  4. Kennedy, C., and Hartig, H. (2019). Response rates in telephone surveys have resumed their decline. Pew Research Center.
  5. Clinton, J., and others (2021). Task Force on 2020 Pre-Election Polling: An Evaluation of the 2020 General Election Polls. AAPOR.
  6. Kennedy, C., and others (2017). An Evaluation of 2016 Election Polls in the United States. AAPOR.
  7. Mercer, A., Caporaso, A., Cantor, D., and Townsend, R. (2015). How Much Gets You How Much? Monetary Incentives and Response Rates in Household Surveys. Public Opinion Quarterly, 79(1), 105-129.
  8. Wojcik, S., and Hughes, A. (2019). Sizing Up Twitter Users. Pew Research Center.
  9. Mercer, A., Kennedy, C., and Keeter, S. (2024). Online opt-in polls can produce misleading results, especially for young people and Hispanic adults. Pew Research Center.

Related guides

Ready to ask?

One question, a handful of answers, a link you can paste anywhere. No account, no limits on votes.

Make a free poll