Why might I not want to maximize sample size?
- Bias: If your sample or data collection is biased in a way that reduces the validity of your conclusions, adding more people will not necessarily improve the accuracy of your results.
- Diminishing Returns: The gain in information beyond a certain sample size becomes minimal.
- Resources: It might require more resources to gather a larger sample.
- Fatigue: If the individuals being sampled are frequently surveyed, additional surveying adds to "community fatigue", which may lead to increased non-response and reduced accuracy in responses.
Why does adding more people not necessarily reduce bias in the results?
Increasing sample size means adding information. However, if the way you sample individuals occurs in a way that adds biased information to your results, then you are simply adding more biased information to your sample that will reinforce the bias in your results.
Adding more people to a survey sample with a goal of decreasing bias in the results only occurs if one is intentionally sampling to correct any existing bias in the current sample. The goal of such intentional sampling would be to improve the chance that the characteristics of interest in the final sample are representative of the characteristics of interest in the larger population.
Why is "more people" not better when it comes to gaining information?
From a practical perspective, many results will "stabilize" after a certain sample size is reached. In other words, the final result or conclusion will essentially not change after a certain number of people respond. At this point, expending additional resources to sample more people might be unnecessary or add to survey fatigue.
For example, one can expect that the percentage of respondents who chose a particular option on a closed-ended question will tend to stabilize within a couple of percentage points around 700 to 1,000 respondents, and sometimes with even smaller sample sizes.
Of course, there are exceptions. For example, if a complex statistical model will be estimated, then the need for a minimum sample size is increased. Additionally, seeing "stability" in your results may be different for categorical questions (e.g., agreement scale) vs. numerical questions (e.g., age).