“What is sampling in the context of social research? Discuss different forms of sampling with their relative advantages and disadvantages.” (2025)
- Sampling is the process of selecting a defined subset of units — individuals, households, villages — from a larger population, in a manner that allows the researcher to draw conclusions about that population without studying every single unit.
- It exists because populations are large and research budgets, and time, are not — the entire justification for sampling is the claim that a properly drawn few thousand cases can stand in for a population of hundreds of millions.
- A.L. Bowley is credited with pioneering the systematic use of sampling in social research in the early twentieth century, and the technique has since split into two fundamentally different logics: probability sampling, licensed by randomisation and capable of statistical inference, and non-probability sampling, which trades statistical representativeness for feasibility, cost, or theoretical purpose.
- The starkest cautionary tale for why the type of sampling matters is the 1948 US presidential election, where major opinion polls relying on quota sampling confidently predicted a Dewey victory over Truman — Truman won, and the failure has stood ever since as the textbook illustration of a non-probability method’s structural risk.
- This answer sets out the sampling frame concept, works through the probability and non-probability families in turn, and closes with a worked comparison of their relative advantages and disadvantages.
The Sampling Frame and the Logic of Probability Sampling
- A sampling frame is the actual list of units from which a sample is physically drawn — a voter roll, a village census list, a school enrolment register — and its defects are inherited by every conclusion drawn from it; no amount of clever randomisation afterward can repair a frame that excludes part of the population from the outset.
- Probability sampling is defined by the randomisation principle: every unit in the population has a known, non-zero, and ideally equal chance of selection, which is what licenses formal statistical inference from sample back to population.
- Simple random sampling: every unit has an equal chance of selection, typically via a lottery or random-number procedure — the unbiased benchmark against which every other design is judged, but it requires a complete frame and can scatter selected units across a wide, costly geography.
- Systematic sampling: every kth unit is selected after a random start on an ordered list — easy to execute in the field, but vulnerable to periodicity in the list (if every kth household happens to be a corner house, or every kth name falls in a particular caste block, the sample is systematically skewed).
- Stratified sampling: the population is first divided into strata that are internally homogeneous on some relevant criterion — caste, religion, income bracket, region — and units are then drawn randomly within each stratum, either proportionately or with deliberate oversampling of small strata (disproportionate stratification) followed by reweighting.
- Its precise logic must not be confused with cluster sampling: strata should be as internally similar as possible, because the entire gain is a reduction in sampling error relative to simple random sampling, and it guarantees that small but relevant groups are not missed by chance.
- Cluster sampling: naturally occurring groups — villages, wards, schools — are sampled first, and then all (or most) units within each selected cluster are studied.
- Its logic runs in the opposite direction from stratification: clusters should ideally be internally heterogeneous, each a rough miniature of the whole population, because only some clusters are ever sampled and the aim is that any given cluster could stand in for the others.
- This is the single most commonly confused distinction in sampling theory: stratify on homogeneity and sample within every stratum; cluster on heterogeneity and sample only some whole clusters.
- Multi-stage sampling: sampling occurs in successive stages — districts, then villages within districts, then households within villages, then individuals within households — usually combining stratification at the higher stages with clustering at the lower ones.
- India’s National Sample Survey, and its successor rounds such as the Periodic Labour Force Survey, are real working examples of this design: a nationally representative sample is built up through stratified selection of first-stage units (villages/urban blocks) followed by systematic or random selection of households within them, precisely because no single frame of every Indian household exists or could be maintained.
Non-Probability Sampling and Its Legitimate Uses
- Non-probability sampling does not give every unit a known chance of selection, does not permit formal statistical generalisation to the population, and relies instead on the researcher’s judgment, availability, or a respondent network.
- Convenience (accidental) sampling: whoever is easiest to reach — a mall intercept, a researcher’s own students — fast and cheap, but the bias is unmeasurable and uncorrectable.
- Purposive (judgmental) sampling: the researcher deliberately selects units judged to be information-rich or typical — entirely legitimate when the goal is theoretical insight rather than statistical estimation, but the sample’s quality is wholly dependent on the soundness of the researcher’s own judgment.
- Quota sampling: the population is divided into categories (age, sex, class) and interviewers are assigned fixed quotas to fill within each — cheap, fast, and able to mirror known population proportions on paper, but the fatal flaw is that within each quota, the interviewer chooses who to approach at their own discretion, which is exactly the crack through which uncorrected bias enters.
- The 1948 Dewey/Truman failure is the sharpest illustration available: pollsters filled their quotas correctly by broad demographic category, but interviewers, left free to choose whom to approach within each quota, systematically over-reached more accessible, generally more affluent Republican-leaning respondents — producing a confident, and completely wrong, prediction.
- Snowball sampling: existing respondents recruit further respondents from their own networks — often the only realistic route into hidden or stigmatised populations (undocumented migrants, members of a banned sect, drug users), but the sample follows social ties, so isolated individuals with no network connection remain structurally invisible.
- Theoretical sampling: the grounded-theory procedure associated with Glaser and Strauss, in which the next case is chosen specifically because the emerging theory needs it, and data collection stops at theoretical saturation rather than at a pre-fixed sample size — no statistical generalisation is claimed or intended; the generalisation is to theory, not to a population.
Fig: Sampling methods — selection logic, advantage, and disadvantage
| Type | How units are selected | Chief advantage | Chief limitation |
|---|---|---|---|
| Simple random | Equal chance for every unit | Unbiased benchmark; full statistical inference | Needs a complete frame; scatters the sample |
| Systematic | Every kth unit after a random start | Easy to execute in the field | Periodicity in the list produces skew |
| Stratified | Random draw within homogeneous strata | Lower sampling error; guarantees representation of small groups | Needs prior knowledge of strata and their sizes |
| Cluster | Whole natural groups sampled, then all units within them studied | Cheap for dispersed populations; no frame of individuals needed | Higher sampling error, since cluster members resemble each other |
| Multi-stage | Clusters within clusters, sampled in stages | Feasible for national populations with no individual frame | Errors compound at each stage; complex to design and weight |
| Convenience | Whoever is available | Fast and cheap | Unknown, uncorrectable bias; no inference possible |
| Purposive | Researcher’s judgment of who is informative | Efficient access to key informants and rare expertise | Wholly dependent on researcher’s judgment |
| Quota | Fixed quotas by category, interviewer picks within quota | Cheap; matches known population proportions | Selection within quota is at interviewer discretion (1948 Dewey/Truman failure) |
| Snowball | Respondents recruit further respondents | Only realistic access to hidden/stigmatised populations | Follows networks; social isolates invisible |
| Theoretical | Next case chosen by what emerging theory needs | Builds theory systematically; principled stopping rule | No statistical generalisation possible |
When Each Family Is the Right Choice
- Probability sampling is the correct choice whenever the research goal is statistical estimation — a prevalence rate, an average income, a population proportion — because only randomisation licenses the calculation of a sampling error (the calculable, unavoidable discrepancy between sample and population).
- Non-probability sampling is not a lesser fallback but the correct choice under specific conditions: when no sampling frame exists and cannot practically be built (the normal condition for hidden populations), when the population is genuinely rare, when the research is exploratory and the goal is to discover categories rather than measure their spread, and whenever — as in most qualitative research — the logic of generalisation is theoretical rather than statistical. Applying probability logic to a grounded-theory study, or dismissing quota sampling for market research where only a rough estimate is needed, are both category errors in the opposite direction.
- A distinction worth holding onto throughout: sampling error shrinks as sample size grows and is calculable only under probability sampling, while non-sampling error — a bad frame, non-response, interviewer bias, coding mistakes — does not shrink with size at all, meaning a very large but badly designed sample can be confidently wrong.
- The continuing relevance of the 1948 failure is not that quota sampling is inherently useless — it remains a serviceable tool for rough, low-stakes estimation — but that interviewer discretion inside a quota is where representativeness quietly dies, a lesson opinion pollsters have had to relearn repeatedly, including in more recent electoral forecasting exercises where exit-poll and pre-poll projections have diverged sharply from actual results, renewing scrutiny of how genuinely representative even carefully quota-managed samples really are.
- Sampling method is never a merely technical choice separable from the research question: a survey question demands a probability design capable of estimating a population parameter, while a question about the meaning of an experience among a rare or hidden population demands the reach that only non-probability sampling can provide.
- The stratified-versus-cluster distinction remains the conceptual fault line separating a sophisticated sampling answer from a rote one — homogeneity within strata, heterogeneity within clusters — and getting this exactly backward is the single most common error a researcher, or a candidate, can make.
