“Explain the probability sampling strategies with examples.” (2019)
- Jerzy Neyman’s formalisation of stratified random sampling in the 1930s supplied the statistical logic that still underwrites every probability sampling strategy used in Indian social research today: a sample is only as trustworthy as the known, calculable chance each unit had of being drawn into it.
- Probability sampling is any sampling strategy in which every unit of the population has a known, non-zero chance of selection — this single condition is what licenses statistical inference and lets a researcher attach a calculable margin of error to an estimate.
- It presupposes a workable sampling frame — a list, however constructed, from which units can actually be drawn at random — and its inferential power collapses if that frame is incomplete or outdated, no matter how sophisticated the drawing procedure applied to it.
- Five strategies are conventionally distinguished, and each solves a different practical constraint (cost, frame availability, need to represent small groups) rather than being interchangeable options — the value of each is best seen through a worked scenario rather than an abstract definition.
Simple Random Sampling
- Every possible combination of n units from the population has an equal chance of being selected, typically executed through a lottery method or a random-number table applied to a numbered list of the full population.
- Worked example: a public-health researcher wants to estimate the prevalence of hypertension among patients registered at a cluster of primary health centres in a district. With a complete, up-to-date patient master list as the frame, the researcher assigns each patient a number and draws the required sample using a random-number generator — every registered patient, regardless of age, sex or visit frequency, has an identical chance of being pulled into the study.
- Its strength is that it is the unbiased benchmark against which every other design is judged; its weakness is that it demands a complete frame, which is precisely what is often unavailable for large, dispersed, or poorly enumerated populations — and a purely random draw can still scatter the sample so widely across a district that fieldwork becomes needlessly expensive.
Systematic Sampling
- After a random starting point within the first interval, every kth unit thereafter is selected, where k is the population size divided by the desired sample size.
- Worked example: to study dropout risk factors, an education researcher takes a school’s admission register — several thousand students listed by enrolment number — picks a random start between 1 and 20, and then selects every 20th student thereafter, quickly generating a sample spread evenly across the entire register without having to generate hundreds of separate random numbers.
- The strategy’s real danger is periodicity: if the sampling frame itself has a hidden cyclical structure that happens to coincide with the chosen interval — for instance, a household listing for a housing-access survey organised floor-by-floor within identical apartment blocks, where every kth unit always lands on the same corner unit on each floor — systematic sampling silently reproduces that pattern instead of neutralising it, yielding a systematically skewed rather than a representative sample.
Stratified Sampling
- The population is first divided into strata that are internally homogeneous on some criterion judged relevant to the study — caste, religion, income category, region — and a random sample is then drawn independently within each stratum.
- Proportionate stratification mirrors each stratum’s actual share of the population in the sample; disproportionate stratification deliberately oversamples a numerically small stratum so that it yields enough cases for reliable analysis, with results reweighted afterward to restore proportional accuracy.
- Worked example: an employment researcher studying labour-force participation across a state stratifies the target population first by social category (Scheduled Caste, Scheduled Tribe, Other Backward Class, general category) and then by rural–urban residence, before drawing a random sample within each cell. A purely simple random draw across the whole state could easily return too few respondents from a small Scheduled Tribe population settled in a handful of blocks to say anything reliable about it — stratification, with disproportionate oversampling of that group, guarantees it is analysable rather than left to chance.
- Its cost is that it requires prior, reasonably accurate knowledge of stratum sizes — a stratification scheme built on outdated demographic assumptions can misallocate the sample as badly as no stratification at all.
Cluster Sampling
- Naturally occurring groups — villages, urban wards, schools — are treated as the sampling units; a random sample of clusters is drawn, and then all units within each selected cluster are studied, rather than sampling individuals scattered across the whole population.
- Worked example: to evaluate implementation of a school mid-day-meal programme across a state, a researcher randomly selects a set of government schools as clusters from a complete list of schools, then surveys every enrolled child and teacher within each selected school — avoiding the impossible cost of drawing individual children at random from across thousands of schools statewide.
- Its logic is the deliberate reverse of stratification: while strata should be internally homogeneous, clusters ideally should be internally heterogeneous — a genuine cross-section of the wider population in miniature — because when cluster members closely resemble one another (as neighbouring households often do), the design’s sampling error rises for a given sample size even though its fieldwork cost falls sharply.
Multi-Stage Sampling — the Working Model Behind India’s Large Surveys
- Clusters are sampled within clusters across successive stages — districts, then villages or urban blocks, then households, sometimes then individuals — usually combined with stratification applied at the higher stages, so that the “textbook” strategies above function as building blocks rather than rival, mutually exclusive choices.
- India’s National Sample Survey framework is the standard real-world illustration: strata are formed first by state and by rural/urban residence; villages or urban blocks are selected within each stratum as first-stage units, generally with probability proportional to population size; households are then listed and sampled within each selected first-stage unit as second-stage units.
- The most recent large-scale application of this exact logic is the Household Consumption Expenditure Survey conducted across 2023–24, whose consumption and poverty-related estimates rest on precisely this stratified multi-stage design — a reminder that these are not archival textbook categories but the active machinery behind the government statistics cited in contemporary policy debate.
- Its advantage is that it makes national-scale probability sampling feasible without ever needing a single frame listing every individual in the country; its cost is that sampling error compounds across each stage, and the design and weighting calculations required to correct for this are considerably more complex than for any single-stage strategy above.
- Every probability strategy shares Neyman’s founding requirement — a known, calculable chance of selection — but each answers a distinct practical problem: simple random sampling is the unbiased baseline where a complete frame exists; systematic sampling trades a little randomness for field convenience, at the risk of periodicity; stratified sampling protects small but analytically important groups from being missed by chance; cluster sampling makes geographically dispersed populations affordable to study; and multi-stage sampling is what actually makes probability sampling possible at national scale.
- The National Sample Survey framework demonstrates that these are not competing choices in practice — a single well-designed national survey routinely combines stratification, clustering, and multiple stages within one sampling plan.
- What is at stake in getting the strategy right is not a statistical nicety but the very legitimacy of extrapolating from a sample of a few hundred thousand households to conclusions about a population of over a billion people.
- No amount of sophistication at the drawing stage compensates for a defective sampling frame — the strategy determines how units are drawn, but the frame determines who could ever have been drawn at all.
