Variables, Sampling, Hypothesis, Reliability, and Validity

Variables

  • A variable is any characteristic that takes on two or more values across the units being studied — age, income, caste category, occupation, attitude toward a policy.
    • If a characteristic does not vary at all, it is a constant, and a constant can explain nothing — causal analysis only ever works by explaining one difference through another difference.
    • Earl Babbie frames a variable more formally as a logical set of attributes; individual cases simply differ in which attribute from that set they happen to possess.
  • Independent and dependent variables are the foundational pair.
    • The independent variable is the presumed cause — manipulated by the researcher in an experiment, or treated as logically prior in a survey.
    • The dependent variable is the presumed effect, whose variation the study is trying to explain.
    • The mnemonic that prevents the most common mix-up: the dependent variable depends on the independent one. Example: a handwashing awareness campaign is the independent variable; the resulting change in handwashing behaviour is the dependent variable.
  • Intervening variables sit between cause and effect as the actual mechanism through which the effect is transmitted — caused by the independent variable, and in turn causing the dependent one.
    • Durkheim’s study of suicide supplies the cleanest illustration: religion is the independent variable, the suicide rate is the dependent variable, and social integration is the intervening variable that does the causal work.
    • Protestantism itself does not cause anyone’s death; it produces comparatively lower social integration, and it is that lower integration which raises the suicide rate.
    • Naming the intervening variable is precisely what converts a bare correlation into a genuine explanation.

Further Variable Distinctions

  • An antecedent variable sits earlier still in the causal chain, prior to the independent variable, and can explain why the independent variable takes the value it does.
  • An extraneous or control variable is not part of the hypothesis being tested but can contaminate the result if left unmanaged, so it must be held constant, matched across groups, or controlled for statistically.
  • A moderator variable changes the strength or direction of the relationship between independent and dependent variables, so an effect holds for one group but not another.
  • Experimental variables describe the researcher’s own deliberate manipulations, while measured variables describe what is simply observed and recorded without manipulation (rural development, for instance, is typically measured through indicators like income growth and literacy).
  • Any manipulated variable is an active variable; one that cannot be manipulated at all — gender, caste at birth — is an assigned variable.
  • Variables are also classed as qualitative (discrete, non-numeric, like religion) versus quantitative (numeric, like age or income), and as dichotomous (only two values, like sex) versus continuous (a full range of values, like intelligence scores).
  • A spurious relationship is the trap the entire apparatus of variable analysis exists to guard against.
    • Two variables correlate not because either causes the other, but because both are independently caused by a third variable.
    • Standard illustration: ice-cream sales correlate with drowning deaths — not because ice cream causes drowning, but because summer weather independently drives up both.
    • In sociology this danger is far less obvious: a correlation between social class and crime rates could mean class drives criminality, or that a criminal record blocks upward mobility, or that neighbourhood policing intensity drives both. A bare correlation is never, by itself, an explanation.

Lazarsfeld and Kendall’s Elaboration Paradigm

  • Paul Lazarsfeld and Patricia Kendall’s elaboration paradigm supplies the systematic procedure for telling a real relationship apart from a spurious one: introduce a test factor into an already-observed relationship between two variables, and see what happens to the original association.
  • Replication — the association persists within every category of the test factor, corroborating the original finding.
  • Explanation — the association disappears, and the test factor is antecedent to the independent variable: the original relationship was spurious.
  • Interpretation — the association disappears, but the test factor is intervening rather than antecedent: the relationship is real, and its mechanism has just been found.
  • Specification — the association holds within one category of the test factor but not another, revealing the conditions under which it operates.
  • The elegant trap: explanation and interpretation look statistically identical, since the association vanishes in both. Only the test factor’s position in time separates a debunked spurious correlation from a genuine mechanism — a judgement supplied by the researcher’s theory, not by the statistics alone.
  • Operationalization — converting an abstract concept like “religiosity” or “social class” into an observable, measurable indicator — is where a sociological variable is ultimately won or lost.
    • The resulting number is only as good as the correspondence between the original concept and the indicator standing in for it.
    • This is exactly where Blumer’s and Cicourel’s critiques of variable analysis bite hardest (developed fully in the articles on qualitative and quantitative methods, and on techniques of data collection): sociology has no generic variables the way physics has mass, so a variable’s meaning can shift across contexts even while its label stays the same.

Sampling

  • Sampling exists because populations of interest are almost always too large, too dispersed, or too costly to study in their entirety. Its entire intellectual justification rests on randomness — nothing else licenses the claim that a properly drawn sample of a few thousand can speak reliably for a population of hundreds of millions.
  • A. L. Bowley is generally credited with the first systematic use of sampling in social research, in 1754.

Core Vocabulary

  • The population or universe is the entire set of units a study wants to draw conclusions about.
  • The sampling frame is the actual list the sample is physically drawn from, and it is never identical to the true population — its defects (an outdated voter roll, a phone directory excluding landline-less households) are silently inherited by every number the study later produces.
  • The sampling unit is the individual element being selected, and the sample size is simply how many units are drawn.
  • Sampling error is the chance discrepancy between a sample and the population it was drawn from, arising purely from the luck of the draw — calculable only under probability sampling, and it shrinks as sample size grows.
  • Non-sampling error is everything else that can go wrong — a flawed sampling frame, non-response, interviewer bias, coding mistakes. It does not shrink as a sample gets larger, so a huge, badly-drawn sample is not more trustworthy than a small one — it is simply confidently wrong at a larger scale.

Probability Versus Non-Probability Sampling

  • Probability sampling requires that every unit have a known, non-zero chance of being selected, decided by chance rather than by the researcher’s own judgement — this is precisely what licenses formal statistical inference and a calculable margin of error.
  • Black and Champion specify the practical requirements: a complete list of the units to be studied, a known size for the universe, a specified desired sample size, and an equal chance of selection for every element.
  • Non-probability sampling makes no such guarantee: the chance of any given unit’s selection is unknown, so no formal inference to the wider population is warranted, and no error margin can be meaningfully computed.

Types of Probability Sampling

  • Simple random sampling gives every unit an equal chance of selection, typically via lottery or a random-number table. It is the theoretical benchmark, but often impractical, since it requires a complete frame of what may be an enormous, scattered population.
  • Systematic sampling selects every kth unit after a random start, where k is population size divided by sample size. Simple to execute, but vulnerable to periodicity — if the frame has a cyclical pattern coinciding with k (every tenth house being a wealthier corner plot, say), the sample is skewed without the researcher noticing.
  • Stratified sampling divides the population into internally homogeneous strata (caste, religion, income, region) and samples randomly within each. Proportionate stratification mirrors each stratum’s actual population share; disproportionate stratification oversamples a small stratum so it can be analysed, then reweights. Its chief virtue is guaranteeing representation of small groups a purely random draw might miss — a serious concern in Indian social research.
  • Cluster sampling works the other way round: it samples naturally occurring groups (villages, wards, schools) and studies every unit within the selected clusters. Strata should be internally homogeneous; clusters should ideally be internally heterogeneous, each a rough miniature of the population. Cheaper for dispersed populations, but carries higher sampling error, since people within one village resemble each other more than a random draw would.
  • Multi-stage sampling nests clusters within clusters — districts, then villages, then households, then individuals — usually combined with stratification at higher stages. This is what national surveys actually do: the National Sample Survey’s design is stratified and multi-stage, a reminder that these types are building blocks, not rival choices.

Types of Non-Probability Sampling

  • Convenience (or accidental) sampling uses whoever happens to be available — cheap and fast, but close to worthless for formal inference.
  • Purposive (or judgement) sampling has the researcher deliberately select especially informative units — legitimate when the goal is theoretical depth rather than representativeness. Herbert Blumer recommended seeking out the most acute, best-informed observers of a group rather than a random cross-section.
  • Quota sampling fixes quotas for known categories but leaves selection within each quota to the interviewer’s discretion — exactly where it fails, since interviewers approach whoever is easiest to approach. Its most cited failure: the 1948 US election, where quota-sampling polls confidently predicted a Dewey victory, and Truman won.
  • Snowball sampling has existing respondents recruit further respondents from their own networks — the only realistic way to reach hidden populations (sex workers, undocumented migrants), though the sample follows existing networks, so anyone socially isolated from them remains invisible.
  • Volunteer sampling invites people to opt in. Those who volunteer are typically unusually engaged with the topic, which can skew the sample away from the wider population.
  • Theoretical sampling, developed for grounded theory, has the next case selected by what the emerging theory needs, ending at theoretical saturation rather than a pre-fixed size — no claim to statistical generalization is made here.
  • Deliberately non-representative sampling can itself be a positive choice: in the spirit of Popper’s falsificationism, a researcher may seek out untypical cases specifically to test whether a theory survives an apparent exception. The finding that societies like the Mbuti of the Congo show no strict biological basis for a rigid sexual division of labour is exactly this kind of falsifying case.

Why Randomness Matters, and Where Its Power Stops

  • It removes the researcher’s own discretion from selection, which is where systematic bias most often creeps in, and ensures every unit — including ones a researcher would never think to include — has a known chance of selection.
  • It makes a sample’s departures from the population random rather than directional, so they tend to cancel rather than accumulate, and it permits calculating sampling error and confidence intervals, letting a study honestly state how wrong its estimate might be.
  • It guarantees representativeness only in expectation, over the long run — any single random sample can still be unrepresentative by chance.
  • It cannot repair a defective sampling frame — a random draw from a bad list is simply a random bad sample — and it does nothing about non-response, which is where most real-world surveys actually fail.
  • Random sampling secures external validity (generalizability), not measurement validity — a random sample answering a badly worded question yields perfectly representative nonsense.

Hypothesis

  • A hypothesis is a tentative, testable statement asserting a relationship between two or more variables. Because it must be empirically testable, it excludes pure opinions, value judgements, and normative claims about what ought to be — none of these can, even in principle, be shown right or wrong by evidence.

How Different Scholars Define a Hypothesis

  • Theodorson and Theodorson: a tentative statement asserting a relationship between certain facts.
  • Kerlinger: a conjectural statement of the relationship between two or more variables.
  • Black and Champion: a tentative statement about something whose validity is, when proposed, unknown.
  • William Goode and Paul Hatt: a proposition that can be put to a test to determine its validity — worth holding onto for its emphasis on testability.

Types of Hypothesis

  • By logical role: the null hypothesis (H0) asserts no relationship or no difference, and it is the hypothesis actually put through statistical testing — a claim can be decisively disproved by contrary evidence, but never conclusively proved by any amount of confirming evidence (Popper’s falsificationism, embedded in statistical practice). The alternative or research hypothesis (H1) is what the researcher actually expects, accepted only by implication when H0 is rejected; a study never proves H1 directly, it only fails to sustain H0.
  • By content:
    • Descriptive hypotheses assert the existence, size, or distribution of a phenomenon.
    • Relational hypotheses assert that two variables co-vary without claiming causation.
    • Causal hypotheses assert that one variable produces change in another.
  • By precision:
    • A directional hypothesis specifies which way a relationship runs.
    • A non-directional hypothesis asserts only that some relationship exists.
  • By stage of development:
    • A working hypothesis is a preliminary assumption used to shape the research design before it is refined further.
    • A scientific hypothesis is derived from sufficient existing theoretical and empirical grounding.
    • A statistical hypothesis reduces the claim to numerical, testable quantities.

Type I and Type II Errors

  • A Type I error rejects a null hypothesis that is actually true — a false positive.
  • A Type II error fails to reject a null hypothesis that is actually false — a false negative.
  • The two trade off: a stricter significance threshold reduces Type I errors but raises Type II errors, and no setting minimizes both at once.

Sources of Hypotheses

  • Existing theory is the most respectable source, since a hypothesis deduced from theory tests that theory when it is tested.
  • Past research, particularly its unexplained anomalies, regularly generates new hypotheses.
  • Personal observation and experience can suggest a pattern worth formally testing.
  • Folk wisdom and common sense are a legitimate source of hypotheses, even though they are a poor place to draw final conclusions from directly (the connective thread to the article on sociology and common sense).
  • Analogy borrowed from another discipline has historically been productive — Spencer and Durkheim both borrowed the organic analogy from biology.
  • A researcher’s own cultural values and social position shape which hypotheses even occur to them, whether or not this is consciously acknowledged.

Middle-Range Theory as the Source Worth Leading With

  • Robert Merton’s contribution is worth leading with: hypotheses should ideally be derived from theories of the middle range.
  • Middle-range theory is general enough to count as genuine theory, but specific enough to generate testable propositions.
  • It occupies the productive space between Talcott Parsons’s grand, highly abstract theory (which generates essentially no testable hypotheses) and purely descriptive “abstracted empiricism” (which tests propositions connected to no larger theory).
  • Merton’s own reference-group theory and anomie/strain theory are middle-range theories precisely because each yields concrete, testable hypotheses.

Qualities of a Workable Hypothesis

  • Conceptually clear — terms defined precisely enough to be operationalizable, not vague or merely evocative.
  • Genuine empirical referents — variables that can actually be tested, not a moral judgement dressed up as a claim.
  • Specific, not sweepingly general — an overly broad hypothesis cannot be tested as stated.
  • Related to available techniques — a hypothesis no existing method can investigate is merely a wish.
  • Related to a body of theory, so confirming or refuting it carries meaning beyond the single study.
  • This list of qualities comes from Goode and Hatt.

When a Hypothesis Isn’t Needed

  • Exploratory research exists precisely because too little is yet known to state a specific hypothesis in advance.
  • Grounded theory explicitly forbids starting with one: Glaser and Strauss argued a fixed prior hypothesis forces incoming data into categories decided before the setting is understood, making the hypothesis an output rather than an input.
  • Ethnography, phenomenology, and ethnomethodology do not test hypotheses at all, because their object of study is meaning, not a causal relationship between variables.
  • A hypothesis is indispensable to explanatory research and inappropriate to exploratory and interpretive research (this connects back to the article on qualitative and quantitative methods) — which is called for depends on how much is already known.

A Worked Example

  • Take the hypothesis “illiteracy of mothers leads to female infanticide”: female infanticide is the dependent variable, maternal literacy is the independent variable.
  • Testing it means operationalizing both terms precisely, then actively trying to falsify the claim — checking the rate of female infanticide among well-educated mothers.
  • If that rate is comparable to the rate among illiterate mothers, the hypothesis is false; if educated mothers show a markedly lower rate, it is tentatively supported.
  • Even then, the exercise should raise the further question of reverse causation, or a third factor like regional poverty, before the finding is treated as settled.

Reliability

  • Reliability is the consistency, stability, and repeatability of a measurement instrument — whether it gives the same result on repetition under the same conditions.
  • If two investigators using the same instrument, or the same investigator on two occasions, arrive at different results, the instrument is unreliable and the data is essentially noise.
  • Reliability concerns the absence of random error, and it is estimated, not directly measured.

Ways of Testing Reliability

  • Test-retest — the same instrument, administered to the same people after an interval, with the two sets of scores correlated. Vulnerable to memory of the first administration and to genuine change in the interval.
  • Split-half — an instrument’s items are divided into two halves, and the two halves’ scores are correlated, testing internal consistency in one sitting.
  • Internal consistency, generalized across every possible split via Cronbach’s alpha.
  • Parallel or alternate forms — two equivalent versions of an instrument given to the same people, avoiding the memory problem.
  • Inter-observer or inter-coder reliability — independent observers or coders record the same material, and their agreement is computed. Indispensable for observation and latent content coding.

Kirk and Miller’s Typology, for Qualitative Research

  • Quixotic reliability — a single method of observation yields the identical measurement repeatedly. In an ethnographic interview, this suspicious consistency can indicate the researcher has only elicited a rehearsed or “politically correct” answer.
  • Diachronic reliability — the stability of an observation across time, the logic behind test-retest. Harder to achieve for fast-changing socio-cultural phenomena than for stable psychological traits.
  • Synchronic reliability — similarity of observations made within the same time period, evaluated by comparing the same data through different methods. Its paradox: it can be more informative when absent, since divergent results can alert a researcher to a dimension of the problem they had not considered.

Validity

  • Validity is whether an instrument actually measures what it claims to measure — the correspondence between the measurement and the underlying concept. It concerns the absence of systematic error, as distinct from reliability’s concern with random error.

Types of Validity

  • Face validity — the measure simply looks right on inspection. The weakest possible form, effectively no real evidence at all.
  • Content validity — the items genuinely cover the full conceptual domain (a religiosity scale asking only about temple attendance has poor content validity, since it omits belief and private practice).
  • Criterion validity — the measure correlates with an independent external criterion, either concurrently or predictively.
  • Construct validity — the measure behaves the way theory predicts: correlating with what it theoretically should (convergent validity), and not with what it should not (discriminant validity).
  • Internal validity asks whether a causal inference within a study is sound.
  • External validity asks whether a finding generalizes beyond the sample.
  • Ecological validity asks whether a finding holds up in natural settings rather than only the artificial setting it was produced in — a standard reason laboratory experiments, and some argue even questionnaires, can be questioned on ecological grounds.
  • Alan Bryman’s four-part scheme largely overlaps with the above under slightly different labels: measurement validity (also called construct validity — does an IQ test really measure intelligence), internal validity, external validity, and ecological validity.

Reliability and Validity Are Not the Same Thing

  • The relationship between reliability and validity is asymmetric, and this is the single most exam-relevant point in the whole topic: reliability is necessary but not sufficient for validity.
  • A genuinely valid measure must also be reliable, but a perfectly reliable measure can still be consistently and precisely wrong.
  • Example: a weighing scale calibrated five kilograms too high gives the identical answer every time — perfectly reliable, and entirely invalid.
  • Reliability is the absence of random error; validity is the absence of systematic error — that single distinction resolves most confusion here.
ReliabilityValidity
Question it answersDoes the instrument give the same result on repetition?Is the instrument measuring what it claims to measure?
Type of error addressedRandom errorSystematic error
How it is assessedTest-retest, split-half, Cronbach’s alpha, parallel forms, inter-coder agreementFace, content, criterion, construct, internal, external, ecological
Logical relationNecessary but not sufficient for validityPresupposes reliability — a valid measure must also be reliable
Favoured byStandardisation — fixed wording, fixed codes, no researcher discretionImmersion — flexibility, probing, natural settings, actors’ own categories
What it costsValidity — imposes the researcher’s own categories (Cicourel’s “measurement by fiat”)Reliability — the researcher is themselves the instrument, so the study cannot be repeated by anyone else

The Reliability-Validity Trade-Off

  • Standardisation buys reliability at the cost of validity, and immersion buys validity at the cost of reliability.
  • A highly standardised instrument is reliable precisely because it removes discretion — but that same rigidity imposes the researcher’s own categories onto respondents whose own categories may differ, and forbids probing (Blumer’s critique of the variable and Cicourel’s “measurement by fiat,” developed fully in the article on qualitative and quantitative methods).
  • Immersive qualitative work achieves the opposite trade: it stays close to actors’ own meanings, producing high validity — but the “instrument” is a particular researcher with a particular biography, so no second researcher could exactly reproduce the study, and reliability becomes effectively unattainable.
  • This follows from the object being studied, not from sociologists’ technical incompetence: meaning is context-dependent, and any instrument standardised enough to stay identical across contexts is, by that fact, insensitive to context.

Triangulation

  • Triangulation is the general strategy for managing this trade-off rather than eliminating it.
  • Norman Denzin identifies four types: data triangulation (multiple sources, time points, or settings), investigator triangulation (multiple researchers, which also supplies a form of inter-observer reliability), theory triangulation (multiple theoretical lenses on the same data), and methodological triangulation (multiple distinct methods on the same question).
  • Harvey and MacDonald identify what methodological triangulation can achieve: gathering genuinely different types of information; having researchers independently apply the same method and compare results; and checking that material gathered one way is both reliable and valid by cross-referencing it against another.
  • The underlying logic, from Eugene Webb’s multiple operationism: every method carries its own flaws, but differently-flawed methods carry different flaws, so convergence across them counts as real evidence, and divergence is informative rather than simply a failure.
  • Triangulation is not a machine for manufacturing certainty: when methods genuinely disagree, there is no algorithm for deciding which to trust — and the assumption that different methods measure the same underlying reality is exactly what interpretivists dispute.
  • Respondent validation — having those studied check the findings — is sometimes proposed as a further check, but Rosaline Barbour notes it does not guarantee validity either, since conflicting interpretations between researcher and respondent are common, and choosing which to treat as “correct” is often itself a value-laden choice.
  • Treat triangulation and respondent validation as ways of making a finding more defensible and its limits more visible — not as proof.

Previous Year Questions

  1. What is a variable in social research? What are their different types? Elaborate. (2025) (10 marks)
  2. What do you mean by reliability? Discuss the importance of reliability in social science research. (2025) (10 marks)
  3. What is hypothesis? Critically evaluate the significance of the hypothesis in social research. (2025) (10 marks)
  4. What is sampling in the context of social research? Discuss different forms of sampling with their relative advantages and disadvantages. (2025) (20 marks)
  5. What are variables? How do they facilitate research? (2023)
  6. Explain the different types of non-probability sampling techniques. Bring out the conditions of their usage with appropriate examples. (2022)
  7. What is reliability? Explain the different tests available to social science researchers to establish reliability. (2022)
  8. Discuss the importance and source of hypothesis in social research. (2020)
  9. Explain the probability sampling strategies with examples. (2019)
  10. Illustrate with example the significance of variables in sociological research. (2017)
  11. How can one resolve the issue of reliability and validity in the context of sociological research on inequality? (2017)
  12. “Hypothesis is a statement of the relationship between two or more variables.” Elucidate by giving examples of poverty and illiteracy. (2016)
  13. What are variables? Discuss their role in experimental research. (2015)
  14. Why is random sampling said to have more reliability and validity in research? (2015)
  15. Write short note on Reliability and Validity, keeping sociological perspective in view. (2011)
  16. Distinguish between probability and nonprobability sampling methods. How many types of sampling designs are there? (2009)
  17. Write short note: Importance and sources of hypotheses in social research. (2008)
  18. What is the importance of sampling in sociological studies? Distinguish between simple random sampling and stratified random sampling. (2008)
  19. Utility of Reliability and Validity in Social Research. (2003)
  20. Write short note: Reliability of a sample. (1998)
guest
1 Comment
Oldest
Newest Most Voted
Bhavs🦋

Rich content