“The difference between information and data in social science is subtle. Comment.” (2022)
- Data, in the ordinary methodological sense, are the raw, unprocessed records a researcher generates through systematic collection — survey responses, census counts, coded observation notes, official registers — while information is data that has been organised, interpreted, and given context so that it answers a question.
- The distinction sounds elementary on first statement, which is exactly why the question calls it “subtle”: a number on a page looks like a plain, self-evident fact, when in reality every datum already carries an interpretive history the moment it is collected.
- Aaron Cicourel’s notion of “measurement by fiat” names the core problem directly: a researcher’s predetermined response categories decide, in advance, what can even be recorded as data — so “raw” data are never actually raw, they are pre-interpreted by the very instrument that produced them.
- The argument developed here: the data/information boundary blurs because data are theory-laden from the moment of collection, and treating data as if they were already self-evidently meaningful information — a persistent positivist tendency — is itself a methodological error worth naming explicitly, not a harmless simplification.
The Standard Distinction, Stated Precisely
- Data are the discrete, unprocessed units generated by an instrument of collection: a tally of responses, a recorded observation, a government register entry, a count.
- Information is what results once data have been organised, classified, cross-tabulated, and interpreted in light of a research question — a table, a rate, a trend line, a finding stated in relation to a hypothesis.
- On this reading, information is simply data plus context and purpose; the movement from one to the other is what an entire research process — coding, classification, analysis — actually does.
Why the Line Is Not as Clean as It Looks
- The apparent self-evidence of a number is precisely where the subtlety lies: a suicide count, a literacy percentage, a crime rate all present themselves as plain facts, when each is the outcome of prior, often invisible, classificatory decisions.
- Cicourel’s measurement by fiat exposes this at the point of collection: when a schedule offers pre-set categories for occupation, caste, or attitude, the respondent’s actual, lived reality is forced into boxes decided before the fieldwork began — the researcher cannot recover, from the resulting “data,” any nuance the categories were not built to capture.
- Because every act of measurement already embeds a theoretical decision about what matters and how it should be classified, no datum is genuinely raw — the positivist habit of treating a number as if it were a transparent window onto reality, requiring no interpretation, mistakes an already-processed artefact for an unprocessed fact.
- This is the methodological error the question is pointing toward: what looks like data (self-evident, requiring no further work) is frequently already functioning as information (an interpreted claim), while what a researcher calls “information” (a published rate, an official statistic) is often simply data from an earlier, hidden round of classification that has not been re-examined.
The Sharpest Illustration: Official Suicide Statistics
- Durkheim’s classic study of suicide treated official suicide statistics as reliable social facts — external, objective indicators of a society’s level of integration and regulation, fit to be correlated with variables like religion and marital status.
- Later critics, most influentially Jack Douglas, challenged this at the root: a “suicide” only enters the official statistic once a coroner, police officer, or medical examiner has classified an ambiguous death as self-inflicted — a classification shaped by the available evidence, the deceased’s social status, family pressure to avoid the stigma of suicide, insurance implications, and the local officials’ own working assumptions about what a “typical” suicide looks like.
- This means the official suicide rate is not a raw social fact waiting to be read off from reality — it is already a social construct, produced through an institutional classification process, before a sociologist ever opens the register to analyse it.
- The implication cuts uncomfortably in both directions: if official statistics are constructed rather than raw, this undermines confidence in a great deal of quantitative sociology that treats such figures as unproblematic data — including, ironically, much of the very critique that draws on other official statistics to make its case.
- Contemporary debates over the reliability of suicide data — where researchers, forensic pathologists, and public-health statisticians continue to flag inconsistent classification practices, under-reporting linked to stigma, and coding differences across jurisdictions as a live problem in interpreting even recent suicide-rate trends — show this is not a dated methodological curiosity but a continuing, practically consequential issue in how such figures are produced and read.
Why the Subtlety Matters Methodologically
- Treating data as if it were already information — the mistake this question is really targeting — has a specific practical consequence: it allows classificatory decisions embedded early in the data-generation process to pass unexamined into a final analysis, disguised as neutral fact.
- Recognising the distinction as genuinely subtle, rather than obvious, obliges a researcher to interrogate the categories and classification procedures behind any dataset before treating its output as settled information — asking not just what the numbers say, but what prior decisions produced the numbers in that form.
- This is also why qualitative and interpretive method retains a distinctive value even for questions that appear quantitative on the surface: understanding how a category like “suicide” or “crime” gets applied in practice is itself indispensable information about the social process generating the data, not a mere footnote to it.
- The corrective is not to abandon quantitative data but to treat every dataset as theory-laden and processed at the point of origin, meaning the analytic work of turning data into trustworthy information begins earlier — at the moment categories are designed — than a naive data/information distinction suggests.
- The difference between data and information is subtle precisely because data present themselves as self-evidently factual while actually being theory-laden from the moment of measurement, a distinction that measurement by fiat makes explicit.
- The suicide-statistics debate remains the sharpest available illustration: what looks like a raw social fact is, on inspection, already a classificatory and institutional construct.
- The methodological lesson is to treat the boundary as porous rather than fixed — information is not simply data plus later analysis, since classificatory “analysis” of a kind has already occurred by the time data exist at all.
- A sociologically literate researcher’s task is therefore to interrogate the categories behind any figure before accepting it as settled information, rather than assuming the movement from data to information is a neutral, purely technical step.
