Statistics studies the full data workflow: collection, organization, analysis, interpretation, and presentation.
Statistics is the discipline concerned with the collection, organization, analysis, interpretation, and presentation of data. In practice, it often begins with a statistical population (the full group of interest) or a statistical model, and then uses data collection plans such as survey design and experiment design. When a census is not possible, statisticians use sampling so that conclusions drawn from a sample can reasonably extend to the population. Statistics covers how data are gathered and studied through both experimental and observational approaches. Experimental studies involve manipulating a system and measuring outcomes to assess causal effects, while observational studies collect data without manipulation and examine associations. Two major branches of data analysis are descriptive statistics, which summarize data (e.g., mean and standard deviation), and inferential statistics, which use probability theory to draw conclusions about populations despite random variation such as sampling error and measurement error. A key part of inferential statistics is hypothesis testing and interval estimation, typically framed around a null hypothesis and an alternative hypothesis. Statistical tests quantify how evidence from data can lead to rejecting the null, while recognizing possible errors (Type I and Type II). Statistics also addresses measurement error, bias, missing data, and censoring, using specialized techniques to reduce or correct their impact.
Statistics studies the full data workflow: collection, organization, analysis, interpretation, and presentation.
When full population data canβt be collected, sampling and experimental/observational study designs are used to support valid inference.
Inferential statistics uses probability theory to make population-level conclusions under uncertainty, including hypothesis testing and interval estimation while accounting for errors and bias.
The complete set of people or objects (and their characteristics) that are of interest for a study.
A mathematical representation of how data are generated, used to support analysis and inference.
The process of selecting a subset of a population to collect data, with the goal of enabling conclusions about the full population.
A study where the researcher manipulates the system (e.g., treatments) and measures outcomes to assess causal effects.
A study where data are collected without experimental manipulation, focusing on associations between variables.
Methods that summarize and describe data using quantities such as mean, dispersion, and other distribution features.
Methods that use sample data and probability theory to draw conclusions or predictions about a population under uncertainty.
A proposed statement of no (or a baseline) relationship that is assessed using statistical tests.
The error of rejecting the null hypothesis when it is actually true (a false positive).
The error of failing to reject the null hypothesis when it is actually false (a false negative).
A descriptive property of a distribution that characterizes its typical or central value.
A descriptive property of a distribution that measures how spread out values are around the center.
βCan you explain what "Statistics studies the full data workflow: collection, organization, analysis, interpretation, and presentation." means in simple terms?β