Probability is a measure of the possibility of an event. It is an intuitive concept, easy to understand, and was defined as a value between zero and one.
Probability is not about data but events.
Events become data when measured or aggregated (counted, summed, averaged, etc.)
A person's age lower than 25 has a certain probability in a country.
When we start collecting information about some persons' age, we create a sample of the country's entire population. At that moment, we started working on a different field, Statistics.
The first interest when handling data is to describe it. Descriptive Statistics is about processing the data to get valuable information like range, mean, and distribution.
Most of the time, we are more interested in using that data sample to make inferences about the full population, Statistical Inference.
Some questions requiring inference could be:
- What is the median age of the population?
- What is the age distribution in the country?
- What is the percentage of the people of working age?
Each of these questions will have an answer with some degree of uncertainty. It is expected because we are using a sample of the population and not a census.
The inference process should provide the answer and a measure of the uncertainty.
Statistics (or its modern definition) was born in the 18th century when computational methods were primitive. As a result, statistical methods tried to avoid complex calculations and relied on a highly developed theoretical foundation.
Statisticians discovered a few standard statistical distributions likely to represent many real-life random variables. These distributions were widely analyzed and used to derive statistical decisions and results for many practical situations.
As good as it appears, there is always a basic assumption that, in the end, made for unreliable results. Every statistical method assumed that the variable under analysis followed some standard distribution. Even testing the premise, the empirical distribution of the variable was never the theoretical one. That mismatch will always put in question the final statistical results.
During the 20th century, computers became part of our lives, and Statistics started taking advantage of them. The initial improvements came from applying known techniques with computers much faster than before and to much bigger datasets.
At some point, computers were so fast that statisticians started getting rid of previous restrictions and taking full advantage of modern tools, powerful hardware, and state-of-the-art software.
Suddenly, it became possible to create generic statistical methods based on intuitive and straightforward bootstrapping methods, supported by massive Montecarlo simulations.
In some way, Statistics went back to basic descriptive properties and foundational concepts.
We prefer to name this evolution Computational Statistics, the creative use of state-of-the-art computing platforms to reshape Statistics as a more intuitive way of looking at random data.