Statistics, as a field of mathematics with broader developments and applications across the sciences, is recent when compared with the other branches of mathematics. Although older writings and treatises with some grounding exist, much of the statistics we use today was developed mostly within the last 400 years (or even less). And contrary to the contemporary view we have of researchers, science and mathematics, the way these areas were cultivated in the past reads almost as antagonistic to the stereotype of the science nerd, the one who holds very specific knowledge, apparently with no practical application in daily life, developed in some obscure corner of a university.
Since I am no historian, I will not make the blunder of offering a definitive diagnosis of this relationship. But the way I like to think about it is that, before the hyper-specialization of scientific development, exploring new fields started, naturally and incipiently, from personal problems and interests. And when there is no established reference framework, these individual interests end up echoing broader interests shared collectively.
With that idea in mind, it seems somewhat logical to suppose that the first developments of statistics happened in fields such as gambling, something commonly seen as a popular interest. And a person with a bit more knack for that kind of reasoning would start looking into the subject. That is exactly what happened. According to Ian Hacking, the outcomes of bets were seen at the time as a direct decision of God, and randomness and chance were not concepts applied to this kind of problem. A gambler with a betting system for dice games that seemed profitable, however, did not fit the perception that God directly decided the outcome. It was from questions like this one that Blaise Pascal and Pierre de Fermat exchanged letters discussing the mathematical grounding of these problems. And although each had his own approach, from these discussions came the first ideas and rules of probability, starring a legal practitioner (Fermat) and a prodigy without formal academic education (Pascal). In a rather unconventional way for the present day, in which mastery over one area seems to preclude going deep into academically distant areas (at least in the modern view of the exact, social and biological sciences), these are our first formal pioneers of statistics. These exchanges, along with other similar developments of the time, later crystallized into the first rules of probability, in which chance stopped being something outside human reach and became, to some degree, predictable.
From that development, or rather, from those letters about the probabilities of a dice game, a little over half a century later Abraham De Moivre used these findings in his gambling manuals, in which he explained mathematically why and how to bet. This way he produced one of the most famous gambling manuals, with some notable findings, such as the approximation of a complex coin game (what we now call the binomial distribution) by a normal curve. Unlike Pascal and Fermat, De Moivre made these developments as a way to earn a living. With no university degree, he supported himself with private lessons and with his large practice as a gambling consultant, and gained enough recognition to join the Royal Society committee that judged who had invented calculus, Leibniz or Newton.
From then on, the development of statistics, or of this pre-statistics, like that of all science, became more institutionalized. Although some figures from outside academia made extremely important contributions, didactically the contributions of Bayes, Gauss and Laplace may represent this phenomenon well. On one side, we have Thomas Bayes, a Presbyterian minister trying to understand how beliefs adjust as new evidence arrives. This led him to create one of the fundamental theorems of modern statistics. Bayes never published this work, but after his death his notes were taken to the Royal Society, which institutionalized the concept.
On the other side, we have mathematicians such as Laplace and Gauss, who consolidated several of these existing ideas. Gauss used the normal curve to ground the method of least squares, and Laplace demonstrated a version of what we now know as the central limit theorem, which shows the importance of this curve. The curve ended up named after Gauss, even though De Moivre's work had already described it. At this point, the curious minister and the career mathematicians coexisted in the same field, but the center of gravity was shifting toward the institutionalization of research.
Once this statistical base and the understanding of numbers, odds and probabilities advanced, the next step would be to apply these ideas to daily life, as a way of understanding natural phenomena through this lens. However, as Jakob Bernoulli had already shown with the law of large numbers, a large amount of data was needed for these laws to work. From the eighteenth century on, governments across Europe began keeping records of births, deaths and crimes, and here we would have the perfect meeting between this new field, which tried to describe phenomena from their occurrence or chance, and a number of records large enough to feed these models. Adolphe Quetelet had a background in mathematics, devoted himself to applying the field to astronomy and was a great lover of the arts, having written poetry and even an opera. His great contribution was noticing that the error curve used in astronomy could be used in a similar way to understand human traits. Analyzing Scottish soldiers and French recruits, he arrived at the idea that nature would be "aiming" at an ideal body type and that individuals would be the errors of that aim, as if the ideal were the center of a distribution and the individuals the spread around it: l'homme moyen (the average man). Along with these measurements, he also created what we now know as the Body Mass Index (BMI), although he called it the Quetelet Index, with the same formula we use to this day.
Following Quetelet's trail, Francis Galton, Darwin's cousin, studied mathematics and medicine (dropping the latter and actually graduating in the former) and, before devoting himself to statistics, was known as an explorer of southwestern Africa. Galton went deep into the other side of the normal curve. Instead of the mean, which holds the errors, he decided to look at the tails, trying to understand heredity and traits out of the ordinary. However, while studying physical traits, he noticed that even parents far from the mean tend to have children closer to it: tall parents have tall children, but slightly shorter than they are (which brings them closer to the distribution's mean), and, in the same way, short parents have slightly taller children. Despite the advances, Galton's motivations in this research are far from the best. Creator of the term "eugenics", Galton used this idea to motivate his research, trying to understand how genes (an anachronistic term here) controlled populations and their traits, and he defended these ideas for a good part of his life, supported somewhat naively by rudimentary understandings of how heredity worked (there were contemporary critics of eugenics, so the judgment is not entirely anachronistic). Standing on Galton's shoulders, Karl Pearson, a mathematics graduate with top honors, formalized his work while trying to prove Darwinian evolution, and also formalized the concept known to this day as standard deviation. With the same enthusiasm, he analyzed data from the Monte Carlo roulette to check the randomness, or the fairness, of that kind of game, which led him to develop the chi-square test, improving the understanding of how real data fit theoretical models. Influenced by Galton's convictions, Pearson was an even more fervent defender of eugenics, founding the journal Annals of Eugenics and defending the idea throughout his life.
The history of statistics unfolds much like that of science as a whole: many scientists with relevant contributions were not specialized people who had passed through academic institutional filters. Across these two and a half centuries, we see names barely celebrated and even barely known, with the most varied stories: a magistrate, a refugee with no degree at all, a curious minister, an eager explorer and mathematicians interested in many subjects, such as astronomy or even Darwin's theory, trying to use this tool to take a step forward in their fields. Statistics was born from interests that were anything but noble or aseptic, such as games of chance, religious curiosity and dark motives like eugenics. And, ironically, it was this mix of human passions that made it one of the most powerful tools of modern science. With the institutionalization that followed, this plurality of profiles dwindled. The field moved closer to the modern conception of specialists with years of academic study, leaving the curious to follow these ideas as a hobby, with little room for relevant contributions.
References
- HACKING, Ian. The Emergence of Probability: A Study of Early Ideas about Probability, Induction and Statistical Inference. 2nd ed. Cambridge: Cambridge University Press, 2006.
- HACKING, Ian. The Taming of Chance. Cambridge: Cambridge University Press, 1990. (Ideas in Context, 17). ISBN: 9780521388849.
- STIGLER, Stephen M. The History of Statistics: The Measurement of Uncertainty before 1900. Cambridge, MA: The Belknap Press of Harvard University Press, 1986. ISBN: 0674403401.