Statistics From Scratch by Rob Brotherton

1  Variables & Measurement

  • Introduction
  • 1  Variables & Measurement
  • 2  Frequency
  • 3  Central Tendency
  • 4  Variability
  • 5  \(z\)-Scores
  • 6  Probability
  • 7  Sampling
  • 8  Hypothesis Testing
  • 9  Statistical Power
  • 10  Confidence Intervals
  • 11  The \(t\)-test
  • 12  Independent Samples t Test
  • 13  Related Samples t Test
  • 14  ANOVA
  • 15  Related Samples ANOVA
  • 16  Factorial ANOVA
  • 17  Correlation
  • 18  Regression
  • References

Table of contents

  • Where do statistics come from?
  • Constructs and operational definitions
  • Describing variables
    • Scales of measurement
    • Gray areas
  • Populations, samples, and inference

1  Variables & Measurement

A colorful animated bar chart hides the word STATISTICS.

This is a statistics textbook: you’re here for some math, right? Well, sure, we’ve got that to look forward to. But I named this book Statistics From Scratch. The idea is that we’ll try to build and scrutinize our statistical intuitions first, before formalizing principles with equations and theorems and things like that. I mean, I like math as much as the next person (so, not that much if I’m being honest), but it’s a tool, and to use it well we need to understand what we’re doing and why. So in this chapter, we’ll explore the parts of statistical thinking that come before the math: the logic, technical terms, and practical considerations that allow us to usefully point statistical techniques at social-scientific questions.

And look, we can have some fun here too. See that colorful graph at the top? It’s not just there to look pretty. It actually contains a secret hidden message! Can you figure it out? Think on it a while and feel free to come back for the answer later if you can’t get it right away. When you think you’ve got it, you can type your answer below and check if you’re right.

What is the hidden meaning of the graph at the top of the page?

Answer

It’s the word “STATISTICS” encoded in the heights of the bars. Each bar corresponds to a letter in the word, and the height of each bar corresponds to the position of that letter in the alphabet (S=19, T=20, A=1, etc.).

At the risk of getting too meta too quickly, statistics can be used both to elucidate and to obfuscate. My mysterious little graph was made from some real numeric data, but by leaving out labels or anything else that would help to explain it, I made it hard to understand. Data can be more or less useful, can even tell different stories, depending on how they’re analyzed and presented.

The choices begin even before we have any data. The student, statistician, researcher, journalist, or anyone else trying to understand or explain something about the world faces countless decisions about how to conceptualize, measure, encode, analyze, visualize, and interpret things. There isn’t always a single obviously correct choice. We need to consider how well a choice captures what we want to study, what evidence supports it, what it reveals or conceals, whether we’re being ethical and transparent. These decisions are not minor technical details. They have real consequences for how we understand the world and each other. So it’s important to be thoughtful and critical about our approach to statistics from the outset.

Where do statistics come from?

Another quick quiz to get you thinking about statistics very broadly: what are statistics, what are they for, and how do they get produced? Obviously these aren’t the kinds of questions that have tidy right or wrong answers. It’s just worth taking the time to jot down a few words as we set out on this journey.

Where do statistics come from?

Explanation

There isn’t one complete answer. Statistics are produced when somebody decides what to study, how to define and measure it, whom to collect data from, and how to summarize and present what they find.

Statistics come in many forms and from many sources and are intended to serve many different purposes. The term statistics comes from the Latin status, meaning state or condition. A “State” as in a political entity also came from the same root, and in the 1700s German scholar Gottfried Achenwall coined Statistik, meaning the study of a nation’s political, social, and economic goings on — the state of the State. Taking a census was an early way of collecting and systematizing data on things like the number of births and deaths in a community, rates of employment, incomes, and so on. These obviously aren’t just abstract facts. They’re political information that is collected and reported for a purpose. When thinking about prototypical “statistics” today, the kind of things that often come to mind are numbers in reports from government agencies that produce, say, crime or employment figures, or the news stories that report those kinds of findings.

Contemporary academic research is an extension of this theme. While it may be tempting to think of social scientists as objective observers of the world, we ought to remember that researchers are active producers of data and statistics. A researcher’s choices about what to study, how to study it, and how to analyze and present the results are just that: choices. Had the researcher chosen differently at any point they might have arrived at a different understanding of things.

Here’s an example of how statistics can tell different stories. Researchers used data from the 2020 National Survey on Drug Use and Health to estimate how many children and adolescents under 18 lived with a parent who met the criteria for a past-year, non-nicotine substance use disorder. But they produced two different estimates, using different definitions of “substance use disorder.” One, the DSM-IV criteria, led to an estimate of 9.3 million affected children. Applying the newer DSM-5 criteria to the same data produced an estimate of 16.9 million (Schepis et al., 2025). That’s a big difference — 81% more! Which answer was right? The statistics themselves cannot say. Each estimate followed from a recognized, defensible set of diagnostic criteria. Those different definitions led to different statistical answers, even when the underlying data didn’t change.

Now consider other kinds of questions researchers might be interested in. Maybe we want to know how many people are smokers. Again, we need to define what we mean by a smoker. Is it someone who smoked one cigarette once? In the last year? In the last month? Somebody who regularly smokes? How many cigarettes do you have to smoke to be a regular smoker?

Or think about claims that the current generation of young adults is psychologically different from previous generations somehow. From Boomers to Zoomers, there have always been stories about how “kids these days” are more or less narcissistic, anxious, introverted, outgoing, or whatever.1 These are all interesting, important, and massively nuanced questions. There are many possible ways to define and measure and test each of these claims, and any statistic you see claiming to offer answers is not just some naturally occurring fact, but the end result of a lot of choices and decisions on the part of the researcher.

Constructs and operational definitions

We’re getting into important issues here, not only for statistics but for research methods more broadly, so let’s introduce some technical terms: constructs and operational definitions.

A construct in the context of social-scientific research is generally an abstract phenomenon, attribute, or characteristic that cannot be directly observed. The operational definition specifies how we actually go about measuring that intangible concept.

It’s easiest to understand these terms by thinking about a few examples. Let’s say we think extraversion is something interesting and important to study. You can’t see somebody’s extraversion. You can’t touch it or weigh it. It’s just an abstract concept that we think describes something that exists: some people are more outgoing than others. So extraversion is a construct in this sense. If we want to study this construct, we need a way of measuring it. One widely used model of personality is the Big Five. As the name suggests, it sees personality as consisting of five traits, including our construct extraversion (the others are openness, conscientiousness, agreeableness, and neuroticism). Quite a few questionnaires have been designed to measure the Big Five traits. Their extraversion items generally ask people how quiet, reserved, talkative, or outgoing they are. Responses to those questions, together with the procedure for combining the responses into a score, provide an operational definition: the means by which we measure the construct of extraversion.

Figure 1.1: An example of a Big Five personality scale item measuring extraversion (Goldberg, 1992, 1999).
A five-point agreement scale for the statement: I am the life of the party.

Another example of a construct is intelligence. A researcher might propose that differences in intelligence or kinds of intelligence between people are real and worth knowing about. But you can’t directly access somebody’s intelligence. You need to create some kind of task that you think captures variation in whatever it is you’re calling intelligence. One example of a type of question that might appear on an IQ test presents the series D · F · I · M · R · ? and asks what letter comes next (Condon & Revelle, 2014).2 A single question cannot measure someone’s intelligence any more than one question about being the life of the party can tell you if someone is an extravert, but it again illustrates one way an operational definition can turn an abstract construct into a recorded response.

Lastly, think about something as simple as height. This might seem more straightforward and objective. Unlike personality or intelligence, it’s a physical attribute, at least. But even so, height needs to be defined and operationalized in some way that fits our purposes and whatever equipment we have available. Most obviously, we might record the distance from the floor to the top of a subject’s head according to a tape measure. But if we’re being careful, that’s not quite all there is to it. Shoes on or off? Feet flat on the ground, or are tippy-toes acceptable? Why not? For that matter, why measure from the bottom of the feet to the top of the head? Why not distance from the floor to fingertips at a full stretch? If what we’re really interested in is whether someone can reach something on a high shelf, maybe that’s the more useful thing to measure. And of course, if you didn’t have a tape measure to hand, you’d need to think of some other operational definition entirely.

Height operational definition

Imagine you want to study height but don’t have a tape measure. What would be a useful operational definition of height?

Answer

Many procedures could work. You might stand each person against a wall and mark the top of their head, compare them with an object of known size, or use a phone app calibrated against a known distance.

Explanation

The procedure should be explicit enough for somebody else to repeat and should produce a value you can record. Different procedures may capture slightly different versions of height, so the useful choice depends partly on what you want the measurement for.

The point is that there are a lot of constructs that we might be interested in. And for any given construct, there are usually a lot of potential operational definitions. How we conceptualize a construct and operationalize it determines the variable we record and, in turn, which statistical analyses might be more or less useful and appropriate.

Describing variables

Once we have worked out the definition of our construct and decided on the operational definition we will use, we can think about the variable it produces. By variable, we just mean something that can vary; in other words, whatever we’re measuring can take on different values. (In Chapter 4 we’ll introduce the idea of variability. The terminology can get a little confusing when we get to talking about how variable (adjective) a variable (noun) is, but let’s just not worry about that.) Time is a variable: if you look at the clock now, and again a bit later, the time will be different. But there are different ways things can vary, and it’s good to be precise about how we’re thinking about our construct and its measurement. What kind of values are possible? And what do they represent? We can introduce some more technical terms here to help nail these things down.

First, a variable can be categorical or quantitative. A categorical variable places cases into groups or categories, whereas the values of a quantitative variable represent amounts.

An example of a categorical variable is eye color. Your driver’s license might record your eye color as BRO (Brown), BLU (Blue), GRN (Green), for example. For this variable, the values tell us only which category was recorded, not how much eye color someone has. This distinction depends on what the values mean, not just how they are encoded. Rather than those three-letter category codes, licenses could use numbers: 1 for blue, 2 for green, and so on. But that still wouldn’t mean that green eyes are greater than blue; the numbers would still just be functioning as category labels.

In contrast, the values of quantitative variables do represent numerical amounts. But there are different kinds: a quantitative variable can be either discrete or continuous.

A discrete quantitative variable is one that can take only certain values, usually whole-number counts. An example would be the number of siblings you have. You might have 0, 1, 2, 3, or more, but you can’t have 0.75 or 2.13 siblings. There are no allowable intermediate values between one count and the next.

A continuous quantitative variable, on the other hand, can take fractional or decimal values. Height, for example, is a continuous variable. The actor Jake Gyllenhaal is reportedly 5 feet 11 inches tall.3 But what does that mean? If we held up a tape measure next to him, would the top of his head line up precisely with the 5’11” mark? Probably not. Maybe he’d be a little over, say by a quarter of an inch. We’re just rounding to the nearest inch for convenience, but in fact that recorded value covers a range of possible heights, from 5’10.5” to 5’11.5”. There are an infinite number of possible heights within a given range; in practice, we’re limited by the precision of our measurement tools and by how much precision is necessary for our purposes.

Scales of measurement

A more detailed way of thinking about the nature of variables was proposed by psychophysicist Stanley Smith Stevens (1946). Stevens described four scales of measurement: nominal, ordinal, interval, and ratio. These scales map onto the categorical/quantitative distinction: nominal and ordinal variables are categorical, whereas interval and ratio variables are quantitative. Interval and ratio variables can, in turn, be either discrete or continuous. The scales focus on which comparisons are supported by the values recorded for a variable. You can think of them as cumulative: each scale supports the comparisons of the previous one and adds something new.

Nominal means we are purely recording named categories. There is no quantitative distinction between the categories, meaning it doesn’t make sense to say any one category conveys more or less of the construct in question than another. They simply belong to qualitatively different groups. Eye color is a good example. Brown eyes are not “more” or “less” than blue eyes. They are just different.

Ordinal scales are still categorical, but in this case the categories are ordered. Here, category membership does convey something about who performed better or whose score is higher or lower. An example is Olympic medals (or, more generally, coming in first, second, third, and so on in a race). We know that the person with the silver medal performed worse than the person with the gold and better than the person who received bronze, but we don’t know exactly how much better or worse anyone did.

Interval scales add that information. They tell us who scored higher or lower, and also by how much. The gaps between neighboring values are all quantitatively the same size. In other words, the intervals are all equal. Take temperature: the difference between 10 degrees Celsius and 11 degrees Celsius is the same as the difference between 11 degrees Celsius and 12 degrees Celsius. A single-degree difference means the same thing all the way along the scale.

Ratio scales provide the same information as interval scales, but with one additional feature: a meaningful, nonarbitrary zero point. Consider the interval Celsius temperature scale. It has a zero point, but that point was chosen for convenience: zero degrees Celsius is the point at which water freezes, rather than the absence of temperature. Twenty degrees Celsius is not twice as much temperature as ten degrees Celsius. Reaction time, in contrast, is a ratio variable. Zero means that no time has elapsed. This makes ratio comparisons meaningful. Two hundred milliseconds is twice as long as one hundred milliseconds. The number of questions answered correctly on a test is a discrete ratio variable: zero means that no questions were answered correctly, and ten correct answers are twice as many as five. For this count, neither negative scores nor fractions are possible (assuming each question is marked simply right or wrong and we’re not giving partial credit).

Table 1.1: Summary of the four scales of measurement
Scale Variable type Characteristics Examples
Nominal Categorical Categories only; no inherent order Eye color; experimental condition
Ordinal Categorical Ordered categories; gaps between categories are unequal or unknown Olympic medals; education level
Interval Quantitative Ordered numeric values; equal intervals; arbitrary zero Temperature (Celsius/Fahrenheit); IQ scores
Ratio Quantitative Equal intervals; nonarbitrary zero; meaningful ratios Height; reaction time; counts

Try classifying some variables

0 of 4
1

Social Security number:

Response options

Nominal

Ordinal

Interval

Ratio

Answer

Nominal

Explanation

A Social Security number is an identifier. The numbers are labels, not quantities.

2

Customer satisfaction recorded as poor, fair, good, or excellent:

Response options

Nominal

Ordinal

Interval

Ratio

Answer

Ordinal

Explanation

The categories have a meaningful order, but we do not know that the gap from poor to fair represents the same amount of satisfaction as the gap from good to excellent.

3

Calendar year (for example, 1990, 2000, or 2010):

Response options

Nominal

Ordinal

Interval

Ratio

Answer

Interval

Explanation

Differences between years are meaningful and use equal units. The year 0 is meaningful as a reference date, but it does not represent the absence of time. So a ratio such as 2000/1000 does not represent a useful ratio of the two points in time.

4

Number of text messages received today:

Response options

Nominal

Ordinal

Interval

Ratio

Answer

Ratio

Explanation

This is a count with a true zero: zero messages means none, and twenty messages is twice as many as ten. It is also discrete, but its measurement scale is ratio.

Gray areas

As we’ll see in later chapters, these distinctions can inform our choice of statistical summary, graph, and test. The terms we just laid out are helpful for thinking about the nature of variables, but I don’t want you to think they provide a complete and rigid taxonomy into which we can neatly slot any variable or operational definition. In practice, distinctions aren’t always clear-cut.

To take an example, we mentioned the personality trait extraversion earlier. The construct might be defined, roughly, as the tendency to be outgoing. (I’m vastly simplifying here, but you get the point.) But what kind of variable captures that construct? We could model it as categorical: either you are extraverted or you are not. Or we could model it as quantitative and continuous: extraversion is a spectrum on which you can score high, low, or anywhere in between.

And once you feel you’ve got that figured out, you’re not out of the woods yet, because you still need your operational definition and scale of measurement. I gave an example extraversion question in Figure 1.1. But most measures have a few questions all asking variations of the same thing, as in Figure 1.2. The name for this kind of statement with ordered response options is a Likert item, named after the psychologist Rensis Likert; a set of Likert items becomes a Likert scale. (I have most often heard it pronounced “LIE-kert”, but apparently the correct pronunciation of the name is “LICK-ert”.)

Figure 1.2: Three extraversion items, each with an agreement response scale numbered from 1 to 5.
Three extraversion items, each with an agreement response scale numbered from 1 to 5.

What scale of measurement does each question reflect? It can’t be nominal, because the categories have a natural order from disagreement to agreement. It isn’t ratio either, because there is no true zero point. So interval or ordinal? Numbering them 1 through 5 doesn’t automatically make them interval: the numbers are equally spaced on paper, but that doesn’t mean that the difference between disagreeing and slightly disagreeing is the same as the difference between slightly disagreeing and being neutral. These are verbal labels that we can’t assume have equal intervals. So really the response categories are ordinal.

But usually researchers are not so much interested in people’s responses to individual questions. Rather, they create an overall score by combining a person’s responses across all the questions. How? If our model of extraversion is categorical, we might just count how many of the items a person agreed with (selecting either slightly agree or agree). If they agree with, say, two or more, we might classify them as an extravert; if they agree with one or none, we might classify them as an introvert. We might use a number along the way, but after applying the cutoff the resulting variable is categorical.

If our model of the construct is quantitative, on the other hand, we might treat the responses as numeric and combine them into a single score by averaging them. This is how many Big Five questionnaires are scored. But does averaging ordinal responses produce an interval scale? Not automatically. Like I said before, the gaps between the response categories cannot be assumed to be equal intervals. If people treat the difference between, say, “slightly disagree” and “neutral” as psychologically bigger or smaller than the difference between “slightly disagree” and “disagree”, a simple average of the numeric responses would not reflect this imbalance. But the if there is important: in practice, a carefully constructed Likert scale can behave enough like an interval scale to be treated that way for many purposes (Carifio & Perla, 2008; Norman, 2010).

Obviously, this is all a matter of some (by which I mean copious) debate. In fact, when Stevens proposed the nominal/ordinal/interval/ratio taxonomy, he was responding to a debate that had been started by a committee of the British Association for the Advancement of Science fourteen years earlier, in 1932. There is a hint of frustration and disappointment when he writes that their years of “deliberation led only to disagreement, mainly about what is meant by the term measurement” (Stevens, 1946). The debate goes on. The point for our purposes is that nothing should be taken for granted. It is good to have helpful terms for things, but we always need to think carefully about how we apply them in any given context.

Populations, samples, and inference

The last technical terms we’ll introduce here underlie a lot of what we’ll talk about later in the course when we come to inferential statistics — the process of drawing general conclusions from our data.

The issue is that, generally speaking, social scientists want to understand a population, meaning some entire group. The population we’re interested in might be all humans, or all Americans, or all newborn babies, or whatever. One way of understanding that population would be to record the data we need from every single member of the population. Maybe you can foresee a problem here. Usually doing that is not feasible. For one thing, we can’t ethically force everyone to take part in our study. For another, we probably don’t have time to go around asking every human being if they’re the life of the party, or whatever it is we want to know about them.

All hope is not lost, however. We don’t need to ask everyone. We can ask a sample, a smaller subset of the population, and use that sample to make claims about the population. Much of the statistical logic that we’ll encounter later in this book hinges on this leap from sample to population. A sample provides a clue about the population it came from, but the clue is inherently imperfect because any one sample is not going to perfectly represent the population from which it’s drawn. So how much trust can we place in any given sample?

Think about it this way. If we were to repeatedly draw sample after sample from a population, each sample would look a little bit different and so would give us a slightly different impression of the population from which it was drawn. In practice, we don’t draw sample after sample. We just use one. But we know that our sample isn’t special. It’s just one possible sample out of virtually infinite possibilities. If we’d happened to end up with a slightly different sample, our inferences might be different.

ActivityDrawing samples from a population
A balanced population of 100 colored dots and repeated random samples of ten dots.

Here’s a population of 100 colorful dots. You’ll notice it has the same number of dots of each color. When you press the next button, we’ll take a sample of 10 dots from that population. If the sample were to be perfectly representative, it would contain two dots of each color. Do you think that’s what will happen?

In this case, the sample is not perfectly representative. This sample has only one blue dot but three orange, for example.

The next sample has three blue and no orange dots! The population stays the same, but each random sample will be different.

The trouble is that generally we don’t know the population’s true characteristics. Without access to the population, the sample provides our best guess, but each guess will be different.

Here’s a third sample. It provides a different impression again of the population’s colors, getting the proportions of orange, red, and purple exactly right but overestimating the prevalence of green and missing the existence of blue entirely.

Keep clicking the button as many times as you like. I’ve revealed the true population again, so you can see that while it never changes, each sample offers a different estimate of its characteristics.

How do we deal with the inherent variability of samples and our desire to make stable inferences about the population? That’s what we’ll spend much of this book figuring out. We’ll set it aside for a few chapters and circle back in Chapter 7. For now, I just want to introduce the terminology that will help us keep these distinctions in mind. A numerical value describing a population is called a parameter. The corresponding value calculated from a sample is called a statistic.

Table 1.2: Summary of the distinction between population parameters and sample statistics.
Feature Population Sample
Term Parameter Statistic
Scope Describes the whole population Describes the observed sample
Common notation Often Greek letters, such as \(\mu\) or \(\sigma\) Often Latin letters, such as \(M\) or \(s\)
Value Fixed for a defined population, though usually unknown Depends on the sample and will usually differ from one sample to another

So, for example, we’ll talk about the mean or average in Chapter 3. A population mean is a parameter, represented by the Greek letter \(\mu\) (mu). A sample mean is a statistic, which we will represent using an uppercase italic \(M\).

These conventions aren’t completely consistent. If we’re talking about the number of individuals in a population or sample, for example, we use uppercase \(N\) for the population and lowercase \(n\) for the sample.

The distinction between populations and samples is related to, but not the same as, the distinction between descriptive and inferential statistics. Descriptive statistics summarize data we have observed. Inferential statistics use sample data to make an informed guess about the population from which the sample was drawn. The same sample statistic can be involved in both. Reporting a sample mean describes the sample; using that sample mean to estimate a population mean is an inference.

This terminology, and the differences in how we represent populations and samples, will be important to keep in mind as we move forward.

Learning Checks

0 of 6
1

Which term refers to the concrete way a researcher measures an abstract construct?

Response options

Construct

Operational definition

Variable

Parameter

Hint

Distinguish the abstract idea, the procedure used to measure it, and the values that procedure produces.

Answer

Operational definition

Explanation

A construct is the abstract characteristic, such as extraversion or intelligence. The operational definition is the actual procedure used to measure it.

2

A researcher records the reaction times of 30 students selected from a university. What do we call the mean reaction time of those 30 students?

Response options

Population parameter

Sample statistic

Sample

Variable

Answer

Sample statistic

Explanation

The mean is a numerical value calculated from the sample, so it is a sample statistic. The true mean reaction time of all students at the university would be a population parameter.

3

In an experiment, the control condition is coded 0 and the treatment condition is coded 1. What kind of variable is condition?

Response options

Nominal categorical

Ordinal categorical

Discrete quantitative

Continuous quantitative

Hint

Ask whether 0 and 1 represent amounts or labels for groups.

Answer

Nominal categorical

Explanation

The codes are labels for two groups. Coding the conditions with numbers does not make the variable quantitative, and neither condition is inherently higher than the other.

4

A questionnaire asks people to rate three extraversion statements from 1 (disagree) to 5 (agree). The individual responses are ordinal. The researcher then averages them. Which conclusion is most careful?

Response options

Treating the average as approximately interval needs further justification

The average is automatically an interval variable

The average must remain ordinal in every analysis

The average becomes a ratio variable

Hint

Distinguish what we know about each response from the modeling choice we might make about a carefully constructed combined score.

Answer

Treating the average as approximately interval needs further justification

Explanation

The numbers attached to the response categories do not guarantee equal intervals, and averaging does not change that automatically. But a carefully constructed scale can sometimes behave enough like an interval variable to make that a useful modeling choice.

5

Age is commonly reported in whole years. What kind of variable is age itself?

Response options

Continuous

Discrete

Answer

Continuous

Explanation

Time is infinitely divisible: between any two ages there is always another possible age. Reporting age in whole years is a choice about measurement precision, not a property of the underlying variable.

6

A sample of 20 students has a mean reaction time of 310 milliseconds. Which statement uses that sample statistic to make an inference?

Response options

The sample mean was 310 milliseconds

The sample contained 20 students

The university mean is probably close to 310 milliseconds

Reaction time was recorded in milliseconds

Answer

The university mean is probably close to 310 milliseconds

Explanation

The first, second, and fourth statements describe the observed sample or its measurement. Estimating the university mean goes beyond those observations to make an inference about the population.

Carifio, J., & Perla, R. (2008). Resolving the 50-year debate around using and misusing Likert scales. Medical Education, 42(12), 1150–1152. https://doi.org/10.1111/j.1365-2923.2008.03172.x
Condon, D. M., & Revelle, W. (2014). The international cognitive ability resource: Development and initial validation of a public-domain measure. Intelligence, 43, 52–64. https://doi.org/10.1016/j.intell.2014.01.004
Goldberg, L. R. (1992). The development of markers for the big-five factor structure. Psychological Assessment, 4(1), 26–42.
Goldberg, L. R. (1999). A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models. In I. Mervielde, I. Deary, F. De Fruyt, & F. Ostendorf (Eds.), Personality psychology in europe (Vol. 7, pp. 7–28). Tilburg University Press.
Norman, G. (2010). Likert scales, levels of measurement and the “laws” of statistics. Advances in Health Sciences Education, 15(5), 625–632. https://doi.org/10.1007/s10459-010-9222-y
Schepis, T. S., Veliz, P. T., West, B. T., McCabe, V. V., Hulsey, E., Kcomt, L., & McCabe, S. E. (2025). US youth exposed to parental substance use disorder in the home: A comparison of DSM-IV and DSM-5 criteria. Journal of Addiction Medicine, 19(6), 722–725. https://doi.org/10.1097/adm.0000000000001469
Stevens, S. S. (1946). On the theory of scales of measurement. Science, 103(2684), 677–680. https://doi.org/10.1126/science.103.2684.677

  1. Seriously, this has been going on for a while. A couple of millennia ago, the Roman poet Horace put a self-deprecating spin on this idea, writing, “Our sires’ age was worse than our grandsires’. We, their sons, are more worthless than they; so in our turn we shall give the world a progeny yet more corrupt.” (Horace, Book III of Odes)↩︎

  2. The answer is X: the jumps through the alphabet increase from one place to two, then three, four, and five.↩︎

  3. This example is inspired by the surprisingly detailed and interesting investigation into Gyllenhaal’s height which was part of the Mystery Show podcast episode Source Code. Available on Apple Podcasts, Spotify, or wherever you get your podcasts.↩︎

Introduction
2  Frequency