🎥 Watch this before starting!

Learning Objectives.

1 Historical Origins of Intelligence Testing

Sir Francis Galton (1822-1911), a cousin of Charles Darwin, was one of the earliest researchers to study and attempt to measure intelligence. Drawing on Darwin’s theory of evolution, Galton believed that certain families were born with innate traits such as intellect and physical strength. As a result, he strongly favored a nature-based view of intelligence. In fact, the phrase nature vs. nurture originates from Galton’s book English Men of Science: Their Nature and Nurture.

Around the same time, the French government passed a law requiring all children to attend school. To implement this policy effectively, officials wanted a method to identify children who were unlikely to benefit from standard classroom instruction. In 1904, psychologist Alfred Binet (1857-1911) was commissioned to develop an objective tool to measure children’s intellectual abilities.

Binet proposed that while all children progress along a similar developmental path, they do so at different rates. To quantify this, he introduced the concept of mental age—the age at which the average child performs at a given level. For instance, if a child performs like the average 8-year-old, their mental age is 8. This is different from a child’s chronological age (i.e., time-based age).

The intelligence quotient (IQ) was later formalized by German psychologist William Stern (1871-1938), who developed the following formula:

\[ IQ = 100 \times \left( \frac{\text{Mental Age}}{\text{Chronological Age}} \right) \]

For example, a 10-year-old child with a mental age of 7 would have an IQ of 70, while one with a mental age of 13 would have an IQ of 130.

Binet believed that intelligence was shaped primarily by nurture, and he intended his test to identify children in need of additional support—not to label or limit them. Unfortunately, standardized intelligence tests have sometimes been used to exclude or marginalize both individuals and entire groups of people, rather than to offer help or resources. We’ll discuss the discredited practice of eugenics later in the lecture.

Later, psychologist Lewis Terman (1877–1956) at Stanford University adapted Binet’s test into what became the Stanford-Binet Intelligence Scale. Terman extended the test to include adult norms and promoted its use as a general measure of intellectual ability across the lifespan.

Building on this tradition, David Wechsler (1896-1981) introduced the Wechsler Adult Intelligence Scale (WAIS) in 1955, evolving from his earlier Wechsler–Bellevue Intelligence Scale (1939). The WAIS offered separate verbal and performance subscales, setting a new standard in adult intelligence testing. It has undergone several major revisions—WAIS-R (1981), WAIS-III (1997), WAIS-IV (2008), and WAIS-5 (2024)—each aimed at improving psychometric properties, updating norms, and incorporating insights from modern neurocognitive science. The WAIS is now part of a broader family of Wechsler intelligence tests, including the Wechsler Intelligence Scale for Children (WISC) and the Wechsler Abbreviated Scale of Intelligence (WASI). Together, these tools represent some of the most widely used and comprehensive assessments of intelligence across the lifespan.

2 Theories of Intelligence

Around the same time that Binet was working in France, English psychologist Charles Spearman (1863–1945) used correlation analysis to examine student performance across various academic subjects. In a paper published in 1901, he found that scores in subjects like classics, French, mathematics, and music were all positively correlated, suggesting the presence of a common underlying factor. Spearman referred to this pattern—what he called a “perfectly constant hierarchy”—as evidence for a general intelligence factor, later known as “g”. Using an early form of factor analysis–a statistical method used to identify the underlying latent variables (or factors) that explain the patterns of correlations among a set of observed variable–he proposed that these correlations could be explained by a single latent ability influencing performance across domains.

🎥 Watch this (optional)! Note that due to the small text in the video, you should press the fullscreen icon button to expand the video.

Learning Objectives.

L.L. Thurstone (1887–1955), a fellow psychometrician, was an early critic of Spearman’s concept of generalized intelligence. Instead, he proposed that intelligence consists of seven distinct clusters of primary mental abilities: verbal comprehension, word fluency, number facility, spatial visualization, associative memory, perceptual speed, and reasoning.

Later, psychologist Raymond Cattell (1905–1998) introduced a distinction between crystallized and fluid intelligence, which was further expanded by his student John Horn (1928–2006). Fluid intelligence (Gf) refers to the ability to solve novel problems, reason abstractly, and adapt to new situations without relying on prior knowledge. In contrast, crystallized intelligence (Gc) reflects accumulated knowledge, skills, and experience—such as vocabulary and factual information—gained through education and cultural exposure. Fluid intelligence tends to peak in early adulthood and decline with age, whereas crystallized intelligence generally increases across the lifespan.

Howard Gardner also challenged the idea of a single intelligence or even a small set of abilities. He proposed a theory of multiple intelligences, arguing that people possess a range of relatively independent cognitive capacities that evolved to solve different types of problems. Gardner’s model includes traditional domains like verbal and mathematical intelligence, but also expands to naturalistic, bodily-kinesthetic, interpersonal, and other forms of intelligence.



Diagram illustrating Howard Gardner's theory of multiple intelligences, including eight distinct types of cognitive strengths.
Figure 1. Howard Gardner’s theory of multiple intelligences proposes that people possess a variety of cognitive strengths. These include verbal-linguistic (word smart), logical-mathematical (logic smart), visual-spatial (picture smart), musical (music smart), bodily-kinesthetic (body smart), interpersonal (people smart), intrapersonal (self smart), and naturalistic (nature smart) intelligences. Source: Wikimedia Commons.



Gardner often highlights the example of savants, such as Kim Peek, the real-life inspiration for the film Rain Man. Peek reportedly memorized more than 7,600 books but required support for many everyday tasks—suggesting highly specialized but uneven cognitive abilities.

🎥 Watch this! Follow this link to a YouTube video that tells the story for Kim Peek.

Critics of Gardner’s theory argue that it broadens the concept of intelligence too far, encompassing traits that might be better classified as talents or personality dimensions. They advocate for keeping intelligence more narrowly focused on academic and cognitive performance.


3 Assessing Intelligence

Intelligence testing plays an important role in both child development and adult neuropsychological assessment. It is often used to identify conditions like Intellectual Disability (ID) or to evaluate cognitive functioning following neurological disease or injury, or psychiatric disorders. In assessment contexts, intelligence can refer to aptitude (natural ability) or achievement (acquired knowledge), concepts that map onto Cattell and Horn’s theory of fluid and crystallized intelligence, respectively. Though distinct, aptitude and achievement are moderately to strongly correlated, as higher aptitude often supports greater learning. However, the relationship is not perfect. Many individuals defy this pattern; that is, show low aptitude but high achievement, or high aptitude but low achievement.

3.1 Examples of Aptitude Tests

Wechsler Adult Intelligence Scale (WAIS-5)
A widely used adult cognitive battery (ages 16:0–90:11) with five primary index scores. Administration time: ~45 min for the 7-subtest FSIQ, ~60 min for the 10 primary index subtests. :contentReferenceoaicite:0

  • Verbal Comprehension (VCI)
    • Similarities — explain how two things are alike (e.g., “How are a poem and a statue alike?”)
    • Vocabulary — define words (e.g., “What does ‘transparent’ mean?”)
  • Visual Spatial Ability (VSI)
    • Block Design — recreate patterns with colored blocks
    • Visual Puzzles — choose pieces that form a target shape
  • Fluid Reasoning (FRI)
    • Matrix Reasoning — select the missing piece in a visual pattern
    • Figure Weights — balance-scale analogical reasoning with figures
  • Working Memory (WMI)
    • Digit Sequencing — repeat numbers in ascending order
    • Running Digits — recall the most recent digits from a continuous stream
  • Processing Speed (PSI)
    • Coding — copy symbol–digit pairs under time limits
    • Symbol Search — quickly mark whether target symbols appear


Photo of the WAIS-R Block Design subtest, showing red-and-white cubes and a pattern card to be recreated.
Figure 2. The Block Design subtest of the WAIS-R requires participants to recreate visual patterns using red-and-white cubes within a time limit. It is a measure of visuospatial reasoning and processing speed, and is commonly used in neuropsychological assessment. Source: Wikimedia Commons.



Raven’s Progressive Matrices: A nonverbal test of abstract reasoning and pattern recognition. Participants select the missing piece that completes a visual matrix.



Example of a Raven's Progressive Matrices problem, showing a 3x3 grid with one missing piece and four answer choices labeled A through D.
Figure 3. Raven’s Progressive Matrices is a nonverbal test used to measure abstract reasoning and fluid intelligence. Participants are asked to identify the missing piece that completes a pattern, making it useful for cross-cultural assessments and minimizing the influence of language or formal education. Source: Wikimedia Commons.



Wonderlic Personnel Test: A short test commonly used in employment settings to assess problem-solving and learning potential. Includes numerical ability, verbal reasoning, and logic questions.

🏈 Fun Fact: The Wonderlic Personnel Test was famously used by the NFL as part of the player evaluation process during the draft. The test was administered to prospects at the NFL Scouting Combine. Its use has declined in recent years, and its predictive value for athletic performance is debated.

Armed Services Vocational Aptitude Battery (ASVAB): The ASVAB is used by the U.S. military to assess aptitude in areas such as mechanical comprehension, arithmetic reasoning, and electronics. Test scores are used to place recruits in roles aligned with their cognitive strengths.

3.2 Examples of Achievement Tests

Achievement tests measure what a person has already learned—often in a school or training context. These tests emphasize crystallized intelligence and are used to assess mastery of specific content areas.

Wechsler Individual Achievement Test (WIAT-4): Measures academic achievement across Reading, Writing, Mathematics, and Oral Language (ages 4:0–50:11). Commonly used to identify learning disabilities and instructional needs.

  • Reading
    • Word Reading / Pseudoword Decoding
    • Oral Reading Fluency
    • Reading Comprehension — read a passage and answer questions
  • Mathematics
    • Numerical Operations — calculations (e.g., fractions)
    • Math Problem Solving — word problems with real-world contexts
  • Writing
    • Spelling — dictated words
    • Sentence Writing Fluency
    • Essay Composition
  • Oral Language
    • Listening Comprehension
    • Oral Expression

Woodcock-Johnson Tests of Achievement: A comprehensive battery for measuring academic achievement and cognitive abilities. Covers areas like reading fluency, calculation, and academic knowledge.

Scholastic Achievement Tests (e.g., SAT Subject Tests, AP Exams): Standardized tests that assess knowledge in specific academic domains (e.g., U.S. History, Biology, Calculus). These are often used for college placement and credit.

Statewide Standards-Based Assessments: Examples include state-level exams like the Colorado Measures of Academic Success (CMAS). These tests assess students’ proficiency in subjects aligned with state curriculum standards.


4 Psychometrics

Psychometrics is the scientific study of psychological measurement. It relies heavily on statistical methods, although some non-statistical approaches are also used. The core purpose of psychometrics is to determine whether psychological measures accurately reflect the abilities, traits, behaviors, or concepts they are designed to assess. Psychometric theory and methods have been so successful that they are now widely used to study measurement in other fields beyond psychology, including education, healthcare, marketing, and organizational science.

Psychometrics developed in close connection with the origins of intelligence testing. For example, figures like Francis Galton (along with his colleague Karl Pearson), Charles Spearman, and L.L. Thurstone were all deeply involved in the early development of psychometric theory.

Psychometrics is primarily concerned with two related concepts: reliability and validity.

Validity can take several forms, including:



Four target diagrams illustrating the relationship between reliability and validity in measurement.
Figure 4. The diagram shows how reliability and validity relate to each other. A measurement can be unreliable and invalid, reliable but not valid, valid but unreliable, or both reliable and valid. The goal in psychological testing is to achieve both high reliability and high validity. Source: Wikimedia Commons.



The psychometric properties of intelligence tests and other psychological measures can be improved through a process called standardization. Standardization involves comparing an individual’s score to the scores of a large, representative sample of people who completed the same measure under identical testing conditions.

For example, most intelligence tests are standardized and require the psychometrist (a trained professional who administers and scores psychological or neuropsychological tests) to follow strict procedures, including reading verbatim prompts and using specific scoring rules. The benefit of this approach is that it allows psychologists to determine whether an individual’s performance is typical or atypical by comparing it to a known distribution of scores—often from a normative sample of 1,000 or more individuals.

More precisely, a standardization sample is expected to yield scores that follow a normal distribution. The normal distribution—commonly referred to as a bell curve—is a symmetrical probability distribution in which most individuals score near the average (center), and fewer score very high or very low.



Normal distribution curve showing percentages of data within 1, 2, and 3 standard deviations from the mean.
Figure 5. The normal distribution is a bell-shaped curve where approximately 68% of values fall within ±1 standard deviation, 95% within ±2, and 99.7% within ±3. This distribution is often used to describe standardized test scores and other naturally occurring variables. Source: Wikimedia Commons.



When we compare an individual’s score to this distribution, we can use a mathematical tool called the cumulative density function (CDF) to determine the proportion of the population scoring below a given value. This allows us to assign a percentile rank. For example:

Standardization allows clinicians, researchers, and educators to provide meaningful, comparative interpretations of test scores, offering valuable context for diagnosis, treatment planning, and feedback.

🌐 Explore this! Follow this link to a Inch Calculator website that has an IQ percentile calculator.


4.0.1 Genetic and Environmental Influences on Intelligence

Even before modern genetics, many thinkers argued that physical and mental traits were inborn (nature) rather than primarily shaped by nurture. Others, however, emphasized environmental influences, so the debate long predated genetics.

Unfortunately, proponents of genetic explanations of intelligence have often used this perspective to justify the systematic exclusion of certain groups. Francis Galton, who coined the term eugenics, promoted ideas aimed at improving the human population—ranging from encouraging reproduction among certain groups to harmful policies like forced sterilization, restrictive immigration, and in extreme cases genocide. The eugenics movement is now discredited for both ethical and scientific reasons, as many claims ignored confounding variables and relied on flawed assumptions. For example, intelligence test data during World War I was misused to suggest ethnic and racial hierarchies, reinforcing pseudoscientific arguments about inherited intelligence and supporting institutional racism.

Decades of research have shown that average IQ differences between socially defined racial or ethnic groups are statistically small and have been narrowing over time. While these differences were historically misused to justify claims of genetic inferiority, the modern psychological consensus emphasizes that individual differences within ethnic and racial groups far exceed those between them. On the other hand, environmental factors—such as access to quality education, socioeconomic status, stereotype threat, and early nutrition—play a major role in shaping cognitive performance. When socioeconomic and contextual variables are considered, racial group differences in IQ tend to shrink and disappear altogether.

Of course, this is not to say that genetic factors have no influence on intelligence—they do. However, Galton both greatly overestimated the role of heredity and incorrectly attributed group differences in intelligence to genetics, ignoring clear and powerful environmental influences. Intelligence is related to genetics in complex ways that are not yet fully understood. This is partly because there are multiple types of intelligence and because there are differing views on what constitutes intelligence across cultural groups.

However, research comparing twins and unrelated individuals suggests that intelligence has a strong genetic component at the level of the individual. For example:

  • Identical twins reared together show a correlation of ~0.85 in IQ.
  • Identical twins reared apart show a correlation of ~0.70.
  • Unrelated individuals reared together show a correlation of ~0.31.

These findings suggest that approximately 50% of the variance in intelligence is attributable to genetic factors (i.e., squaring a correlation tells us the variance accounted for). The remaining variance is influenced by environmental factors and gene–environment interactions. In other words, both nature and nurture play important roles in shaping intelligence.

🎥 Watch this! Follow this link to a TED-Ed video that expands on the dark history of IQ testing.

5 Bias in Testing



Cartoon image showing an American yelling 'FEET!' and a European yelling 'METERS!' to illustrate a cultural disagreement.
Figure 6. A humorous example of a cultural clash—here between the use of imperial versus metric systems—highlighting how even simple conventions like units of measurement can reflect deeper cultural differences. Cartoon by Cox & Forkum, © 2003. Source: Cox & Forkum.



A concern often raised of intelligence and other standardized testing is that of bias. Bias occurs if score differences on the indicators of a particular construct do not correspond to differences in the underlying trait or ability. Bias in psychological testing typically falls into three main categories:

1. Construct Bias: Construct bias occurs when the construct being measured is not conceptually equivalent or valued across cultural or demographic groups. A prominent example is the concept of intelligence. Most intelligence measures are rooted in Western/European traditions, and even within Western psychology, definitions of intelligence are highly debated, and multiple competing theories exist. As an example, Kenyans define the more prized trait of wisdom as a combination of rieko (knowledge and skills), luoro (respect), winjo (comprehension of how to handle real-life problems), and paro (initiative) (Grigorenko et al., 2001). Only one of these terms (rieko) matches closely with the Western concept of intelligence as measured on most tests.

2. Method Bias: Method bias refers to problems arising from the methodology used in assessment. It includes: Sample bias: Occurs when comparison groups differ in ways unrelated to the primary construct being studied (e.g., comparing cognitive test scores between two populations when one group has been affected by war, malnutrition, chronic stress, etc.); Instrument bias: Results from differences in familiarity with test content, response procedures, or response styles (e.g., social desirability bias); and Administration bias: Arises from examiner-examinee interactions (e.g., language barriers, cultural misunderstandings, or examiner expectancy effects).

3. Item or Measurement Bias: Item bias—also referred to as measurement bias—occurs when a test item reflects not just the intended ability but also knowledge of specific information that is highly related to demographic factors such as race, gender, or cultural background. In these cases, performance on an item is influenced by extraneous variables, undermining the fairness and validity of the measure. As an example, history items related to war have been shown to give males positively biased scores, presumably because of gender stereotypes that expose males to more war knowledge and trivia (e.g., through watching war movies). Thus, if we were to heavily emphasize war knowledge on a general history exam, this would, in theory, favor males.


🎥 Watch this before ending!

Lecture Summary