← back to the library 🍬 Psychology Candy

Six Months Later, They Said No

In the early 1930s a Stanford sociologist named Richard LaPiere drove across the United States with a young Chinese couple, and kept a written record of every place they stopped for the night or for a meal. It was not a kind decade for that particular trip. Anti-Chinese sentiment was thick enough in American life to have been legislated once already, and a Chinese couple walking into a rural hotel lobby with a white professor was going to be noticed. LaPiere logged 251 hotels and restaurants over roughly two years of travel. They were turned away exactly once. Then he went home, waited six months on purpose, and mailed every one of those businesses a question.

🧠 251 Stops, One Refusal

The travel notes are almost boring, which is the point. Some hosts looked curious. A few were warmer than usual. One establishment said no. LaPiere’s own reading of the pattern was that the friendliness had nothing to do with race and everything to do with presentation: the couple was neatly dressed, they smiled, and a polite party of three at the front desk is not a difficult transaction for a business that needs the room filled.

Six months after the last stop, he sat down with the survey. One question, three boxes:

Will you accept members of the Chinese race in your establishment? Yes / No / Depends upon the circumstances

128 of them wrote back. 92 percent of the replies said No.

He also mailed the same questionnaire to a comparison group of hotels and restaurants they had never visited. The answers came back about the same, which was its own kind of finding: the people who had actually served this couple were not some unusually prejudiced subset of the industry. They were the industry.

🤔 Why the Letter Said No

The easy reading of this study is that people are hypocrites, or that surveys are worthless because everyone lies. LaPiere was after something sharper, and it is the reason the study is still taught.

The two events were not the same question. The letter asked about a category — a named group, appearing in the abstract, with no face and no date attached. The doorway asked about a situation: two specific people, well dressed, a white man standing with them, a table that needed filling, tonight. When a person evaluates a category, what comes online is norms, upbringing, the going social consensus of the place. When a person handles a concrete situation, what comes online is what is visibly in front of them. Both answers were honest. They were just answers to different questions, given by different parts of the same person.

That reframing turned the discrepancy from a moral failure into a measurement failure, and it detonated the field. Wicker’s 1969 review of the attitude literature concluded that the link between what people say and what they do was so thin that the concept of an “attitude” might not be worth keeping. Decades later, Kraus ran the meta-analysis that put the pieces back together: 88 attitude-behavior studies, attitudes significantly predicting future behavior, mean r = .38. Not nothing, not destiny. The conditions mattered more than the headline — the relationship got much stronger when the attitude was stable, easily recalled, and built on direct experience, and stronger again when the way you measure the attitude matches the behavior in specificity.

Read that last condition against LaPiere’s letter. Ask whether someone would host Chinese guests at their hotel next Tuesday and you are closer to predicting Tuesday. Ask about “members of the Chinese race” and you have asked a different question entirely. That is also where Fishbein and Ajzen’s theory of planned behavior came from: stop predicting behavior from attitudes, predict intention first, then break intention into attitude toward the act, what other people expect, and whether the person believes they can pull it off.

Academia eventually named the underlying mistake. Using what people say as evidence of what they do is the attitudinal fallacy. The example sentences in the literature are unkind: Americans report attending religious services about twice as often as attendance records show, while European self-reports match the records closely. Employers say they would interview formerly incarcerated young Black men and then don’t. People in food studies under-report what they eat. Subjects in bystander-effect research claim a far higher willingness to intervene than onlookers show in real emergencies.

The honest caveats matter here too, because this is a landmark and also a mess. All 251 observations came from one traveler with no coding scheme and no second observer. Only 128 of the businesses replied, and nobody knows about the rest. The question asked about a category while the visit involved a specific couple with a white escort. And LaPiere’s own writing carries the racial assumptions of his time, including his habit of reading curiosity as approval. The study is a beautifully clear demonstration wrapped in one man’s field notes.

🔗 What Asking Cannot Measure

Every survey that asks people about their behavior is that letter. In product and technology research the split shows up constantly: a study measuring whether someone would use a tool is attitude data, and a retention graph is behavior data. They will disagree, and when they disagree the graph wins. The failure mode is quiet, because intention surveys are cheap, fast and quotable, and it is very easy to write up a study about adoption on top of data that measured enthusiasm.

LaPiere’s real methodological gift is the order he chose: watch first, ask second, and ask only after a delay long enough that people can no longer reconstruct the specific occasion. Behavior first, questions second. Most research runs it backwards.

The same gap organizes fiction. A character’s stated values are questionnaire data. The same character, five chapters later, under pressure, with something to lose, is the door — and the space between those two things is where a person exists on the page. A protagonist who behaves exactly as their principles predict is a description wearing a name. The gap is not a flaw to be smoothed over. It is the material.

And it applies to any system designed to model a human being. A personality written into a description document is self-report. What the thing shows when a specific situation arrives with an emotional cost attached is behavior. Design documents and evaluations that only ever collect the first kind are, in LaPiere’s terms, mailing letters and calling them evidence.

🎲 The Two-Week Test

You can run this on yourself, and it is genuinely uncomfortable.

Write down one thing you believe about how you live — the kind of sentence that goes in a journal. Then, without looking at the past, predict what you will actually do the next time a specific situation tests it: a specific day, a specific person, a named decision. Write the prediction down with a date on it. Then record what happened.

The literature’s guidance is to keep the two measurements at the same altitude. “I value my health” is a category question. “I will walk on Tuesday evening after the meeting ends” is a doorway question. Only the second one can be checked.

There is one more small historical joke worth having. Richard LaPiere spent the rest of his career writing sociology and novels — he won a California Book Award silver medal in 1941 for fiction — and he died in 1986. Stanford’s sociology department still gives out an annual award named after him for the best graduate student paper. It is decided by a committee of readers, not by asking the students how good they think they are. The prize named for the man who proved that self-report is unreliable is the one award in the department that refuses to rely on it.