History of Randomised Experiments
Experiments are the engine of behavioral economics. Understanding how a well-designed experiment isolates cause and effect—and how poor design can mislead—requires knowing where the method came from. The history of randomized experiments spans medicine, social policy, and economics, and each domain contributed key design principles.
1. Medical Origins: From Lemons to the Polio Vaccine
The earliest controlled trial is credited to James Lind (1774). To test whether lemons cure scurvy, he recruited eight sailors and split them into two groups: four received lemons (treatment), four did not (control). Although small and non‑random in the modern sense, it established the idea of comparing treated and untreated groups.
Ronald Fisher (1920s–30s) introduced explicit randomization as the foundation for causal inference, arguing that random assignment eliminates systematic bias. His two books—Statistical Methods for Research Workers and The Design of Experiments—formalised this logic.
Why randomization matters: Without it, treatment and control groups may differ in ways that affect the outcome, making it impossible to attribute differences solely to the treatment.
- 1948 – Streptomycin trial (Austin B. Hill et al.): A true randomised controlled trial (RCT) to test streptomycin against tuberculosis. Limited supply and ethical concerns motivated the trial: rather than giving the drug to all patients (which would waste scarce doses if ineffective), randomisation allowed fair evaluation. This became the blueprint for modern medical RCTs.
- 1954 – Salk polio vaccine (Jonas Salk): A double‑blind RCT (neither patients nor doctors knew who received vaccine or placebo). Salk used Fisher’s exact test to analyse the results, which showed a clear reduction in polio risk.
2. Social Experiments: Health Insurance and Cash Transfers
Randomised trials moved into social policy in the 1970s, testing large‑scale government programs.
RAND Health Insurance Experiment (1971–1982)
- Question: Does free healthcare lead to over‑use and excessive spending?
- Design: 7,700 participants (age < 65) assigned to three treatment conditions:
- Free care (status quo)
- Plans with varying degrees of cost‑sharing (copay)
- Health Maintenance Organization (HMO) plan (restricted network of doctors)
- Results: Cost‑sharing reduced physician visits, dental visits, hospitalizations, and overall spending. It also reduced initiation of care (people avoided seeing a doctor). The policy implication: copay can curb demand.
Negative Income Tax Experiments (1970s, US and Canada)
- Question: Does providing cash transfers to the poor reduce their willingness to work?
- Design: Five separate social experiments testing a negative income tax (subsidy that tops up earnings).
- Results: No effect on individual labour supply, but an unintended increase in divorce rates.
- Critique (Josh Angrist): Too many treatment arms made the design underpowered to detect credible effects.
| Experiment | Years | Sample | Key Finding | Design Issue |
|---|---|---|---|---|
| RAND Health Insurance | 1971–82 | 7,700 | Cost‑sharing reduces utilisation and spending | — |
| Negative Income Tax | 1970s | Multiple sites | No work disincentive; divorce rate rose | Underpowered, many arms |
3. Lab Experiments in Economics
While social experiments tested policy, lab experiments probed individual and strategic decision‑making. Early work focused on indifference curves, utility, and social dilemmas.
| Researcher(s) | Year | Experiment / Contribution |
|---|---|---|
| Thurstone | 1931 | First lab experiment in economics: used choices between goods (hats, shoes, coats) to deduce indifference curves. |
| Mosteller & Nogee | 1951 | Lab experiment measuring the subjective value of additional money income (restricted setting). |
| Von Neumann & Morgenstern | 1944 (book) | Proposed four uses of lab experiments: (a) let subjects take/refuse gambles with real money, (b) construct utility curves from behaviour, (c) make predictions about future choices, (d) test those predictions with more complex tasks. |
| Dresher & Flood (RAND) | 1950 | Implemented the first Prisoner’s Dilemma game in the lab. Participants played 100 repetitions. The Nash equilibrium (row2, column1) yields lower joint payoff than the social optimum (row1, column2). Empirically, subjects reached neither; their behaviour deviated from rational prediction, sparking decades of research on cooperation. |
| Guth et al. | 1982 | Introduced the Ultimatum Game: a bargaining game between a proposer (divides a sum) and a receiver (accepts or rejects). The receiver’s ability to reject unfair offers (even at own cost) violates the rational self‑interest prediction. After the dictator game, it is the most widely used experimental paradigm in economics. |
Exam tip: The RAND Health Insurance experiment is a classic example of a field experiment with clear policy implications. The negative income tax experiment illustrates the dangers of too many treatment arms—a design flaw that reduces statistical power.
How the Pieces Fit Together
Key takeaways
- Randomised experiments emerged from medicine (Lind, Fisher, Hill, Salk) and were later adopted in social policy and economics.
- Fisher formalised randomisation as essential for clean causal inference.
- Social experiments (RAND, Negative Income Tax) tested real‑world policies; the latter showed the risk of underpowered designs and unintended consequences.
- Lab experiments in economics began with Thurstone (indifference curves) and expanded to strategic games (Prisoner’s Dilemma, Ultimatum Game) that revealed systematic deviations from rational choice.
- The Prisoner’s Dilemma lab results directly challenged the Nash equilibrium prediction and spurred the field of behavioural economics.
Behavioral vs. Experimental Economics
Behavioral economics is a sub-discipline that uses experiments as a method. Experimental economics is the broader field of designing and running experiments to test economic theories. Experiments are tools applicable beyond behavioral economics—to psychology, market design, and other disciplines.
What is an Experiment?
An experiment creates identical environments and changes one thing at a time to measure causal inference.
An experiment is a way through which you develop identical environments, but change one thing at a time to measure causal inference.
Advantages of Experiments
- Control over data collection – Allows precise measurement of variables consistent with psychological factors that are unobservable in observational data (e.g., beliefs, preferences, norms).
- Clean causal identification – The researcher manipulates the treatment, enabling clear inference of cause and effect.
Typology of Experiments
Experiments can be classified into four main types, each with trade-offs in control, generalizability, and cost.
Lab Experiments
Lab experiments take place in a controlled laboratory environment, typically with student participants.
Market experiments – Test price discovery and market mechanisms.
- Example: Plott & Pogorelskiy (2017) – how quickly prices converge to fundamental values.
- Auction mechanism tests (Kagel & Levin, 1993) – compare trading structures:
| Auction Type | Bidding Rule | Winner Pays |
|---|---|---|
| First Price Sealed Bid (FPA) | Highest bid wins | Own bid |
| Second Price Sealed Bid (SPA) | Highest bid wins | Second-highest bid |
| Third Price Sealed Bid (TPA) | Highest bid wins | Third-highest bid |
Outcomes measured: bid values and revenue (related to the revenue equivalence theorem).
Behavioral experiments – Test heuristics, biases, risk, time, and social preferences.
- Linda problem (heuristics), Prospect Theory gambles, present bias discounting, social preferences (altruism, reciprocity, trust, deception).
Example: Corruption and Trust (Banerjee, Journal of Public Economics) Correlational data shows a negative relationship between corruption and trust across countries. To test causality, a lab experiment with three treatments was designed:
- Baseline – play the trust game only.
- Treatment 1 – play a bribery game (corruption), then a trust game.
- Treatment 2 – play a strategically equivalent ultimatum game (no corruption), then a trust game. Result: corruption caused a decrease in social trust (measured by trust-game outcomes). A subsequent experiment measured social norms.
Ultimatum game – A staple lab experiment.
- Proposer splits a pie (e.g., ₹10): offers (0–10).
- Responder accepts (gets , proposer gets ) or rejects (both get 0).
- Subgame perfect Nash equilibrium: proposer offers tiny amount, responder accepts any positive offer.
- Lab evidence: average offer ≈ ₹4 (40%); 40–50% of proposers offer 50-50; offers below ₹2 (20%) rejected.
- Why reject? Dislike of unfair outcomes – a psychological motivation not observable in field data.
Advantages of lab experiments:
- Full control over data collection.
- Clean identification (treatment manipulated by researcher).
- Creative, small changes possible.
Disadvantages:
- Subject pool: typically university students (WEIRD: Western, Educated, Industrialized, Rich, Democratic).
- Low external validity / generalizability.
- Incentives are small (but sufficient for students).
Lab-in-the-Field (Artefactual Field) Experiments
To address the WEIRD problem, experiments are moved to the field – a lab constructed in a real-world setting (e.g., a village panchayat building). Participants are non-standard (e.g., villagers, farmers).
Examples:
- Affirmative action and competitive choice across caste categories – treatments vary cost-based affirmative action policies, comparing decision-making of participants from different caste groups.
- Farmers’ behaviour before vs. after harvest.
- Teacher expectations: compare teachers’ expectations about students’ ability (by caste/gender) with actual performance – requires setting up labs in schools, eliciting expectations, and linking to outcomes.
Advantages:
- Greater generalizability than lab experiments.
- Still maintain control relative to pure field experiments.
Disadvantages:
- Less generalizable than large-scale field experiments.
- More expensive than lab experiments (need to set up “paraphernalia”).
Online Experiments
Conducted on platforms such as Amazon MTurk, Prolific, Qualtrics, TGM. Participants complete tasks online.
Applications:
- Expectations about the future economy.
- Role of narratives.
- Subjective models of the economy among people and experts.
Advantages:
- Low cost – fraction of lab experiment cost.
- Large-scale, participants from different parts of the country.
- Can achieve quota-based representativeness (e.g., 50% men, 50% women) – not a true representative sample but captures key demographic characteristics.
Disadvantages:
- Data credibility concerns – AI tools / LLMs may impersonate human participants.
- Time constraints – participants unlikely to spend >20–30 minutes; limits complexity and control.
Survey Experiments
A special type of online experiment where a treatment condition is embedded within a survey. Usually unincentivized, enabling very large samples.
Example (Stefanie Stantcheva, Nathan Nunn, et al.):
- Study how people perceive and form attitudes toward economic policies (e.g., gains from trade, tariffs).
- Manipulate salience of gains/losses from trade to examine impact on policy support.
- Recent paper on Zero-Sum Thinking used 20,000 US residents.
Advantages: Very large scale, can study attitudes and perceptions directly.
Disadvantages: Unincentivized – responses may be less reliable than incentivized choices.
Key Takeaways
- Experiments are defined by constructing identical environments and varying one factor to infer causality. They offer control and clean identification.
- Lab experiments provide maximum control but suffer from WEIRD subject pools and low external validity. They remain useful for testing theory (e.g., auction mechanisms, social preferences).
- Lab-in-the-field experiments improve generalizability by recruiting non-student participants in natural settings, at higher cost.
- Online experiments are cheap, large-scale, and can achieve quota representativeness, but face data quality risks and time constraints.
- Survey experiments embed treatments in surveys, enabling massive samples to study attitudes (e.g., trade policy, zero-sum thinking) without monetary incentives.
- The ultimatum game robustly demonstrates that fairness concerns (rejection of low offers) cannot be explained by standard equilibrium – a key result from lab experiments.
- When choosing an experiment type, trade off control, generalizability, cost, and credibility of data.
Field Experiments and Randomized Control Trials
Field experiments and randomized controlled trials (RCTs) test hypotheses by observing outcomes in naturally occurring environments – students in schools, patients in hospitals, labourers in factories. Participants operate in their natural habitat, often unaware they are part of a randomized experiment.
Distinction: Field Experiments vs RCTs
While the terms are sometimes used interchangeably, the key difference lies in purpose:
| Feature | Field Experiment | Randomized Controlled Trial (RCT) |
|---|---|---|
| Aim | Tests a precise theoretical prediction | Tests a hypothesis with direct policy relevance |
| Example | Does social pressure drive charitable giving? | Do treated bed nets reduce malaria? |
Both use random assignment in the field, but RCTs are explicitly designed to inform policy.
Examples
-
John List et al. – altruism and social pressure
- Competing motivations: pure altruism vs. dislike of saying “no”.
- Design: control (door-to-door campaign), treatment (pre-announcement that a door-to-door campaign will occur).
- Result: people avoid being home when they know a solicitor will come → the dislike of saying no is a major driver of giving.
-
Bertrand & Mullainathan – racial discrimination in hiring
- Sent résumés with stereotypically white names (Emily, Greg) vs. African-American names (Lakisha, Jamal).
- Result: callback rate for interviews differed by race → evidence of discrimination.
-
Kremer & Miguel – treated bed nets (RCT)
- Tested effect of insecticide-treated bed nets on malaria, income, child mortality.
- Policy relevance: directly informs public health interventions.
- Michael Kremer (with Banerjee and Duflo) won the 2019 Nobel Prize for using RCTs in development economics.
Advantages
- External validity – results are generalizable because they come from real-world settings.
- No Hawthorne effect / social desirability bias – participants are unaware they are being studied, so behaviour is natural.
Disadvantages
- Expensive and logistically challenging – difficult for a single researcher to run.
- Lack of control – cannot isolate precise mechanisms as easily as in a lab.
- Hard to add extra treatments – cost limits the number of conditions; often relies on survey measures to infer mechanisms.
RCT in Policy: Banerjee, Cole, Duflo & Linden – Education in India
- Problem: how to raise learning levels in low-income countries.
- Two programs evaluated via RCT:
- Balsakhi program: young women tutors helped lagging students.
- Computer-assisted learning: computers to improve numeracy.
- Each program was compared to a control condition.
- Direct policy relevance: successful programs could be scaled up by government.
Exam tip: The key distinction between a field experiment and an RCT is not always sharp; exam questions may ask you to classify an example by its purpose (theory-testing vs. policy evaluation).
Key takeaways – Field experiments & RCTs
- Both involve random assignment in natural settings; RCTs have explicit policy focus.
- Advantages: high external validity, natural behaviour, no Hawthorne effect.
- Disadvantages: costly, less mechanistic insight, limited treatment flexibility.
- Classic examples: List (social pressure), Bertrand/Mullainathan (discrimination), Kremer/Miguel (bed nets), Banerjee et al. (education programs).
Non-Randomised Experiments
Experiments can be used without randomization – researchers carefully measure outcomes in situations where assignment is not under their control. These rely on natural or quasi-experimental variation.
Examples
-
Blouin & Mukand (2019) – Rwandan nation-building radio
- After the genocide, the government broadcast radio programs to promote inter-trust and harmony.
- Terrain determined which villages received the signal → non-random exposure.
- Researchers visited exposed and unexposed villages and measured social trust using a trust game.
-
Mani et al. – financial scarcity and cognitive bandwidth
- Studied farmers before harvest (poor) and after harvest (rich).
- No randomization – same farmers at two time points.
- Result: scarcity reduces cognitive performance.
-
Gneezy, List et al. – gender differences in competition
- Compared behaviour (competitiveness, self-confidence, risk) across matrilineal and patriarchal societies; these societies are similar in wealth but differ in women’s roles.
- Found that stereotypical gender differences reversed in matrilineal societies → evidence that nurture, not nature, drives these gaps.
- No randomization – natural variation across societies.
-
Babcock et al. (2017) – gender and low-promotability tasks
- Measured likelihood of volunteering for tasks with low career payoff.
- Non-randomized but carefully calibrated measurement.
-
Gautam Rao – intermixing in Delhi private schools
- A court mandate forced schools to admit students from lower socioeconomic status backgrounds.
- Compared students exposed to lower-SES peers vs. those not exposed.
- Used dictator games and other field experiments to measure effects on discriminatory preferences.
- Illustrates how experimental data can uncover mechanisms behind social change.
Purpose of Non-Randomised Experiments
- Allow careful measurement of otherwise unobservable outcomes (trust, cognitive load, discrimination).
- Complement other data sources to understand mechanisms behind societal phenomena.
- They do not allow the same causal inference as randomized experiments, but still provide valuable evidence.
Exam tip: Non-randomised experiments are often used when randomization is impossible or unethical. The key is to identify the source of variation (geography, timing, policy) and argue why it is plausible as a natural experiment.
Key takeaways – Non-randomised experiments
- Randomization is not a requirement for an experiment to be useful.
- Examples leverage natural variation (terrain, season, culture, court ruling).
- Weakness: weaker causal identification; strength: feasible in settings where random assignment is not possible.
- Valuable for uncovering mechanisms (e.g., Rao on social exposure, Gneezy on culture vs. gender).
Ethics in Research – Historical Context
Why ethics matter in experimental economics: experiments use human beings as subjects. Historical violations of basic human rights led to strict modern protocols. The core ethical principles are codified in the Belmont Report (1979), resting on three pillars: respect for persons, beneficence, and justice.
Historical Violations (Four Key Cases)
| Experiment | When | What happened | Why it was unethical |
|---|---|---|---|
| Nazi twin experiments (Josef Mengele) | 1943–45 | ~1500 sets of twins (children) subjected to eye-colour changes, amputations, deliberate infections under pretext of genetic research. | No consent; horrific physical and psychological harm; murder. |
| Stanford Prison Experiment (Zimbardo) | 1971 | 24 male students randomly assigned as prisoners or guards in a mock prison; stopped after 6 days because guards became sadistic, prisoners severely stressed. | Lack of protection from harm; no provision for withdrawal; profound psychological distress. |
| Tuskegee Syphilis Study | 1932–72 | 600 African American men in rural Alabama; 399 had syphilis. Penicillin (effective treatment available from 1947) was deliberately withheld. | Deception; exploitation of a vulnerable population; denial of life-saving treatment. |
| Milgram Obedience Experiment | 1961–62 | Participants were ordered to administer electric shocks (up to 440V) to "learners" for wrong answers. Most obeyed despite hearing screams. | Severe psychological stress; deception about the nature of the shocks; potential long-term harm. |
Exam tip: The Milgram study is often used to illustrate the tension between scientific knowledge and participant welfare – it produced valuable insights about authority but inflicted distress. Modern IRBs would likely reject it.
Modern Requirements (Derived from Past Violations)
- Informed consent – participants must formally agree after being told risks, procedures, and that participation is voluntary.
- Mandatory ethics training for all experimenters (e.g. CITI Program).
- Institutional Review Board (IRB) approval – the entire experimental design must be reviewed.
Core Ethical Principles – The Belmont Report (1979)
-
Respect for persons
- Treat participants as autonomous agents; they have the right to say yes or no.
- Protect those with diminished autonomy (children, cognitively impaired, prisoners).
- Informed decision-making must be documented in a consent form.
-
Beneficence
- Do no harm to participants.
- Maximise benefits while minimising risks.
- Benefit-risk assessment must be independently evaluated by the IRB.
- Researchers benefit (publications, career); participants bear the risk.
-
Justice
- Fair distribution of benefits and burdens.
- Equal treatment of all participants.
- No exploitation of vulnerable populations.
Elements of a Proper Informed Consent Form
- Purpose of the research
- Duration of the study
- Procedures involved
- Risks and discomforts
- Potential benefits
- Confidentiality measures – how data is anonymised and stored
- Compensation – essential in economic experiments
- Voluntary participation and right to withdraw at any time
- Contact information for questions
Special Considerations
- Minors – require both parental consent and child assent.
- Vulnerable populations (prisoners, pregnant women, cognitively impaired, economically/educationally disadvantaged) – extra layers of protection must be specified in the protocol.
- Deception in economics – generally not allowed. Giving objectively false information is prohibited by the discipline's dominant norm. If deception is unavoidable (rare), participants must be debriefed and the IRB informed. Deception reduces credibility in economics.
- Digital consent – online experiments use digital consent forms; consent is ongoing throughout the experiment.
Institutional Review Board (IRB) – Purpose & Process
Composition: Minimum 5 members – scientists, non-scientists, and community members – who evaluate protocols independently.
Review categories:
Most economics experiments fall into exempt or expedited; researchers should be prepared for a full review if the board deems necessary.
Submission requirements to an IRB:
- Detailed IRB form (methodology, research protocol)
- Consent form and recruitment materials
- Risk-benefit analysis (the researcher's own assessment)
- Data protection plan – who has access, cloud security, anonymisation/de-identification
- Clear policies on data retention and destruction after publication
Data Management
- Anonymise or de-identify data unless the research question requires identifiable data (justification needed).
- Securely store and password-protect data.
- Specify retention period (often a minimum of 3 years after study completion).
Ongoing Ethics Responsibility
- Report any adverse events immediately to the IRB.
- Annual review for studies lasting more than one year.
- Any protocol changes must be communicated to the board via a corrigendum.
- Researchers must maintain accurate records for at least 3 years and publish results responsibly.
Exam tip: The shift from historical violations to the Belmont principles is a classic exam question. Be able to list the three principles and how each maps to a modern requirement (e.g., respect for persons → informed consent; beneficence → risk-benefit analysis; justice → fair selection of participants).
Key takeaways
- Four landmark historical violations (Nazi twins, Stanford Prison, Tuskegee, Milgram) drove creation of modern ethical standards.
- The Belmont Report's three pillars: respect for persons, beneficence, justice.
- Informed consent must include purpose, duration, procedures, risks, benefits, confidentiality, compensation, and right to withdraw.
- IRB reviews studies in three categories: exempt, expedited, or full review.
- Deception is largely prohibited in economics; if used, debriefing is mandatory.
- Ethics is an ongoing responsibility – adverse event reporting, annual reviews, and careful data management are required.