📚 Introductory Econometrics with MORISTAT

A Practical Guide to Empirical Analysis

Naguib Lallmahomed · naglal@linux-mauritius.com · 2026-06-14


Chapter 12: Epilogue — The Responsible Use of Statistics and Econometrics

🎯 Learning Objectives

Upon completing this Epilogue, you will be able to:

"All Knowledge is, in final analysis, History.
All Sciences are, in the abstract, Mathematics.
All Judgments are, in their rationale, Statistics."

C.R. Rao (1920–2023)

12.1 Introduction: The Power and Peril of Numbers

Throughout this book, you have learned the technical tools of econometrics: OLS regression, hypothesis testing, diagnostic tests, panel data, time series, and even machine learning. These are powerful tools. But like any powerful tool, they can be used for good or for ill.

This Epilogue is about responsibility. It draws on three traditions:

Key Insight: The goal of this Epilogue is not to make you cynical about statistics. It is to make you vigilant — to empower you to ask the right questions and to demand integrity in empirical research.

12.2 Huff's Warning: How to Lie with Statistics

Darrell Huff's 1954 classic remains the most accessible introduction to statistical deception. Its lessons are timeless.

12.2.1 The Sample with Built-in Bias

Huff warns: “A sample is a sample, not the whole population.” Yet many statistical claims are based on biased or non-random samples.

Example: In the 1936 US presidential election, the Literary Digest polled 10 million people and predicted a landslide victory for Alf Landon over Franklin D. Roosevelt. The poll was wrong because the sample was drawn from telephone directories and automobile registrations — which in 1936 overwhelmingly favoured wealthier (and Republican) households.

Key Question to Ask: How was the sample selected? Is it truly random?

12.2.2 The Well-Chosen Average

Huff famously points out that there are three kinds of averages: mean, median, and mode. The choice of which one to report can dramatically alter the impression conveyed.

Example: A real estate agent tells you that the "average" income in a neighbourhood is $100,000. What she doesn't tell you is that this is the mean, pulled up by a few wealthy executives. The median — the income of the typical resident — might be only $60,000.

Key Question to Ask: Which average is being used? Is it the mean, median, or mode?

12.2.3 The Little Figures That Are Not There

Many statistical claims omit crucial information: sample size, standard error, confidence intervals, or p-values.

Key Question to Ask: What's missing? Is the sample size large enough? Are the standard errors reported?

12.2.4 The Gee-Whiz Graph

Graphs can be manipulated by changing the scale of the axes, starting the Y-axis above zero, or using 3D effects that exaggerate differences.

Example: A graph showing a small increase in profits can look dramatic if the Y-axis starts at 90% instead of 0%. The visual impression is one of explosive growth, even if the actual increase is only a few percentage points.

Key Question to Ask: Does the graph distort the data? Look at the axes carefully.

12.2.5 The Semiattached Figure

Huff warns about comparing apples and oranges — using one result to imply something else that doesn't logically follow.

Key Question to Ask: Is the comparison legitimate? Or is the subject being changed?

12.2.6 Post Hoc Rides Again

Huff's most famous lesson: Correlation does not imply causation.

Example: A study found that cancer was more common in milk-drinking regions. But milk-drinking regions also had longer life expectancies — and cancer is a disease of older age. The correlation was spurious.

Key Question to Ask: Does X really cause Y, or could there be a third factor (or reverse causation)?

12.2.7 Huff's Final Advice: How to Talk Back to a Statistic

Huff ends with five questions every critical reader should ask:

  1. Who says so? — What is the source? Is there a conflict of interest?
  2. How does he know? — What is the methodology? Is the sample adequate?
  3. What's missing? — Are standard errors, sample size, or other key details omitted?
  4. Did somebody change the subject? — Is the comparison legitimate?
  5. Does it make sense? — Does the conclusion pass the "common sense" test?

12.3 Espasa on Hendry's Methodology: A Defence Against Malpractice

Antoni Espasa's paper Avoiding Malpractice in Econometrics presents David Hendry's methodology as a systematic framework for rigorous empirical research.

12.3.1 The Problem: Traditional Modelling Can Be Misleading

Espasa argues that traditional econometric modelling — where researchers "specify" a model based on theory alone — can lead to malpractice: omitted variables, incorrect functional forms, and biased estimates.

12.3.2 Hendry's Solution: A Discovery Process

Hendry's methodology treats econometric modelling as a discovery process, not a one-step specification. The key principles are:

"The specification of econometric models are not known at the beginning and a modelling discovery process is required."

— Antoni Espasa, summarising Hendry's methodology

12.3.3 Why Hendry's Methodology Prevents Malpractice

Potential MalpracticeHendry's Safeguard
Omitting relevant variablesThe GUM includes all plausible variables
Ignoring outliers or breaksIndicator Saturation Estimation (ISE)
Data mining (p-hacking)Automated, systematic search with rigorous testing
OverfittingDiagnostic checks and parsimonious encompassing
Ignoring non-linearitiesNon-linear selection capabilities
Ignoring dynamic structureLag selection and equilibrium correction mechanisms

Key Insight: Hendry's methodology does not eliminate the need for economic theory. Rather, it nests theory within a rigorous empirical framework — allowing theory to inform the GUM while letting the data speak.

12.4 Franses on Ethics in Econometrics

Philip Hans Franses' Ethics in Econometrics (2024) provides a comprehensive guide to ethical research practice.

12.4.1 The Ethical Landscape

Franses argues that econometricians face unique ethical challenges because:

12.4.2 Specific Ethical Pitfalls

12.4.3 Franses' Ethical Guidelines

12.5 Synthesising the Lessons: A Code of Responsible Practice

Drawing on Huff, Hendry, and Franses, here is a Code of Responsible Practice for econometricians:

1. Be Skeptical — Especially of Your Own Results

2. Be Transparent

3. Use a Systematic Methodology

4. Embrace Uncertainty

5. Focus on Substance, Not Just Significance

12.6 Practical Exercises

Exercise 12.1: Spot the Lie

Consider the following claim: "A study of 1,000 people found that 90% of those who drank red wine daily lived to be over 90 years old."

  1. What questions would you ask about the sample?
  2. What questions would you ask about the methodology?
  3. Is there a possible confounding factor (a third variable)?
  4. Does correlation imply causation here?

Solution:

  • Sample questions: Who was sampled? Was it random? What was the sample size? Was there a control group?
  • Methodology questions: How was consumption measured? Was it self-reported? Were other factors (e.g., diet, exercise, wealth) controlled for?
  • Confounding factors: People who drink red wine daily may also be wealthier, have better healthcare, or have other healthy habits.
  • Causation: No — correlation does not imply causation. A controlled experiment would be needed.

Exercise 12.2: Average or Averages?

You are presented with the following data on salaries in a company: $25,000, $28,000, $30,000, $32,000, $35,000, $40,000, $45,000, $50,000, $1,000,000.

  1. Calculate the mean, median, and mode.
  2. Which average would a manager use to attract recruits? Why?
  3. Which average would a union representative use to argue for higher pay? Why?
  4. Which average best represents the "typical" salary?

Solution:

  • Mean: (25+28+30+32+35+40+45+50+1000)/9 = 1285/9 = $142,778
  • Median: The middle value (5th of 9) = $35,000
  • Mode: No repeated values = None
  • Manager: Would use the mean ($142,778) to make salaries look generous.
  • Union: Would use the median ($35,000) to show most workers earn much less.
  • Best representation: The median ($35,000) best represents the typical salary, as the mean is distorted by the outlier.

Exercise 12.3: The Misleading Graph

You are shown a bar chart where the Y-axis starts at 80% instead of 0%, making a small increase from 82% to 86% look like a dramatic rise.

  1. Why is this misleading?
  2. What should the graph look like to be honest?
  3. How would you redraw the graph to show the true change?

Solution:

  • Why misleading: Starting the Y-axis above zero exaggerates visual differences. A 4% increase looks like a 33% increase (from 82 to 86 is a 4.9% increase, but visually it's distorted).
  • Honest graph: The Y-axis should start at 0% to show the true proportion.
  • Redraw: Use a Y-axis from 0% to 100%. The bars will be much smaller and the change will be shown in proper perspective.

Exercise 12.4: Post Hoc Reasoning

A researcher finds that ice cream sales and shark attacks are positively correlated. She concludes that eating ice cream causes shark attacks.

  1. Why is this conclusion flawed?
  2. What is the likely confounding factor?
  3. How would you explain this to a non-statistician?

Solution:

  • Flaw: Correlation does not imply causation. The researcher has committed the post-hoc fallacy.
  • Confounding factor: Hot weather. Both ice cream sales and shark attacks increase in summer.
  • Explanation: "Ice cream sales and shark attacks both go up in summer because more people are swimming in the ocean. The ice cream doesn't cause the shark attacks — the common cause is the season."

Exercise 12.5: Applying Hendry's Methodology

You are modelling the effect of education on wages. You have data on 1,000 individuals with variables: education, experience, age, gender, and region. You suspect there may be outliers, structural breaks, or omitted variables.

  1. How would you apply Hendry's General-to-Specific approach?
  2. What would be the GUM (General Unrestricted Model)?
  3. How would you use Indicator Saturation Estimation (ISE) to detect outliers?
  4. What diagnostic tests would you run to ensure congruence?

Solution:

  • Gets approach: Start with a very general model that includes all plausible variables (education, experience, age, gender, region, interactions, polynomials). Then systematically simplify using hypothesis tests and diagnostic checks.
  • GUM: Wage = β₀ + β₁Education + β₂Experience + β₃Age + β₄Gender + β₅Region + β₆Education² + β₇Experience² + β₈Education×Experience + ... + u (including all possible variables and transformations).
  • ISE: Include impulse indicators for each observation to detect outliers. The algorithm will retain those that are statistically significant, revealing potential outliers or breaks.
  • Diagnostic tests: Test for heteroskedasticity (BP, White), serial correlation (DW), normality (JB), functional form (RESET), and multicollinearity (VIF).

12.7 Key Terms

Statistical Manipulation Statisticulation Sample Bias Mean Median Mode Standard Error P-Hacking HARKing Post-Hoc Fallacy General Unrestricted Model (GUM) General-to-Specific (Gets) Indicator Saturation Estimation (ISE) Autometrics Data Mining Cherry-Picking Reproducibility Transparency Ethical Guidelines

12.8 Further Reading

🔄 Ready to connect with MORISTAT?