Writing

Survey methodology,
worked properly

Long-form notes on the statistics behind survey research — the derivations, the simulations, and the places where standard software quietly gives you the wrong answer. Every empirical claim here is backed by code you can run.

Methodology

Sampling Weights in Significance Tests

Weights make your point estimates right. They do not make your standard errors right — and most survey software silently gets this wrong.

18 min read
Methodology

How Many of Your Banner's Significant Results Are Noise?

A 200-test banner produces about ten significant findings when nothing is going on. Correcting for that is easy; the hard part is that the correction costs almost all your power.

12 min read
Methodology

The Confidence Interval in Your Tables Is Probably Wrong

The textbook interval for a proportion achieves 33% coverage where it promises 95%. Here is the exact arithmetic, and the one-line fix that has been available since 1927.

11 min read
Weighting

Should You Trim Your Weights? Look at One Number First

Trimming is sold as a free reduction in variance. Whether it helps or hurts depends almost entirely on the correlation between the weight and the outcome — and you can check that in a line of code.

12 min read
Methodology

Your Tracker Is Throwing Away Half Its Power

If waves share respondents and you test them as independent samples, the test is valid but badly conservative — 1.1% type I error where you asked for 5%, and a quarter of the movements you could have detected go unreported.

10 min read
Methodology

Nearly Every NPS Confidence Interval Is Too Narrow

Promoters and detractors come from the same respondents, so their estimates are negatively correlated. Ignore that and your standard error is 16-26% too small, and your group comparisons reject twice as often as they should.

11 min read
Methodology

The "±3%" On Your Report Applies To One Number

A single margin of error is quoted once and read as though it governs every figure in the document. For a realistic subgroup on a low-incidence item, the real interval is more than twice as wide.

9 min read
Methodology

5,000 Interviews Across 120 Sampling Points Is Not 5,000 Interviews

A tiny intra-cluster correlation of 0.012 turns 5,000 interviews into an effective 3,200. The multiplier is the cluster size, and face-to-face designs have large clusters.

10 min read
Weighting

Your Response Rate Says Almost Nothing About Your Bias

A 70% response rate can carry a 5.8 point bias while a 15% rate carries none. The quantity that governs it is a covariance, and the response rate is only its denominator.

10 min read
AI & Research

How to Validate an LLM That Codes Your Open Ends

Raw agreement of 92.7% can mean a kappa of 0.90 or 0.65 depending on nothing but category balance. And validating on 50 items gives a confidence interval three times too wide to act on.

12 min read
Data Quality

SPSS Rounds Your Weighted Counts. Here's How Much That Actually Matters

Less than you have been told — about 0.09pp on a base of 400. But the rounding is a symptom of a much larger error sitting directly underneath it, and that one is worth 26%.

8 min read