Writing
Survey methodology,
worked properly
Long-form notes on the statistics behind survey research — the derivations, the simulations, and the places where standard software quietly gives you the wrong answer. Every empirical claim here is backed by code you can run.
Sampling Weights in Significance Tests
Weights make your point estimates right. They do not make your standard errors right — and most survey software silently gets this wrong.
How Many of Your Banner's Significant Results Are Noise?
A 200-test banner produces about ten significant findings when nothing is going on. Correcting for that is easy; the hard part is that the correction costs almost all your power.
The Confidence Interval in Your Tables Is Probably Wrong
The textbook interval for a proportion achieves 33% coverage where it promises 95%. Here is the exact arithmetic, and the one-line fix that has been available since 1927.
Should You Trim Your Weights? Look at One Number First
Trimming is sold as a free reduction in variance. Whether it helps or hurts depends almost entirely on the correlation between the weight and the outcome — and you can check that in a line of code.
Your Tracker Is Throwing Away Half Its Power
If waves share respondents and you test them as independent samples, the test is valid but badly conservative — 1.1% type I error where you asked for 5%, and a quarter of the movements you could have detected go unreported.
Nearly Every NPS Confidence Interval Is Too Narrow
Promoters and detractors come from the same respondents, so their estimates are negatively correlated. Ignore that and your standard error is 16-26% too small, and your group comparisons reject twice as often as they should.
The "±3%" On Your Report Applies To One Number
A single margin of error is quoted once and read as though it governs every figure in the document. For a realistic subgroup on a low-incidence item, the real interval is more than twice as wide.
5,000 Interviews Across 120 Sampling Points Is Not 5,000 Interviews
A tiny intra-cluster correlation of 0.012 turns 5,000 interviews into an effective 3,200. The multiplier is the cluster size, and face-to-face designs have large clusters.
Your Response Rate Says Almost Nothing About Your Bias
A 70% response rate can carry a 5.8 point bias while a 15% rate carries none. The quantity that governs it is a covariance, and the response rate is only its denominator.
How to Validate an LLM That Codes Your Open Ends
Raw agreement of 92.7% can mean a kappa of 0.90 or 0.65 depending on nothing but category balance. And validating on 50 items gives a confidence interval three times too wide to act on.
SPSS Rounds Your Weighted Counts. Here's How Much That Actually Matters
Less than you have been told — about 0.09pp on a base of 400. But the rounding is a symptom of a much larger error sitting directly underneath it, and that one is worth 26%.