What the Father of Eugenics Can Teach Us About Data Science


Ok, so that’s a super click bait-y title. But, as the kids on Reddit would say, hear me out.

Meme: Barbie Land is "The Dashboard" and Oppenheimer's mushroom clouds is "The Pipeline."

A couple of weeks ago, I had the pleasure of moderating an AVIXA Power Hour on the subject of “Decoding Your Data.” Darren Coleman and Matt Ford talked us through the principles of data science. I always learn something new when I host Power Hours, but this one was especially eye-opening for me.

Darren and Matt talked me through the process of a proper data project. You can’t just throw a bunch of numbers at the wall and see what sticks. You have to start with a plan for what you want to capture. What question(s) do we want answered? How do we quantify that? And then, when we’re done, how long do we keep that data before disposing of it?

(As a digital hoarder, that last one might be the hardest for me to put into practice).

In the age of LLMs, there’s a tendency towards capturing all of the data, because “we can just put it into a model.” The problem is, if you don’t start with a plan you don’t actually know what you’re measuring.

How to lie with data, a series of charts:

This is where I remind you that correlation โ‰  causation. In other words, just because two things happen at (roughly) the same rate, that doesn’t mean that they are directly related. A classic example of this is the chart of global average temperature plotted against the number of pirates in the world. Temperatures have increased as piracy has decreased. That doesn’t mean that solving climate change is a function of distributing Jolly Roger flags. In fact, a more recent update to said chart would show that, even as temperatures have continued to climb, piracy has seen a comeback.

Humans love to see patterns in things. It’s how we can distinguish predators when they’re hiding in the woods. But sometimes numbers are just numbers. As Darren and Matt reminded me, you have to design your data experiment(s) first.

A real-life example of where data had to be thought through is the chart of damage inflicted to WWII bombers. On the image of a plane, we see clusters of red dots showing bullet holes in planes that limped back to base. Statistician Abraham Wald mapped out areas where the planes needed to be reinforced, taking into account that he only had access to airplanes that made it home. If you plotted out the data without taking into account the data’s origin, you’d reinforce all the wrong areas. Those red dots show areas where a plane can be hit and still make it back. Presumably, the planes with damage in the non-dotted areas were all lying in pieces in the German countryside and at the bottom of the ocean.

Our final example shows where data can be more nefarious. Tyler Vigen has an excellent explainer, where he purposefully organized data to show a serious of spurious correlations. By playing with the scale of the y-axis and using line graphs, he’s able to make the correlations appear far stronger than the data supports. The correlations themselves are a result of throwing a bunch of data together to see what lines up. As Vigen writes,

ย I have 25,237 variables in my database. I compare all these variables against each other to find ones that randomly match up. That’s 636,906,169 correlation calculations! This is called โ€œdata dredging.” Instead of starting with a hypothesis and testing it, I instead tossed a bunch of data in a blender to see what correlations would shake out. Itโ€™s a dangerous way to go about analysis, because any sufficiently large dataset will yield strong correlations completely at random.

Unfortunately, people lie using charts like this all the time. Be wary of news sources that don’t label their x and y axes.


This brings us back to Francis Galton, the originator of Eugenics. I listened to a podcast about him right after our Power Hour and I was struck by what a perfect example he is of bad data science. Galton was a data nerd who pioneered statistical ideas that we still use to day. He was also a virulent classist and racist. In creating his loathsome theories, he relied on spurious data-gathering. He trumpeted whatever numbers backed up the story he wanted to tell and hand-waved away the rest.

One of Galton’s most lasting contributions to statistics was around the concepts of correlation and causation. Ironically, his awful ideas rest on showing the correlation between data sets. (He could never have showed causation, because his ideas were as wrong as they were evil). His “proof” that you could somehow breed humans like dogs rested on things like survey data that showed that “great men” often had “great men” as fathers. He was quite adept at, you know, ignoring the educational advantages in growing up with all of the resources and then having access to all of your dad’s friends and peers when looking for work and other opportunities. To say nothing of the biased selection process he used to choose his survey subjects.

In other words, he came up with a theory and then engineered the data to fit.

Also, garbage in = garbage out.

We could hand wave Galton away as a crank, if it weren’t for the inconvenient fact that his ideas were incredibly popular at the time. His most notorious fan might have been Hitler, but the people who applauded his work ranged from Winston Churchill to Alexander Graham Bell and Teddy Roosevelt.

Galton was always going to lie with numbers, because he had an agenda. But we can do better. From what I learned in the Power Hour, if you want to use data to make better decisions, you need to do the following:

  1. Figure out what problem you’re trying to solve
  2. Choose exactly what data you’re going to collect
  3. Have a plan for keeping that data safe from prying eyes. Nobody should have access to more data than they need to do their work
  4. Analyze your data using the methods that you set forth at the beginning of your project
  5. Dispose of your data according to your data retention plan

We can’t always count on our better natures or the inherent “truth” in data to guide us. Be honest about your biases, design your experiments thoughtfully, and always think about where your numbers are coming from.

If you missed the Power Hour, you can watch the recording here: https://www.avixa.org/events/webinars/enterprise-it-power-hour-decoding-your-data-september-2026

Leave a Reply

Your email address will not be published. Required fields are marked *