Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Thursday, June 12, 2014

Are School Shootings Becoming the New Norm?

It's a familiar story: a string of low-probability/high-consequence events capture the front pages and they become the subject of public discourse, including plenty of editorials. Once it was airplane hijackings or serial killings. Now, it's school shootings. We need to do something to stop this trend of horrible events. In most cases, it's beyond debate whether or not the events are horrible. But are they indicative of trends (i.e., an increase over the prior frequency level of similar horrible events)?

Recently, a statistic has been making the rounds on blogs, social media, and in the mainstream news: 74 school shootings have occurred since the shootings at Sandy Hook in late 2012. The statistic comes from a group that has a rather explicit agenda: Everytown for Gun Safety. While it does give us some information about an important social phenomenon, my initial feeling about this statistic is that when taken out of context (and it is almost always shown without any context), it is apt to mislead.

First, there's the question of how we define "similar events". Do we count incidents in which guns were discharged but didn't kill anyone, or injure anyone? Do we count homicides related to drug deals? Do we count suicides? Do we count accidental gun discharges? Do we count colleges as well as elementary and secondary schools? In the case of the above statistic, Everytown for Gun Safety has counted all of these as "school shootings".

Why does it matter how inclusive our definition of "school shootings" is? Aren't all of these shootings horrible events that we should seek to avoid? In my opinion, yes, absolutely. However, if we're trying to understand a certain social phenomenon and whether a string of recent events are part of a trend, broad definitions only muddy the water. The social and psychological processes involved in suicide-by-gun at or near schools is likely to be different in important ways than those processes as they are applied to individuals like Adam Lanza and Elliot Rodger. Obviously the common denominators are "school" and "gun", and if you're of the mind that gun control is the ONLY solution to the problem of any kind of school shooting, then you may not care about the differences. But you should. Even if you're in favor of gun control and you think it will cut down on the number of people who die each year in school shootings, it won't help to ignore other factors, and I think overly-broad definitions of "school shootings" encourage this kind of ignorance. Is the increase in school shootings due to an increase in angry-young-males lashing out at the world, an increase in drug/gang-related shootings, or both? In order to address the issue, it's important to know.

Is it really a trend? If you just tell someone how many incidents occurred in one year, that doesn't really give them a good idea of whether or not it is part of a potentially alarming trend. This is the real reason I'm writing about this. I was genuinely curious. I wanted to know if the headlines were part of a familiar kind of hysteria about a random rash of low-probability/high-consequence events or if they're on to something. It wasn't easy to tell just by reading the news or the blogs. But there IS data on this. It ain't perfect data, but I think it may help me get an answer, this is one of the things I absolutely love about today's Internet - you have access to data and can see for yourself whether there is evidence to support a conclusion.

I used this list from Wikipedia of school shootings in the U.S., defined as "incidents in which a firearm was discharged at a school infrastructure, including incidents of shootings on a school bus." It included K-12 schools as well as colleges and universities, so the definition is pretty similar to that used by Everytown for Gun Safety. The list draws from newspaper archives. By virtue of their newsworthiness, school shootings seem unlikely to have ever gone unreported, so I think we can safely assume that this list is a fairly accurate record of school shootings and doesn't contain much in the way of systematic measurement error. 

Before I show you what I found, take a moment to guess: how many school shootings do you think there were in 1880? How about 1904? 1970?

(scroll down for the answers!)














Here's what I found:



So yeah, it does seem like there's something crazy going on in the last 18 months if you define school shooting this way. Why do I designate the last 18 months as the time period to pay attention to? It's about the time that the Adam Lanza killings occurred, and it's also near the beginning of the year, but these aren't good reasons to use this date as a cut-off. But if you look at the number of shootings per month, they really seem to go up in January of 2013.

There is plenty of seemingly random year-to-year variation across the whole 164 year period, but the number of incidents tend to fluctuate between 0 and 5 per year. It's worth noting that the last year there were no school shootings was 1981, so you could use that as a the beginning of the trend if you really wanted to.

It was weird for me to read about children shooting themselves in the early 1900's at school. It doesn't fit with my conception of the culture at the time. So school shootings aren't unprecedented, but after looking at this data, it seems hard to deny that the problem has gotten much worse in the past few years (or possibly since 1995 or 2005). If there were, say, 10 school shootings in 2013, I think you could possibly say that it was due to random variation. But there have been 31 in the first half of 2014!

What if we dig a bit deeper and look at the circumstances around each shooting. Do we see differences between the shootings from, say, the early 1900's and the shootings of today? In my search through the shootings since the 1850's in schools, I tried to isolate ones that I thought were similar to Sandy Hook: not suicide; not accidental; and not directed at one particular individual for reasons of, say, revenge or a rejection; not gang-related. I call this kind of shooting "mass school shootings". I include cases in which an individual clearly attempted a mass school shooting but was thwarted (though there are only a few of these, so they don't make much of a difference). Again, I must emphasize: I'm not saying that those other kinds of shootings aren't problems that need to be solved, but only that I want to isolate Sandy-Hook-ish shootings to see if there is indeed a trend. Here's what I found:




Note that the Y axis is different than the Y axis in the first graph. If I had used the same Y axis in both graphs (or included both lines in the same graph), it would've been hard to see the mass shootings line, so this is why I used two different y axes.

If you look at most of the school shootings pre-1966, they're directed at a particular individual and the motives are typically revenge (for a bad grade or being rebuffed by a would-be lover). Again, it was weird to read about mass school shootings similar to Sandy Hook that took place in the 1800's (though there were just two, so it really was anomalous).

As you can see, there's a similar pattern to the pattern in the overall school shootings: not much happens until recently. But what do we mean by "recently"? You could easily put the starting point of the trend at 1985, the first year there were more than two separate incidents of mass school shootings. What about since Sandy Hook? Let's take a closer look at the data, starting at 1985:



So, 1985 was a bad year, and so were 1988, 1998 (the 90's in general were pretty bad for this type of thing) and 2006. 2013 (the only post-Sandy-Hook year in our data set) is bad, but not much worse than these other years. In the first half of 2014, there have been three mass school shootings, but it would be dangerous to extrapolate from that and estimate that there will be six for the year. Extrapolation with such small sample sizes rarely maps on to reality.

In the end, here's what I get from this little exercise. The truth, such as I can determine it from available evidence, is that school shootings in general have become a lot more common in the last 18 months, but it seems unwarranted to say that Sandy-Hook-style mass school shootings have become a lot more common in that same time period. If you wanted to identify a starting point for the cultural phenomenon of mass school shootings, you'd be better off going back to 1985. If I'm looking for a culprit or cause of this phenomenon, I wouldn't look at things that are going on right now in our culture and in our laws. I would look at things that have been around and haven't changed much since 1985.

As would be the case with any phenomenon, my attitude would change if new evidence warranted such a change. If there were no more school shootings this year, I'd stick to my interpretation: nothing new in the world of mass school shootings since '85. BUT if there were 3 or more, then I'd reconsider my outlook.

From looking at this data, I also get the impression that there is an urgent problem with shootings at schools, but that this problem isn't akin to the problem of Sandy Hook or Columbine. Like any compassionate person, I'm appalled at that rapid increase in school shootings you see in the first graph, and I want to know more about why it's occurring and what we can do to stop it. A lot of these school shootings are committed by sane people who have a deep disagreement with another person and (importantly) access to guns. If anything, I think this analysis makes the case of those who choose to pin the problem of school shootings on a lack of proper mental health care and not on gun availability (I'm looking at you, Wayne LaPierre!) a lot weaker. School shootings are skyrocketing, according to the evidence, and most incidents don't involve mental health issues as such. So, maybe addressing the issues of gun availability, conflict resolution, and, perhaps, a cultural component may be an effective way to lessen the number of overall school shootings.

This analysis does NOT make for an easily conveyed, pithy soundbite. Because they need something pithy, the news and bloggers and others have latched on to a pithy-put-misleading alternative - the "74 since Sandy Hook" stat. On the one hand, this stat may grab more people's attention, gets them to click and gets them to post. On the other hand, based on all the evidence I can see, it IS misleading. It's use opens up those trying to convince others of the severity of the issue to attacks based on their use of misleading statistics.

This is, I would say, a familiar story in the use of statistics related to emotionally-loaded, low-probability/high-consequence events. It's a story that's worth returning to when discussing media literacy.

One final note: in my search for information about this topic, I found this article on CNN that actually DID dig a little deeper and found results similar to my own. They found a few more "Sandy-Hook-like" shootings because their definition differed slightly from mine, and (importantly) they had no historical comparison so they can't really talk about trends, but still, it gave me hope for more context when reporting stats in news. Big ups to CNN (though you really need to stop with the auto-playing videos).

Data source: http://en.wikipedia.org/wiki/List_of_school_shootings_in_the_United_States#cite_ref-47

Monday, April 07, 2014

(Mis)Understanding Studies

Nate Silver and his merry band of data journalists recently re-launched fivethirtyeight.com, a fantastic site that tries to communicate original analyses of data relating to science, politics, health, lifestyle, the Oscars, sports, and pretty much everything else. It's unsurprising that articles on the site receive a fair amount of criticism. In reading the comments on the articles, I was heartened to see people debate the proper way to explain the purpose of a t-test (we're a long way from the typical YouTube comments section), but a bit saddened that the tone of the comments made them seem more like carping and less like constructive criticism. Instead of saying someone is "dead wrong", why not make a suggestion as to how their work might be improved?

One article on the site got me thinking about a topic I've already been thinking about as I begin teaching classes on news literacy and information literacy: how news articles about research misrepresent findings and what to do about this phenomenon. The 538 piece is wonderfully specific and constructive about what to do. It provides a checklist that readers can quickly apply to the abstract of a scientific article, and advises readers to take into account this information, along with their initial gut reaction to the claims, when deciding whether or not to believe the claims, act on them, or share the information. It applies to health news articles in the popular press, but I think it could be applied to articles about media effects.

Now, the list might not be exhaustive, and there might be totally valid findings that don't possess any of the criteria on the list, but I think this is a good start. And really, that's what I love about 538. I recognize it has flaws, but it is a much needed step away from groundless speculations based on anecdotes that are geared toward confirming the biases of their niche audience (i.e., lots of news articles and commentary). And they appear to be open to criticism. Through that, I hope, they will refine their pieces to develop something that will really help improve the information literacy of the public.

The piece got me thinking about the systematic nature of the ways in which the popular press misleads the public about scientific findings. They tend to follow a particular script: The researchers account for most likely contributors to an outcome in their studies and test these hypotheses in a more-or-less rigorous fashion. The popular press does not mention the fact that they accounted for certain possible contributing factors because of limited space and the need to attract a large, general audience. When people read the news article about the research study, they think "well, there's clearly another explanation for the finding!" But in most (not all, but most) cases, researchers have already accounted for whatever variable you imagine is affecting the outcome.

In other cases, the popular press simply overstates either the certainty that we should have about a finding or the magnitude of the effect of one thing on another thing. Again, if we look at a few things from the original research article (like the abstract and the discussion section), we should be able to know whether or not the popular press article was being misleading, and we wouldn't even have to know any stats to do this. 

The popular press benefits from articles and headlines that catch our eyes and confirm our biases. That's just the nature of the beast. Instead of just throwing out the abundant information around us, it's worth developing a system for quickly vetting it, and taking what we can from it. 

Friday, July 05, 2013

The Two Webs

There are two dialogs on human behavior (which includes political, economic, and social behavior) taking place on the web. In effect, there are two webs.

One web consists of data on human behavior and commentary about this data. This one connects some folks in academia to folks in policy circles and the private sector around the world. This is Big Data.

The advantages of this mode of inquiry is that it harnesses the power of new media technologies to provide more information to help improve our predictive power when trying to understand something as enormously complex as individual and collective human behavior. Many of the critiques of quantitative study of human behavior were grounded in the fact that studies simply didn't have enough information to predict and explain the variance in behavior. Whereas other sciences (physics, chemistry) had enough information about a system to predict outcomes within that system, social sciences did not. But if we were to assume, for a moment, that a team of researchers had access to every single bit of information about every human thought, feeling, or behavior for thousands of years, then that team's ability to predict human behavior would be comparable, I think, to those in other sciences. With better predictions come better answers to questions: how best to minimize suffering, or the spread of disease, or human's impact on the environment, or whatever.

Of course, this mode of inquiry is not without its flaws (or at least perceived flaws). The collection and analysis of so much data on human behavior is viewed as being exploitative in some way: those collecting the information benefit and those who are the subjects do not. There are privacy issues: privacy is seen as a prerequisite to mental and emotional health as well as a means of maintaining some power over determining the course of your life (the actual value of privacy would be difficult to determine within this purely quantified conversation about human behavior). There is also the fear that someone with enough information about human behavior will be able to manipulate people to suit their ends (but if those collecting and analyzing data discuss findings freely and don't hoard secrets, this critique doesn't make much sense to me). It can easily be abused by people thinking there's a causal relationship when then there is only a correlational one. Statistics, when misused, create the illusion of certainty. Statistics could always be misused, but the more powerful and widespread they become, the more likely misuse might be and the damage that could be done.

You might call this the rational web. It views behaviors and events as probablistic (or it should, anyway) and it takes into account the degree to which outcomes are affected. If, say, it was determined that people's political party identification determined the amount they were paid when controlling for occupation, abilities, etc., you could find an answer to the question of how much difference it made in terms of pay (maybe Republicans make $4000 more on average when controlling for other relevant variables). In fact, there is an expectation that you answer the "how much" question.

The other web consists of rhetoric: emotional appeals to pre-existing, deeply held beliefs about human behavior. The most commonly used technique here is to select a few emotionally charged stories and try to get the audience to empathize with them. As access to the web has increased, it has become easier to find a subject (that is, the individual at the center of the story) whose situation exemplifies the pre-existing beliefs about human behavior held by the author of the story and the intended audience. Its easier to find the emotionally charged stories, the one or 3 or 100 personal stories that, when you look at parts of them the right way, support your pre-existing belief that, say, capitalism or socialism is harmful or that a certain policy does more harm than good. This technique connects some other folks in academia with the public at large, particularly disenfranchise members of the public worldwide.

Here, I think one possible danger of the wide-spread use of this technique might be that the echo chamber effect (where certain factions become less able to take the perspective of others and become more hostile towards others) gets stronger. Confirmation bias runs amok, and fewer people take into account new information in order to make better decisions. Any holder of an opinion, no matter how wacky, can find others supporting their opinion. This social support, this sense that one is not alone in one's beliefs, is essential to the persistence or propagation of an idea or ideology. Even if one is in the minority, all one has to do is draw an analogy to a group that was in the minority that eventually became the majority (the rebelling colonists in early America or civil rights crusaders in the later 1950's) in order to justify one's beliefs.

You might call this the emotional web. Rather than being probablistic, it is principled. It rarely asks the question "how racist is a statement?" or "how much privacy is being sacrificed?" or "how much freedom is the right amount of freedom"? In this way, it seems irreconcilable with the rational web.

You can't easily categorize certain websites as one or the other. Two of my favorite news sites - the New York Times and Slate - have some stories that appeal to statistics and analyses of statistics and other stories (usually editorials) that appeal to emotion by cherry-picking individual stories. In fact, a journalistic standard seems to be to combine the two: start with an individual's story and then zoom out to the larger trend. Hook the audience with emotion and convince their inner skeptic with data.

Still, I can see, at the very least, certain blogs that are more emotional or more rational, and it would be interesting to see if certain people gravitated toward either emotional/rhetoric arguments or rational/data-driven ones. Last month, I saw a great paper at the International Communication Association's annual conference by Brian Weeks titled "Partisan enclaves or diverse repertoires? A network approach to the political media environment" that suggested that the self-selection ideological bias (dems watch only MSNBC, repubs watch only Fox News) is a misconception and that personal media repertoires are more diverse, at least ideologically, than many believe. They may be diverse (or rather, balanced) in terms of their emotional or rational content as well. But maybe they are not, in which case we really are two different groups of people having two fundamentally different conversations about human behavior. Definitely an avenue worth exploring.

Tuesday, April 14, 2009

Data Mining as Psychotherapy: How the petabyte age could help us to know our selves and why that's nothing to be scared of


Let's start with a problem. Make it a personal problem. I mean "personal" in two senses: "personal" as in something that you wouldn't want to talk about with anyone besides a very close friend or a therapist, and "personal" as in specific to you and only you. Say you're depressed, or that you're engaging in what you know to be a self-destructive pattern of behavior. How can you solve this problem, or avoid repeating it in the future?

Now, let's pretend you had a magical machine that could track every bit of thought and experience you had from your birth to this moment. The data produced by the machine tells you everything that led up to that negative outcome moment. Well, not everything. Only the things that you were a part of. If a butterfly flapped its wings in China the day after you were born, the machine would not record that. It is possible that things that you were not directly a part of could have a profound effect on you, leading you to become depressed or engage in shitty behavior (see sensitive dependence on initial conditions). Nevertheless, if you have to limit your recording of information somehow due to technical limits, of which there will always be some, and your goal was to find out the cause of a problem that relates to you specifically, then it would be good to start with all experience and thought related directly to you.

Now that you've got all that information, you could look for patterns in that data, or maybe the machine could look for them for you. You notice a recurring pattern of actions or thoughts that you keep choosing that lead up to that negative outcome you want to avoid. We have an intuitive grasp of the kinds of behavior or thought that lead up to such outcomes (e.g. I'm depressed because I looked at a picture of a dead relative, which reminded me of how much i missed them and of my own mortality. Maybe I shouldn't look that picture so often). But sometimes, we can't see those patterns, either b/c we don't want to acknowledge that somethings that produce pleasure in the short-term might bring us displeasure in the long term, or b/c we just can't remember everything that we thought, felt, or did over the course of our entire lives.

The first problem is an objectivity/subjectivity problem, solved by asking a trusted friend or a therapist for advice. The second problem is a surveillance problem - no other person is there for our every waking moment, and even if they are, they can't see inside our heads (the closest you could come to that would be a parent or a lover). This creates a trade-off: you're the only person with access to your entire history and your thoughts and feelings, but you're not especially objective about them. The third problem is a cognitive load problem: no one person can hold that much information about one person's life in his/her head.

Enter the machine. The machine is capable of holding a much larger amount of information. The machine also makes the data available to you or to trained professionals (though, once you were able to see the patterns that linked the short-term pleasurable actions with long-term displeasurable consequences, you wouldn't need anyone else to tell you to cut it out). The machine could be with you every waking minute. Heck, it could even record what's going on in your head while you're asleep.

I'm not saying that the machine would allow you to see with complete certainty what led to that negative outcome, but it would be on a par with what we think to be causes and effects in the hard sciences. In other words, they would be able to reduce the probability of there being some other explanation for your misery to virtually nil. So, you take the data, you make note of the patterns of thought and behavior that led up to the negative outcomes, and you choose not to think or do the things that led up to those outcomes. Even better, you look at the patterns that lead up to your happiest moments, and you repeat those.

The machine does not exist. yet. But I think we've got a prototype: google. Google tracks our searches, which might tell us a little more about our patterns of behavior than we might know ourselves. A prototype also exists in the form of spyware that tracks our every click on the internet. As more and more of our desires and thoughts and feelings and actions are conveying on computers, the closer we come to having something like "the machine." (if you included mobile tech and its ability to track us throughout the day, you'd have an even better approximation of the machine).

Is the machine to be feared? I'd guess that most people would say yes, but I wouldn't agree. To me, that's like saying that you're afraid of knowing yourself, or afraid of knowledge in general. What we're afraid of is the misuse or misinterpretation of information. But should that keep us from garnering what we know to be more accurate information about our selves? There is such a thing as responsible data interpretation. In order to engage in responsible interpretation, it is essential to start with this assumption: the information we're dealing with is imperfect and incomplete, and yet it may offer us insight into our thoughts and actions that is superior to (or supplements) what we're currently working with. We need to engage in systematic testing of the circumstances in which this information does provide us with insight, and we need to identify misuse and misinterpretation and discourage it.

The other choice is one that I think too many people choose, mostly out of fear and laziness (its easier to dismiss the entire enterprise of data mining than learn how to do it responsibly and teach people how to interpret data properly and how to tell if someone else is interpreting data properly). I think that, on some fundamental level, we fear data, not the corporations or the governments or the scientists who are gathering it and using it, but the data itself. Given the rate at which behavioral data is piling up, this stance is becoming increasingly irresponsible each day. We can either let someone else aggregate all this data and learn why we do things and how to manipulate us, or we can take control of our own destinies and learn how our minds work so as to beat the others to the punch, to alter our behaviors so as to become less predictable. There's no going back to the pre-data age.

And really, the machine is just an extension of established sciences that spring from our desire to know ourselves more fully. Psychology, sociology, and really all of the social sciences are imperfect versions of the machine. They look for patterns leading up to outcomes we judge to be good or bad, but they have huge blind spots. But the blind spots are shrinking. Sociology and psychology have always played the red-headed stepchild to "hard sciences" like physics and chemistry. A great deal of this ill-will comes from the fact that social science can not deliver the levels of certainty that are the norm in hard sciences (hence the tolerance of smaller effect sizes in social science). That, too, may change in the petabyte age.