Skip to content
← Back to BlogMarch 18, 2014

Evaluating Empirical Evidence: The False Promise of Statistical Studies

By Christian Chessman

The strongest literature on the March topic is statistical. Statistics – and statistical studies – are endowed with an aura of scientific legitimacy that might be the edge in a close debate round. That's why the ability to analyze studies – and explain that analysis effectively to a lay judge audience – is of paramount importance both generally as a debater and specifically on this topic.

Be Skeptical of Studies

Statistical studies claim to draw causal links between one variable and a second. People citing the study trust that the authors have eliminated (or "controlled for") all other possible factors that might explain the connection they identify. They also claim to identify data points that accurately represent the concepts they are trying to measure – for example, they take "scores on X standardized test" as representative of student achievement. While social scientists regularly attempt to control for confounding variables, and attempt to identify data that represents the measured concepts, they often fail to do so because of limited data availability and concept complexity.

Valid Studies Have…

Social scientists performing statistical analysis have several checks that ensure the reliability of their findings. Statistical studies rely on a small subset of a group – called a sample – to represent the entire group. Samples are not necessarily reliable though – if the sample is unrepresentative of the entire group, then results from analyzing that sample are also unrepresentative, and don't lead to valid conclusions. To maximize the reliability of sample representativeness, statistical studies should include two factors; random samples and large–N samples. Random sampling increases the validity of statistical studies by decreasing the likelihood that another variable could explain the results the study indicates. The Yale Statistics Project gives an illustrative example1:

Since it is generally impossible to study an entire population (every individual in a country, all college students, every geographic area, etc.), researchers typically rely on sampling to acquire a section of the population to perform an experiment or observational study. It is important that the group selected be representative of the population, and not biased in a systematic manner. For example, a group comprised of the wealthiest individuals in a given area probably would not accurately reflect the opinions of the entire population in that area. For this reason, randomization is typically employed to achieve an unbiased sample.

While randomness is an important element of sampling, it is not sufficient by itself to ensure that a study is reliable. Stephen Morris, a statistician who specializes in teaching medical research professionals how to use statistics, explains that a large sample is also necessary to ensure reliability. "In addition to randomness, the sample must be large–N (or large sample size). N is the universally accepted variable for sample size (how many data points you have, whether those data points are test scores or interviews)"2. As the size of a sample increases, its validity increases because it is more likely to accurately represent the population. An extreme example makes the point; if you wanted to tell the average feeling of a student towards single–gender classes, and you used random sampling but only selected one individual (study size n=1), then your sample is unrepresentative of the general student population despite random selection. The importance of large N studies is also magnified by the complexity of an issue. If a particular question can only have two answers – such as a yes or no – it's easier to get a representative sample because there are only two choices. Not all variables can be reduced to a simple yes or no, though – for example, there are literally dozens of indicators of educational quality, and many of those variables change over time. A single datapoint at a single time of a single indicator is less likely to wholly represent educational quality than a variable which is an unchanging yes or no answer. As such, the size of a study is a determinative factor in establishing the validity of a study. It's important to note that these are only two indicators of study reliability – studies with large–N, random samples are not necessarily good studies. The presence of these two indicators is a necessary condition of a good study, are not themselves sufficient to establish a study's validity. It's possible that a study has both of these indicators but is otherwise a fundamentally defective study. Studies may suffer other defects in design, modeling framework choice, variable selection, variable representativeness, and data insufficiency – all factors which can debilitate a study with otherwise random sampling and a large–N sample. Debaters analyzing a topic should search for possible confounding variables or alternative explanations on a topic that is statistics–heavy, and then ask their opponents to demonstrate where their study indicates it controls for that variable. This is both strategically effective because it sets an implicit burden that assumes the study is invalid until proven differently, and because it is rhetorically powerful to sit in silence for the awkward 30 seconds as your opponents feverishly flip through their study attempting to find proof their study accounts for your variable. The concepts of sample size and sample randomness establish important analytical tools which debaters can wield to win rounds. Direct evidence comparison is an underutilized skill in debate; rather than analyzing their opponents' evidence, debaters may simply abandon an argument as a "wash" when there are conflicting studies. The ability to argue that a study is methodologically superior can clear deadlocks between studies and give debaters an advantage that their opponents may not expect. The ability to challenge the methodological validity of a study is particularly important when debaters do not have evidence directly refuting a study presented by their opponents. With these concepts, debaters can argue that despite their lack of countervailing evidence, the evidence offered by their opponents is fundamentally insufficient to establish their claims. There are two types of statistical studies that debaters typically cite; direct studies that examine a particular dataset and indirect studies that survey the published literature, aggregate studies on the topic, and then synthesize those studies to form a conclusion. These studies – called meta–analyses – often attempt to reconcile disagreements in the field and streamline or summarize the conclusion or consensus demonstrated by majority of the evidence. For that reason, debaters tend to treat meta–analyses as more reliable than a direct study. Unfortunately, the problems we discussed above with individual studies grow exponentially when grouped together.

Meta–Analysis – Scaling the Problems

While meta–analyses are intuitively appealing, the problems with establishing reliable studies with accurately representative samples become more severe when scaled across studies. There are three major limitations to meta–analyses which decrease their reliability, and can serve as reasons to reject in a debate round. The first major limitation is incomparability. The most difficult part of establishing a valid study is deciding what outcomes effectively represent a particular concept, and what data points effectively represent an outcome. For example, I may want to study educational quality (concept), so I decide to use standardized test scores (outcome) from the 2004 National Assessment of Educational Progress test (data). A different scholar may choose different data to represent the same outcome by analyzing data from state standardized tests instead of national standardized tests. A third scholar may choose to analyze the growth of test scores over time (examining improvement from one year to the next) while a fourth scholar may choose to analyze the rate of growth (the relative rate at which students are improving rather than the absolute increase or decrease in score). There are merits to each dataset, which is why even good scholarship often produces radically different results. The variance demonstrated above is simply the variance within datasets – some scholars may think standardized tests are an inappropriate measure of educational quality. They may choose to measure educational quality by looking at the amount of per capita funding each student receives. They could argue that per capita funding is a broader representation of the environment in which students learn (such as teacher quality, classroom resource availability, extracurricular presence and quality) than the limited demonstration of a standardized test (which can be confounded by a number of factors, like test taking nerves, student illness, and other problems related to using a single examination to represent a whole year's learning). Still other scholars could measure educational quality by the number of students who complete grades and graduate. These scholars could argue that the overall best indicator of educational quality is whether or not the students actually receive the education, and measure the percentage rate at which students drop out of a given school or type of school. The vast diversity of outcomes and data sets demonstrates the complication of performing a meta–analysis. Because different studies measure different outcomes and/or use different data, they lack a common measure for comparison. Comparing a study of test score growth against a study of dropout rate leads is comparing (educational) apples and oranges. That means the study analysts must make a studied value judgment about which measure is better, or arbitrarily discard one of the studies. For reasons of funding and time limits, the latter is not unusual –meta–analyses that do not include a survey of outcomes and data as well as their analysis are especially likely to have simply been arbitrary choices by the researcher. That leads comfortably into the second major limitation: the arbitrary inclusion and exclusion of studies. A second weakness that debaters can note is selection choice. The researchers who are conducting a meta–analysis have complete discretion in terms of choosing which studies to include and which studies to exclude. Researchers who purport to represent the entire body of academic research may very well simply omit studies that contradict the result they want. The human choice element adds a necessarily subjective component to meta–analysis that detracts from their objectivity. Most debaters who cite and deploy a meta–analysis will be unable to prove surveys the entire literature on a subject, which means they cannot rule selection bias out as an explanation for their evidence. Framing a meta–analysis that way – by pointing out the various potential ways human bias might have affected that result – decreases the persuasive power of such evidence. A final limitation applies even when the studies are comparable and the selection of measures is justified: study weakness. Meta–analyses may examine a series of studies which are themselves weak for whatever reason – for example, some topics lack the data required to make statistically valid conclusions. This is true of single–gender classes, whose data is inherently non–random because federal law requires them to be opt–in (which makes the data non–random because it's uniquely composed of students who wanted to be in single–gender classes in the first place). A meta–analysis can be likened to a chain – a group of studies are connected together to form a stronger conclusion. But if the individual links in that chain – like the individual studies in a meta–analysis – are weak, then the whole will also be weak. In short – a compilation of weak evidence is a weak compilation. These and other limitations strongly apply to the topic of single–gender schools.

Single–Gender Study Limits

The American Psychological Association effectively summarizes the data limits of single–gender schools in the United States:

"The entire literature on single–sex schooling is confounded by the possible presence of student and school selection biases," says Rebecca S. Bigler, PhD, a psychology professor at the University of Texas at Austin who studies gender role development and racial stereotyping. "You can't conclude a thing about single–sex schooling if you don't check for and control those two potential biases." Research on single–sex education is also complicated by the legal requirement that assignment to single–sex classes must be completely voluntary.

These factors have led social scientists who examined the studies regarding single–gender education to conclude: ""there is no well–designed research showing that single–sex education improves students' academic performance."4 This meta–analysis examined every study published in journals with peer review, and could not find a single study that was both well designed and demonstrated an increase in performance in single–gender classrooms. There are a number of confounding variables that make well–designed statistical analysis nearly impossible. First, the vast majority of single–gender schools (and schools with single–gender classes) are private schools with a predominantly wealthy student body.5 Studies involved in those schools are both questionable in method (the effect of wealth on student performance cannot be ruled out) and untopical – the March 2014 topic exclusively examines public schools. This problem eliminates the vast majority of studies on the topic:

The problem, many experts say, is that it's nearly impossible to compare apples to apples when it comes to single–sex versus coeducation. Most research on single–sex education has been done with private schools, not on single–sex classes in U.S. public schools. In addition, it's rare for any studies on the topic to use random assignment.

Analysis of public schools is fraught with difficulty as well. "Even if they are public — and not charter or magnet — schools often also make academic changes when they switch to a single–sex format, making it hard to attribute gains or falls to any one measure"7. As a result, gains in schools which implement single–gender classes "can be attributed to selection bias—the children involved were more committed students—and to the motivation and sense of mission among teachers and staff"8. Other variables that may jeopardize the validity of studies – variables which can be inquired about on cross examination – include cross–cultural variables, foreign or domestic sampling, the school's focus as a vocational versus collegiate preparatory school, the type of classes, the relative presence or absence of emphasis on extra–curricular activities, the structure of classes (block scheduling vs single–rooms), and whether or not the school is located in a rural or urban area9. As this post hopefully shows, the promise of statistical studies is fundamentally a false promise. Statistical studies amount to no more than fallible human researchers making highly subjective decisions while working against deadlines with limited data that is subject to manipulation. Studies in debate have often been the decision factor in rounds – this post hopefully has given debaters several tools to diminish the impact of studies on judge decisions, both on the March 2014 topic specifically and in debate rounds generally.

Works Cited

  1. Yale University. "Sampling." Yale Statistics Deparment, no date given. Web. http://www.stat.yale.edu/Courses/1997–98/101/sample.htm.
  2. Morris, Stephen. "The Importance of N (sample Size) in Statistics." Statistics for the Terrified. N.p., no date given. Web. http://www.conceptstew.co.uk/PAGES/nsamplesize.html.
  3. American Psychological Association. "Coed vs Single–sex ed" article in the APA Journal Monitor on Psychology, Vol 42. No. 2. Published February 20. https://www.apa.org/monitor/2011/02/coed.aspx
  4. Strauss, Valerie. "Kids Don't Learn Better in Single–sex Classes — Meta Analysis." The Washington Post Online. The Washington Post, 11 Feb. 2014. Web. http://www.washingtonpost.com/blogs/answer–sheet/wp/2014/02/11/kids–dont–learn–better–in–single–sex–classes–meta–analysis/.
  5. Kimmel, Michael. "Don't Segregate Boys and Girls in Classrooms." CNN. Cable News Network, 03 Feb. 2014. Web. http://www.cnn.com/2013/08/09/opinion/kimmel–single–sex–classes/.
  6. American Psychological Association. "Coed vs Single–sex ed" article in the APA Journal Monitor on Psychology, Vol 42. No. 2. Published February 20`. https://www.apa.org/monitor/2011/02/coed.aspx
  7. Ibid
  8. Talbot, Margaret. "The Case Against Single–Sex Classrooms." The New Yorker. N.p., 12 July 2012. Web. 17 Mar. 2014. http://www.newyorker.com/online/blogs/comment/2012/07/single–sex–education.html.
  9. Kimmel, Michael. "Don't Segregate Boys and Girls in Classrooms." CNN. Cable News Network, 03 Feb. 2014. Web. http://www.cnn.com/2013/08/09/opinion/kimmel–single–sex–classes/.

More free resources