Showing posts with label Rethinking statistical criteria of effectiveness. Show all posts
Showing posts with label Rethinking statistical criteria of effectiveness. Show all posts

Friday, May 6, 2016

Are the Recommendations of the What Works Clearinghouse Useful/Reliable for Practitioners?

In 2002 the US Congress established the What Works Clearinghouse in the US Department of Education’s Institute of Education Science. What Works Clearinghouse’s role is to set rigorous quantitative standards based on the best science to provide guidance for practitioners as to what works by assessing the quality of research evidence supporting a given intervention. The What Works Clearinghouse then issues a rating on whether the evidence behind an intervention meets its standard of evidence or does not. The What Works Clearinghouse performs the same function in education that the Food and Drug Administration (FDA) does in medicine.

Can leaders trust the recommendations of the What Works Clearinghouse? Are they valid?  

The above text on quantitative analysis raises questions about the quality of the WWC's reviews (see Chapter 3). Specifically the text criticizes WWC for validating the research behind 'Success for All' when there was lots of published contrary evidence, and for concluding that 'charter schools' are more effective than 'traditional public schools' despite the tiny Effect Sizes (ESs). Are these problems anomalies—or are there deep seated problems with WWC's recommendations?


An amazing recent article suggests that the problems with the recommendations of the WWC are deep seated. Ginsburg and Smith (2016) examined the evidence for all the math programs certified by the What Works Clearinghouse as having evidence of effectiveness as provided by the most rigorous “gold standard” research design—Randomized Controlled Trials (RCT). They reviewed all 18 math programs that had been certified by WCC which contained 27 approved RCT studies. They found 12 potential threats to the usefulness of these studies, and concluded that “…none of the RCT’s provides useful information for consumers wishing to make informed judgments about what mathematics curriculum to purchase." 


One of the key problems across the What Works Clearinghouse math studies was that where Ginsburg and Smith (2016) were able to determine the error effects of a threat, the error generated by even one of those threats was at least as great as the Effect Size favoring the treatment group. In others words, the 'gold-standard' of research is not so golden.


BOTTOM LINES

  • Leaders cannot not trust the recommendations of the What Works Clearinghouse (WWC). Leaders need to conduct their own due diligence using the techniques in the book—both on research in general and on the recommendations of the WWC. 
  • The fact that the WWC requires the most rigorous research methodology and statistical evidence, and there are still a dozen potential "threats" to the results, suggests that he real world of practice is too complex for the traditional experimental approach and its reliance on relative measures of performance.  In other words, it is virtually impossible to establish full internal validity in applied experimental research regardless of how rigorous the research standards are. (Simpler alternatives for assessing the effectiveness of interventions are discussed in Chapter 5 of my methodology book.)   

Ginsburg, A., & Smith, M.S., (2016). Do randomized control trials meet the “Gold Standard”? A study of the usefulness of RCTs in the What Works Clearinghouse. The URL to this article is:  
            
            http://www.aei.org/wp-content/uploads/2016/03/Do-randomized-controlled-trials-meet-the-gold-standard.pdf 

Wednesday, October 7, 2015

60 % of psychology findings do not replicate: Implications for educational practice and research

Late august somewhat shocking results were published wherein efforts to replicate the findings of some of the most important studies in psychology failed. Replication of findings is one of the most important aspects of science. Results should not be taken seriously until they are replicated. However, most studies never et replicated.

What this means is that a majority of the most important findings in the field of psychology, ones that had been widely trusted and used by therapists and psychiatrists, are not valid.

Implications for Education Research

When you decide to implement something based on a research finding you are assuming that the results from the research will be replicated in your schools. However, if laboratory-based research cannot be replicated in the laboratory, it is even more unlikely that the results will be replicated in your school(s) where even less control exists: i.e., it is unlikely that you will see the same positive effects in your school?

Does this mean that leaders should not trust research findings? No!

The good news is that the glass is almost half full given that forty percent of the studies were replicated. What was the key characteristic of the replicable studies—especially since there were no indications of fraud? The key difference was that the replicated studies had larger effects, more positive results/greater effects/greater benefits, than those were not replicable. This validates a key point of Chapter 3 which discusses criteria for determining the practical significance of research. The basic conclusion is that you should only trust research with BIG effects and Chapter 3 provides guidance in how to identify such research. If you follow the recommendations in that chapter, chances are better that any practice that you adopt on the basis of research will be in the category of the 40% that will replicate the positive effects in your school(s).