Wednesday, April 27, 2016

Two New Case Studies for Chapter 3

Here are two additional cases that can be incorporated as assignments after covering Chapter 3—with suggested solutions:

CASE STUDY #1 From the Headlines:

In April 27 2016 the New York Times described the latest results from the 2015 National Assessment of Educational Progress (NAEP) as follows:

"The results, from the National Assessment of Educational Progress, showed a drop in the percentage of students in private and public schools who are considered prepared for college-level work in reading and math. In 2013, the last time the test was given, 39 percent of students were estimated to be ready in math and 38 percent in reading; in 2015, 37 percent were judged prepared in each subject.

“This trend of stagnating scores is worrisome,” said Terry Mazany, the chairman of the governing board for the test... 

The math tests are scored from zero to 300, and in 12th grade, the average dropped to 152 in 2015 from 153 in 2013, a statistically significant decline. " (Zernike, 2016)

How do you interpret these results? Do they have practical significance? Should we be bemoaning a decline in student performance?

SOLUTION: These results have no practical importance for policy or practice. This is a classic case where a big deal is made of small differences simply because a large sample caused them to be statistically significant. At the same time, one must be vigilant to see if there is additional decline on the next NAEP test.


CASE STUDY #2 The Effectiveness of Cognitive Tutor

Is the Cognitive Tutor curriculum for improving student performance in Algebra I in both middle and high schools? Clearly, passing Algebra is a major hurdle for many students and is a gateway to developing the higher level math and science skills that are of increasing importance. As a result, increasing the passing rates in Algebra is also a national priority to as critical to increasing equity. This curriculum combines traditional types of materials with the latest of technology, and individualized intelligent math tutor.

In order to assess the practical significance of this intervention for your school(s), consider the following research:

Pane, J. F., Griffin, B. A., McCaffrey, D. F., and Karam, R. (2014). Effectiveness of cognitive   tutor algebra I at scale. Educational Evaluation and Policy Analysis, 36(2), 127-144.

Step 1. Download the article either from your universities library or from http://www.ncpeapublications.org/index.php/doc-center) site.

Step 2. Read the entire article but ignore all technical terms that you do not understand.

Step 3. Critique the study by.

a. Listing the positives of the study’s methodology
b. Identifying what the key missing piece(s) of data are in the results provided is?
c. Assess the results provided and whether you found it sufficiently compelling to consider adopting it for your schools—and explain why in terms of the specifics of the results.

SOLUTION:  

a. Positives:


A major positive of this study is that it has a high level of transparency and reflectiveness. The researchers clearly present the data in an impartial and open manner. They point out negative results for the intervention. There are also places where they point out that a conclusion of theirs can be interpreted differently. For example, in the first column on page 40, in trying to equate what an ES of .2 means in terms of national expected growth, and they note that it is comparable, they do note that their ES did not include summer loss which the expected national summer loss did.
A second positive is that no members of the sample were dropped even when there was not all the information needed to perfectly match them. In addition, the sample seems to be large and diverse enough to make the study relevant.
A third positive is how they reported the outcome data. Like the study in the previous case study the researchers adjusted final scores via Analysis of Covariance. They are clearly making such adjustment more carefully and openly discuss the problems associated with such adjustment and try to use advanced techniques, and three different methods of adjustment to minimize potential problems. Table 4 & 5 show that there is no difference in outcomes between the three different levels of adjustment (Models 2,3 & 4).
A fourth positive is that they analyze the results separately for middle and high school. As a result it is clear that there is no rationale for middle schools to adopt this intervention. As a result, the rest of the analysis will focus just on the high school results. 
A fifth positive is that they conducted the study over two cohorts of students so that teachers were able to gain a year of experience, and indeed the high school results were better for Cohort II.
A sixth positive is that the researchers report both the unweighted results, i.e., the “no covariate” Model 1 in Table 4, and the fully weighted Model 4. They acknowledge on page 139 (Column 2) that the treatment effects for the unweighted results were “substantially lower than for the models with adjustments.” One minor criticism is would have been better for the researchers to simply state that “there was no significant treatment effect for the unweighted results.”     
As an aside, one interesting thing to note is that the researchers thought that the weighted high school results improved in Cohort II because the teachers reverted back to traditional approaches to teaching.
b.  Problem:
There is one big problem. There is no data on the actual post-test performance of the experimental and comparison groups. Everything is once again a description of the relative differences. We do not know how well the students did in Algebra or on any math test. How many passed? How many moved on to the next level of math? Why are we not told how the students actually did on the Algebra Proficiency Exam in a non-standardized form? Why not tell us how many of the 32 items on the post-teach each group got correct?
Instead, they try to make the ES of .2 seem important by:
·       First indicating that it is equivalent to students move from the 50th to 58th percentile. But that does not mean that anyone scored at the 58th percentile. Did they move from the 10th to the 18th? Did the lower performing students make any equivalent percentile gains?
·       Second, they compare this ES to typical achievement gains nationally in the affected grades. Not only was the ES reported in this study slightly smaller it did not include summer loss so it is even smaller than what is typically achieved nationally. 
In any event, this study compares the tradition of trying to make an ES at the smallest end of Cohen’s range seem more important that it actually is.
Should you adopt the program?
c.   Given that:
·       The program is expensive,
·       The intervention was not compared to traditional individualized drill and practice software,
·       The progress was less than what students typically achieve nationally,
·       The low ES that does not meet recommended minimums for potential practical significance, and
·       No absolute data on how the students actually did,
I would not recommend that anyone adopt this program based on this research.





Sunday, April 17, 2016

Happenings at the 2016 AERA Conference

I just returned from the 2016 AERA (American Education Research Association) conference in Washington DC. I set up a booth in the exhibitor area to highlight the book (Authentic Quantitative Analysis for Leadership Decision-Making). In a bit if irony I was across from the Harvard Education Press booth and they probably had a 100 books on display—my booth only had the one. At the same time, mine was the only booth dedicated to the EdD, and I was pleasantly surprised at the large number of people who stopped by.

I must have had conversations with individuals from 60-70 EdD programs from around the country. These conversations confirmed that many are concerned about the state of how quantitative research is taught. There is a sense that something is wrong as students are increasingly turning to qualitative research. There is nothing wrong with qualitative research if it is being used for the right reason—as opposed to students feeling that quantitative methods are too difficult and that they cannot master it, as well as not seeing the relevancy of the traditional complex quantitative methods for their practice.

There was a tremendous response to the ideas in the book and many were drawn to moving quantitative methods from being a course on statistics to one that focuses on leadership decision-making. It is also becoming clearer the forms of statistical analyses used for PhD programs to test theory and those used to inform leadership decision-making are different. It is not that the statistics are different, but the degree of statistical methodological control and criteria for interpreting the results are different. The big problem for practice is that the forms of statistical analyses typically found in published quantitative research tend to over-estimate the importance of the findings for improving practice in the real world. In other words, the methods used in published research on the effectiveness of practices being tested are overly complex and unintelligible, and then in the end the results are misleading.

When I talk to professors who specialize in policy and practice , but who are not methodologists, about these problems the typical comment is that they do not get involved with, or understand, quantitative research. The result is that we have as a profession have abdicated responsibility for making important decisions about what is effective and left that to statisticians. The statisticians/methodologists have developed powerful techniques that enable them to declare small differences as having important practical importance. But the reality is that these do not have actual real world importance (see the earlier post about the problems that small differences are causing in psychology).


That is why the focus of the book is on much simpler and more accurate ways to determine the practical importance of quantitative research evidence. It is not just that this is important for making better decisions as to what practices are likely to be effective in your settings, it is important for our profession to retake responsibility for interpreting evidence as to effective practices. This book is a critical tool for enabling each and every one of us who are the ones who best understand the dynamics within schools and the needs and abilities of students and teachers to stop being cowed by quantitative evidence and to critically embrace the valuable information contained within quantitative data.  

Wednesday, October 7, 2015

60 % of psychology findings do not replicate: Implications for educational practice and research

Late august somewhat shocking results were published wherein efforts to replicate the findings of some of the most important studies in psychology failed. Replication of findings is one of the most important aspects of science. Results should not be taken seriously until they are replicated. However, most studies never et replicated.

What this means is that a majority of the most important findings in the field of psychology, ones that had been widely trusted and used by therapists and psychiatrists, are not valid.

Implications for Education Research

When you decide to implement something based on a research finding you are assuming that the results from the research will be replicated in your schools. However, if laboratory-based research cannot be replicated in the laboratory, it is even more unlikely that the results will be replicated in your school(s) where even less control exists: i.e., it is unlikely that you will see the same positive effects in your school?

Does this mean that leaders should not trust research findings? No!

The good news is that the glass is almost half full given that forty percent of the studies were replicated. What was the key characteristic of the replicable studies—especially since there were no indications of fraud? The key difference was that the replicated studies had larger effects, more positive results/greater effects/greater benefits, than those were not replicable. This validates a key point of Chapter 3 which discusses criteria for determining the practical significance of research. The basic conclusion is that you should only trust research with BIG effects and Chapter 3 provides guidance in how to identify such research. If you follow the recommendations in that chapter, chances are better that any practice that you adopt on the basis of research will be in the category of the 40% that will replicate the positive effects in your school(s).


Friday, June 26, 2015

Hi Everyone:

This blog is dedicated to having conversations with readers of my book:

Authentic Quantitative Analysis for Education Leadership Decision-Making and EdD Dissertations: A Practical, Intuitive, and Intelligible Approach 

How to Critique and Apply Quantitative Research to Improve Practice, and Develop a Rigorous and Useful EdD Dissertation


The goal of this blog is to share reactions, critiques, suggestions, and experiences with using this book in an EdD program with each other. The advantages of publishing with NCPEA, is that in addition to keeping the price low for students, their print on demand model makes it possible to make frequent revisions.  

In addition, the goal is to start a conversation among us as a community about the use of quantitative analysis in EdD programs. Hopefully, those who join in will not only be the individuals who teach the methodology courses, but a wide range of faculty and program directors. Please share your ideas with me and the others who join in. 

Stanley Pogrow