'Value Added' Concept Proves Beneficial to Teacher Colleges
Only because there is a poverty in imagination and a real lack of understanding of the limitations of this approach. Oh yes, and a desire to control teachers, corporatize, privatize, etc.
-Angela
This blog on Texas education contains posts on higher education, as well as preK-12 policy accountability, testing, bilingual education, immigration, school finance, race, class, and gender issues at both the state and national level. It also represents my digital footprint, of life and career, as a community-engaged scholar in Texas.
Showing posts with label value-added. Show all posts
Showing posts with label value-added. Show all posts
Monday, February 27, 2012
Thursday, February 16, 2012
Leading mathematician debunks ‘value-added’
Excellent analysis of "value-added" modeling and why it's a bad idea. Quotes from within:
"When value-added models were first conceived, even their most ardent
supporters cautioned about their use [Sanders 1995, abstract]. They were
a new tool that allowed us to make sense of mountains of data, using
mathematics in the same way it was used to understand the growth of
crops or the effects of a drug. But that tool was based on a statistical
model, and inferences about individual teachers might not be valid,
either because of faulty assumptions or because of normal (and expected)
variation."
"People recognize that tests are an imperfect measure of educational
success, but when sophisticated mathematics is applied, they believe the
imperfections go away by some mathematical magic. But this is not
magic. What really happens is that the mathematics is used to disguise
the problems and intimidate people into ignoring them—a modern,
mathematical version of the Emperor’s New Clothes."
"Of course we should hold teachers accountable, but this does not mean we have to pretend that mathematical models can do something they cannot. Of course we should rid our schools of incompetent teachers, but value-added models are an exceedingly blunt tool for this purpose. In any case, we ought to expect more from our teachers than what value-added attempts to measure."
-Angela
05/09/2011
Leading mathematician debunks ‘value-added’ http://www.washingtonpost.com/blogs/answer-sheet/post/leading-mathematician-debunks-value-added/2011/05/08/AFb999UG_blog.html
--
This was written by John Ewing, president of Math for America, a nonprofit organization dedicated to improving mathematics education in U.S. public high schools by recruiting, training and retaining great teachers. This article originally appeared in the May Notices of the American Mathematics Society. It gives a comprehensive look at the history, current use and problems with the value-added model of assessing teachers. It is long but well worth your time.
By John Ewing
Mathematicians occasionally worry about the misuse of their subject. G. H. Hardy famously wrote about mathematics used for war in his autobiography, A Mathematician’s Apology (and solidified his reputation as a foe of applied mathematics in doing so). More recently, groups of mathematicians tried to organize a boycott of the Star Wars [missile defense] project on the grounds that it was an abuse of mathematics. And even more recently some fretted about the role of mathematics in the financial meltdown.
But the most common misuse of mathematics is simpler, more pervasive, and (alas) more insidious: mathematics employed as a rhetorical weapon—an intellectual credential to convince the public that an idea or a process is “objective” and hence better than other competing ideas or processes. This is mathematical intimidation. It is especially persuasive because so many people are awed by mathematics and yet do not understand it—a dangerous combination.
The latest instance of the phenomenon is valued-added modeling (VAM), used to interpret test data. Value-added modeling pops up everywhere today, from newspapers to television to political campaigns. VAM is heavily promoted with unbridled and uncritical enthusiasm by the press, by politicians, and even by (some) educational experts, and it is touted as the modern, “scientific” way to measure educational success in everything from charter schools to individual teachers.
Yet most of those promoting value-added modeling are ill-equipped to judge either its effectiveness or its limitations. Some of those who are equipped make extravagant claims without much detail, reassuring us that someone has checked into our concerns and we shouldn’t worry. Value-added modeling is promoted because it has the right pedigree — because it is based on “sophisticated mathematics.”As a consequence, mathematics that ought to be used to illuminate ends up being used to intimidate. When that happens, mathematicians have a responsibility to speak out.
Background
Value-added models are all about tests—standardized tests that have become ubiquitous in K–12 education in the past few decades. These tests have been around for many years, but their scale, scope, and potential utility have changed dramatically.
Fifty years ago, at a few key points in their education, schoolchildren would bring home a piece of paper that showed academic achievement, usually with a percentile score showing where they landed among a large group. Parents could take pride in their child’s progress (or fret over its lack); teachers could sort students into those who excelled and those who needed remediation; students could make plans for higher education.
Today, tests have more consequences. “No Child Left Behind” mandated that tests in reading and mathematics be administered in grades 3–8. Often more tests are given in high school, including high-stakes tests for graduation.
With all that accumulating data, it was inevitable that people would want to use tests to evaluate everything educational—not merely teachers, schools, and entire states but also new curricula, teacher training programs, or teacher selection criteria. Are the new standards better than the old? Are experienced teachers better than novice? Do teachers need to know the content they teach?
Using data from tests to answer such questions is part of the current “student achievement” ethos—the belief that the goal of education is to produce high test scores. But it is also part of a broader trend in modern society to place a higher value on numerical (objective) measurements than verbal (subjective) evidence. But using tests to evaluate teachers, schools, or programs has many problems. (For a readable and comprehensive account, see [Koretz 2008].) Here are four of the most important problems, taken from a much longer list.
1. Influences. Test scores are affected by many factors, including the incoming levels of achievement, the influence of previous teachers, the attitudes of peers, and parental support. One cannot immediately separate the influence of a particular teacher or program among all those variables.
2. Polls. Like polls, tests are only samples. They cover only a small selection of material from a larger domain. A student’s score is meant to represent how much has been learned on all material, but tests (like polls) can be misleading.
3. Intangibles. Tests (especially multiple-choice tests) measure the learning of facts and procedures rather than the many other goals of teaching. Attitude, engagement, and the ability to learn further on one’s own are difficult to measure with tests. In some cases, these “intangible” goals may be more important than those measured by tests. (The father of modern standardized testing, E. F. Lindquist, wrote eloquently about this [Lindquist 1951]; a synopsis of his comments can be found in [Koretz 2008, 37].)
4. Inflation. Test scores can be increased without increasing student learning. This assertion has been convincingly demonstrated, but it is widely ignored by many in the education establishment [Koretz 2008, chap. 10]. In fact, the assertion should not be surprising. Every teacher knows that providing strategies for test-taking can improve student performance and that narrowing the curriculum to conform precisely to the test (“teaching to the test”) can have an even greater effect. The evidence shows that these effects can be substantial: One can dramatically increase test scores while at the same time actually decreasing student learning. “Test scores” are not the same as “student achievement.”
This last problem plays a larger role as the stakes increase. This is often referred to as Campbell’s Law: “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to measure” [Campbell 1976]. In its simplest form, this can mean that high-stakes tests are likely to induce some people (students, teachers, or administrators) to cheat ... and they do [Gabriel 2010].
But the more common consequence of Campbell’s Law is a distortion of the education experience, ignoring things that are not tested (for example, student engagement and attitude) and concentrating on precisely those things that are.
Value-Added Models
In the past two decades, a group of statisticians has focused on addressing the first of these four problems. This was natural. Mathematicians routinely create models for complicated systems that are similar to a large collection of students and teachers with many factors affecting individual outcomes over time.
Here’s a typical, although simplified, example, called the “split-plot design.” You want to test fertilizer on a number of different varieties of some crop. You have many plots, each divided into subplots. After assigning particular varieties to each subplot and randomly assigning levels of fertilizer to each whole plot, you can then sit back and watch how the plants grow as you apply the fertilizer. The task is to determine the effect of the fertilizer on growth, distinguishing it from the effects from the different varieties. Statisticians have developed standard mathematical tools (mixed models) to do this.
Does this situation sound familiar? Varieties, plots, fertilizer ...students, classrooms, teachers?
Dozens of similar situations arise in many areas, from agriculture to MRI analysis, always with the same basic ingredients—a mixture of fixed and random effects—and it is therefore not surprising that statisticians suggested using mixed models to analyze test data and determine “teacher effects.”
This is often explained to the public by analogy. One cannot accurately measure the quality of a teacher merely by looking at the scores on a single test at the end of a school year. If one teacher starts with all poorly prepared students, while another starts with all excellent, we would be misled by scores from a single test given to each class.
To account for such differences, we might use two tests, comparing scores from the end of one year to the next. The focus is on how much the scores increase rather than the scores themselves. That’s the basic idea behind “value added.” But value-added models (VAMs) are much more than merely comparing successive test scores.
Given many scores (say, grades 3–8) for many students with many teachers at many schools, one creates a mixed model for this complicated situation. The model is supposed to take into account all the factors that might influence test results — past history of the student, socioeconomic status, and so forth. The aim is to predict, based on all these past factors, the growth in test scores for students taught by a particular teacher. The actual change represents this more sophisticated “value added”— good when it’s larger than expected; bad when it’s smaller.
The best-known VAM, devised by William Sanders, is a mixed model (actually, several models), which is based on Henderson’s mixed-model equations, although mixed models originate much earlier [Sanders 1997]. One calculates (a huge computational effort!) the best linear unbiased predictors for the effects of teachers on scores. The precise details are unimportant here, but the process is similar to all mathematical modeling, with underlying assumptions and a number of choices in the model’s construction.
History
Such cautions were qualified, however, and one can see the roots of the modern embrace of VAMs in two juxtaposed quotes from William Sanders, the father of the value-added movement, which appeared in an article in Teacher Magazine in the year 2000. The article’s author reiterates the familiar cautions about VAMs, yet in the next paragraph seems to forget them:
Sanders has always said that scores for individual teachers should not be released publicly. “That would be totally inappropriate,” he says. “This is about trying to improve our schools, not embarrassing teachers. If their scores were made available, it would create chaos because most parents would be trying to get their kids into the same classroom.”
Still, Sanders says, it’s critical that ineffective teachers be identified. “The evidence is overwhelming,” he says, “that if any child catches two very weak teachers in a row, unless there is a major intervention, that kid never recovers from it. And that’s something that as a society we can’t ignore” [Hill 2000].
Over the past decade, such cautions about VAM slowly evaporated, especially in the popular press. A 2004 article in The School Administrator complains that there have not been ways to evaluate teachers in the past but excitedly touts value added as a solution:
“Fortunately, significant help is available in the form of a relatively new tool known as value-added assessment. Because value-added isolates the impact of instruction on student learning, it provides detailed information at the classroom level. Its rich diagnostic data can be used to improve teaching and student learning. It can be the basis for a needed improvement in the calculation of adequate yearly progress. In time, once teachers and administrators grow comfortable with its fairness, value-added also may serve as the foundation for an accountability system at the level of individual educators [Hershberg 2004, 1].”
And newspapers such as The Los Angeles Times get their hands on seven years of test scores for students in the L.A. schools and then publish a series of exposés about teachers, based on a value-added analysis of test data, which was performed under contract [Felch 2010]. The article explains its methodology:
“The Times used a statistical approach known as value-added analysis, which rates teachers based on their students’ progress on standardized tests from year to year. Each student’s performance is compared with his or her own in past years, which largely controls for outside influences often blamed for academic failure: poverty, prior learning and other factors.
Though controversial among teachers and others, the method has been increasingly embraced by education leaders and policymakers across the country, including the Obama administration.”
It goes on to draw many conclusions, including:
“Many of the factors commonly assumed to be important to teachers’ effectiveness were not. Although teachers are paid more for experience, education and training, none of this had much bearing on whether they improved their students’ performance.”
The writer adds the now-common dismissal of any concerns:
“No one suggests using value-added analysis as the sole measure of a teacher. Many experts recommend that it count for half or less of a teacher’s overall evaluation.
“Nevertheless, value-added analysis offers the closest thing available to an objective assessment of teachers. And it might help in resolving the greater mystery of what makes for effective teaching, and whether such skills can be taught.”
The article goes on to do exactly what it says “no one suggests” — it measures teachers solely on the basis of their value-added scores.
What Might Be Wrong with VAM?
As the popular press promoted value-added models with ever-increasing zeal, there was a parallel, much less visible scholarly conversation about the limitations of value-added models. In 2003 a book with the title Evaluating Value-Added Models for Teacher Accountability laid out some of the problems and concluded:
“The research base is currently insufficient to support the use of VAM for high-stakes decisions. We have identified numerous possible sources of error in teacher effects and any attempt to use VAM estimates for high-stakes decisions must be informed by an understanding of these potential errors [McCaffrey 2003, xx].”
In the next few years, a number of scholarly papers and reports raising concerns were published, including papers with such titles as “The Promise and Peril of Using Valued-Added Modeling to Measure Teacher Effectiveness” [RAND, 2004], “Re-Examining the Role of Teacher Quality in the Educational Production Function” [Koedel 2007], and “Methodological Concerns about the Education Value-Added Assessment System” [Amrein-Beardsley 2008].
What were the concerns in these papers? Here is a sample that hints at the complexity of issues.
• In the real world of schools, data is frequently missing or corrupt. What if students are missing past test data? What if past data was recorded incorrectly (not rare in schools)? What if students transferred into the school from outside the system?
• The modern classroom is more variable than people imagine. What if students are team-taught? How do you apportion credit or blame among various teachers? Do teachers in one class (say mathematics) affect the learning in another (say science)?
• Every mathematical model in sociology has to make rules, and they sometimes seem arbitrary. For example, what if students move into a class during the year? (Rule: Include them if they are in class for 150 or more days.) What if we only have a couple years of test data, or possibly more than five years? (Rule: The range three to five years is fixed for all models.) What’s the rationale for these kinds of rules?
• Class sizes differ in modern schools, and the nature of the model means there will be more variability for small classes. (Think of a class of one student.) Adjusting for this will necessarily drive teacher effects for small classes toward the mean. How does one adjust sensibly?
• While the basic idea underlying value-added models is the same, there are in fact many models. Do different models applied to the same data sets produce the same results? Are value-added models “robust”?
•Since models are applied to longitudinal data sequentially, it is essential to ask whether the results are consistent year to year. Are the computed teacher effects comparable over successive years for individual teachers? Are value-added models “consistent”?
These last two points were raised in a research paper [Lockwood 2007] and a recent policy brief from the Economic Policy Institute, “Problems with the Use of Student Test Scores to Evaluate Teachers”, which summarizes many of the open questions about VAM:
“For a variety of reasons, analyses of VAM results have led researchers to doubt whether the methodology can accurately identify more and less effective teachers. VAM estimates have proven to be unstable across statistical models, years, and classes that teachers teach. One study found that across five large urban districts, among teachers who were ranked in the top 20% of effectiveness in the first year, fewer than a third were in that top group the next year, and another third moved all the way down to the bottom 40%. Another found that teachers’ effectiveness ratings in one year could only predict from 4% to 16% of the variation in such ratings in the following year.
“Thus, a teacher who appears to be very ineffective in one year might have a dramatically different result the following year. The same dramatic fluctuations were found for teachers ranked at the bottom in the first year of analysis. This runs counter to most people’s notions that the true quality of a teacher is likely to change very little over time and raises questions about whether what is measured is largely a “teacher effect” or the effect of a wide variety of other factors [Baker 2010, 1].”
In addition to checking robustness and stability of a mathematical model, one needs to check validity. Are those teachers identified as superior (or inferior) by value-added models actually superior (or inferior)? This is perhaps the shakiest part of VAM. There has been surprisingly little effort to compare valued-added rankings to other measures of teacher quality, and to the extent that informal comparisons are made (as in the LA Times article), they sometimes don’t agree with common sense.
None of this means that value-added models are worthless—they are not. But like all mathematical models, they need to be used with care and a full understanding of their limitations.
How Is VAM Used?
Many studies by reputable scholarly groups call for caution in using VAMs for high-stakes decisions about teachers.
A RAND research report: The estimates from VAM modeling of achievement will often be too imprecise to support some of the desired inferences [McCaffrey 2004, 96].
A policy paper from the Educational Testing Service’s Policy Information Center: VAM results should not serve as the sole or principal basis for making consequential decisions about teachers. There are many pitfalls to making causal attributions of teacher effectiveness on the basis of the kinds of data available from typical school districts. We still lack sufficient understanding of how seriously the different technical problems threaten the validity of such interpretations [Braun 2005, 17].
A report from a workshop of the National Academy of Education: Value-added methods involve complex statistical models applied to test data of varying quality. Accordingly, there are many technical challenges to ascertaining the degree to which the output of these models provides the desired estimates [Braun 2010].
And yet here is the LA Times , publishing value-added scores for individual teachers by name and bragging that even teachers who were considered first-rate turn out to be “at the bottom”. In an episode reminiscent of the Cultural Revolution, the LA Times reporters confront a teacher who “was surprised and disappointed by her [value-added] results, adding that her students did well on periodic assessments and that parents seemed well-satisfied” [Felch 2010]. The teacher is made to think about why she did poorly and eventually, with the reporter’s help, she understands that she fails to challenge her students sufficiently. In spite of parents describing her as “amazing” and the principal calling her one of the “most effective” teachers in the school, she will have to change. She recants: “If my student test scores show I’m an ineffective teacher, I’d like to know what contributes to it. What do I need to do to bring my average up?”
Making policy decisions on the basis of value-added models has the potential to do even more harm than browbeating teachers. If we decide whether alternative certification is better than regular certification, whether nationally board certified teachers are better than randomly selected ones, whether small schools are better than large, or whether a new curriculum is better than an old by using a flawed measure of success, we almost surely will end up making bad decisions that affect education for decades to come.
This is insidious because, while people debate the use of value-added scores to judge teachers, almost no one questions the use of test scores and value-added models to judge policy. Even people who point out the limitations of VAM appear to be willing to use “student achievement” in the form of value-added scores to make such judgments. People recognize that tests are an imperfect measure of educational success, but when sophisticated mathematics is applied, they believe the imperfections go away by some mathematical magic. But this is not magic. What really happens is that the mathematics is used to disguise the problems and intimidate people into ignoring them—a modern, mathematical version of the Emperor’s New Clothes.
What Should Mathematicians Do?
The concerns raised about value-added models ought to give everyone pause, and ordinarily they would lead to a thoughtful conversation about the proper use of VAM. Unfortunately, VAM proponents and politicians have framed the discussion as a battle between teacher unions and the public.
Shouldn’t teachers be accountable? Shouldn’t we rid ourselves of those who are incompetent? Shouldn’t we put our students first and stop worrying about teacher sensibilities? And most importantly, shouldn’t we be driven by the data?
This line of reasoning is illustrated by a recent fatuous report from the Brookings Institute, “Evaluating Teachers: The Important Role of Value-Added” [Glazerman 2010], which dismisses the many cautions found in all the papers mentioned above, not by refuting them but by asserting their unimportance. The authors of the Brookings paper agree that value-added scores of teachers are unstable (that is, not highly correlated year to year) but go on to assert:
“The use of imprecise measures to make high-stakes decisions that place societal or institutional interests above those of individuals is widespread and accepted in fields outside of teaching [Glazerman 2010, 7].”
To illustrate this point, they use examples such as the correlation of SAT scores with college success or the year-by-year correlation of leaders in real estate sales. They conclude that “a performance measure needs to be good, not perfect”. (And as usual, on page 11 they caution not to use value-added measures alone when making decisions, while on page 9 they advocate doing precisely that.)
Why must we use value-added even with its imperfections? Aside from making the unsupported claim (in the very last sentence) that “it predicts more about what students will learn ... than any other source of information,” the only apparent reason for its superiority is that value-added is based on data. Here is mathematical intimidation in its purest form—in this case, in the hands of economists, sociologists, and education policy experts.
A number of people and organizations are seeking better ways to
evaluate teacher performance in new ways that focus on measuring much
more than test scores. (See, for example, the Measures of Effective Teaching project
run by the Gates Foundation.) Shouldn’t we try to measure long-term
student achievement, not merely short-term gains? Shouldn’t we focus on
how well students are prepared to learn in the future, not merely what
they learned in the past year? Shouldn’t we try to distinguish teachers
who inspire their students, not merely the ones who are competent?
When we accept value-added as an “imperfect” substitute for all these things because it is conveniently at hand, we are not raising our expectations of teachers, we are lowering them. And if we drive away the best teachers by using a flawed process, are we really putting our students first?
Whether naïfs or experts, mathematicians need to confront people who misuse their subject to intimidate others into accepting conclusions simply because they are based on some mathematics. Unlike many policy makers, mathematicians are notbamboozled by the theory behind VAM, and they need to speak out forcefully. Mathematical models have limitations. They do not by themselves convey authority for their conclusions. They are tools, not magic. And using the mathematics to intimidate — to preempt debate about the goals of education and measures of success — is harmful not only to education but to mathematics itself.
References
Audrey Amrein-Beardsley, Methodological concerns about the education value-added assessment system, Educational Researcher 37 (2008), 65–75. http:// dx.doi.org/10.3102/0013189X08316420
Eva L. Baker, Paul E. Barton, Linda Darling-Hammond, Edward Haertel, Hellen F. Ladd, Robert L. Linn, Diane Ravitch, Richard Rothstein, Richard J. Shavelson, and Lorrie A. Shepard, Problems with the Use of Student Test Scores to Evaluate Teachers, Economic Policy Institute Briefing Paper#278, August 29, 2010, Washington, DC. http://www.epi.org/publications/entry/bp278
Henry Braun, Using Student Progress to Evaluate Teachers:
A Primer on Value-Added Models, Educational Testing Service Policy Perspective, Princeton, NJ, 2005. http://www.ets.org/Media/Research/pdf/PICVAM.pdf
Henry Braun, Naomi Chudowsky, and Judith Koenig, eds., Getting Value Out of Value-Added: Report of a Workshop, Committee on Value-Added Methodology for Instructional Improvement, Program Evaluation, and Accountability; National Research Council, Washington, DC, 2010.
http://www.nap.edu/catalog/12820.html
Donald T. Campbell, Assessing the Impact of Planned Social Change, Dartmouth College, Occasional Paper Series, #8, 1976. http://www.eric.ed.gov/PDFS/ED303512.pdf
Jason Felch, Jason Song, and Doug Smith, Who’s teaching L.A.’s kids?, Los Angeles Times, August 14, 2010. http://www.latimes.com/news/local/la-me-teachers-value-20100815,0,2695044.story
Trip Gabriel, Under pressure, teachers tamper with tests, New York Times, June 11, 2010.
http://www.nytimes.com/2010/06/11/education/11cheat.html
Steven Glazerman, Susanna Loeb, Dan Goldhaber, Douglas Staiger, Stephen Raudenbush, Grover Whitehurst, Evaluating Teachers: The Important Role of Value-Added, Brown Center on Education Policy at Brookings, 2010. http://www.brookings.edu/reports/2010/1117_evaluating_teachers.aspx
Ted Hershberg, Virginia Adams Simon and Barbara Lea Kruger, The revelations of value-added: An assessment model that measures student growth in ways that NCLB fails to do, The School Administrator, December 2004.
http://www.aasa.org/SchoolAdministratorArticle.aspx?id=9466
David Hill, He’s got your number, Teacher Magazine,
May 2000 11(8), 42–47.
http://www.edweek.org/tm/articles/2000/05/01/08sanders.h11.html
Cory Koedel and Julian R. Betts, Re-Examining the Role of Teacher Quality in the Educational Production Function, Working Paper #2007-03, National Center on Performance Initiatives, Nashville, TN, 2007.
http://economics.missouri.edu/working-papers/2007/wp0708_koedel.pdf
Daniel Koretz, Measuring Up: What Educational Testing Really Tells Us , Harvard University Press, Cambridge, Massachusetts, 2008.
E. F. Lindquist, Preliminary considerations in objective test construction, in Educational Measurement (E. F. Lindquist, ed.), American Council on Education, Washington DC, 1951.
J. R. Lockwood, Daniel McCaffrey, Laura S. Hamilton, Brian Stetcher, Vi-Nhuan Le, and Felipe Martinez, The sensitivity of value-added teacher effect estimates to different mathematics achievement measures, Journal of Educational Measurement 44(1) (2007), 47–67. http://dx.doi.org/10.1111/j.1745-3984.2007.00026.x
Daniel F. McCaffrey, Daniel Koretz, J. R. Lockwood, and Laura S. Hamilton, Evaluating Value-Added Models for Teacher Accountability, RAND Corporation, Santa Monica, CA, 2003.
http://www.rand.org/pubs/monographs/2004/RAND_MG158.pdf
Daniel F. McCaffrey, J. R. Lockwood, Daniel Koretz, Thomas A. Louis, and Laura Hamilton, Models for value-added modeling of teacher effects, Journal of Educational and Behavioral Statistics 29(1), Spring 2004, 67-101.
http://www.rand.org/pubs/reprints/2005/RAND_RP1165.pdf
RAND Research Brief, The Promise and Peril of Using Value-Added Modeling to Measure Teacher Effectiveness, Santa Monica, CA, 2004. http://www.rand.org/pubs/research_briefs/RB9050/RAND_RB9050.pdf
William L. Sanders and Sandra P. Horn, Educational Assessment Reassessed: The Usefulness of Standardized and Alternative Measures of Student Achievement as Indicators of the Assessment of Educational Outcomes, Education Policy Analysis Archives, March 3(6) (1995). http://epaa.asu.edu/ojs/article/view/649
W. Sanders, A. Saxton, and B. Horn, The Tennessee value-added assessment system: A quantitative outcomes-based approach to educational assessment, in Grading Teachers, Grading Schools: Is Student Achievement a Valid Evaluational Measure? (J. Millman, ed.), Corwin Press, Inc., Thousand Oaks, CA, 1997, pp 137–162.
-0-
Follow The Answer Sheet every day by bookmarking http://www.washingtonpost.com/blogs/answer-sheet. And for admissions advice, college news and links to campus papers, please check out our Higher Education page. Bookmark it!
This was written by John Ewing, president of Math for America, a nonprofit organization dedicated to improving mathematics education in U.S. public high schools by recruiting, training and retaining great teachers. This article originally appeared in the May Notices of the American Mathematics Society. It gives a comprehensive look at the history, current use and problems with the value-added model of assessing teachers. It is long but well worth your time.
By John Ewing
Mathematicians occasionally worry about the misuse of their subject. G. H. Hardy famously wrote about mathematics used for war in his autobiography, A Mathematician’s Apology (and solidified his reputation as a foe of applied mathematics in doing so). More recently, groups of mathematicians tried to organize a boycott of the Star Wars [missile defense] project on the grounds that it was an abuse of mathematics. And even more recently some fretted about the role of mathematics in the financial meltdown.
But the most common misuse of mathematics is simpler, more pervasive, and (alas) more insidious: mathematics employed as a rhetorical weapon—an intellectual credential to convince the public that an idea or a process is “objective” and hence better than other competing ideas or processes. This is mathematical intimidation. It is especially persuasive because so many people are awed by mathematics and yet do not understand it—a dangerous combination.
The latest instance of the phenomenon is valued-added modeling (VAM), used to interpret test data. Value-added modeling pops up everywhere today, from newspapers to television to political campaigns. VAM is heavily promoted with unbridled and uncritical enthusiasm by the press, by politicians, and even by (some) educational experts, and it is touted as the modern, “scientific” way to measure educational success in everything from charter schools to individual teachers.
Yet most of those promoting value-added modeling are ill-equipped to judge either its effectiveness or its limitations. Some of those who are equipped make extravagant claims without much detail, reassuring us that someone has checked into our concerns and we shouldn’t worry. Value-added modeling is promoted because it has the right pedigree — because it is based on “sophisticated mathematics.”As a consequence, mathematics that ought to be used to illuminate ends up being used to intimidate. When that happens, mathematicians have a responsibility to speak out.
Background
Value-added models are all about tests—standardized tests that have become ubiquitous in K–12 education in the past few decades. These tests have been around for many years, but their scale, scope, and potential utility have changed dramatically.
Fifty years ago, at a few key points in their education, schoolchildren would bring home a piece of paper that showed academic achievement, usually with a percentile score showing where they landed among a large group. Parents could take pride in their child’s progress (or fret over its lack); teachers could sort students into those who excelled and those who needed remediation; students could make plans for higher education.
Today, tests have more consequences. “No Child Left Behind” mandated that tests in reading and mathematics be administered in grades 3–8. Often more tests are given in high school, including high-stakes tests for graduation.
With all that accumulating data, it was inevitable that people would want to use tests to evaluate everything educational—not merely teachers, schools, and entire states but also new curricula, teacher training programs, or teacher selection criteria. Are the new standards better than the old? Are experienced teachers better than novice? Do teachers need to know the content they teach?
Using data from tests to answer such questions is part of the current “student achievement” ethos—the belief that the goal of education is to produce high test scores. But it is also part of a broader trend in modern society to place a higher value on numerical (objective) measurements than verbal (subjective) evidence. But using tests to evaluate teachers, schools, or programs has many problems. (For a readable and comprehensive account, see [Koretz 2008].) Here are four of the most important problems, taken from a much longer list.
1. Influences. Test scores are affected by many factors, including the incoming levels of achievement, the influence of previous teachers, the attitudes of peers, and parental support. One cannot immediately separate the influence of a particular teacher or program among all those variables.
2. Polls. Like polls, tests are only samples. They cover only a small selection of material from a larger domain. A student’s score is meant to represent how much has been learned on all material, but tests (like polls) can be misleading.
3. Intangibles. Tests (especially multiple-choice tests) measure the learning of facts and procedures rather than the many other goals of teaching. Attitude, engagement, and the ability to learn further on one’s own are difficult to measure with tests. In some cases, these “intangible” goals may be more important than those measured by tests. (The father of modern standardized testing, E. F. Lindquist, wrote eloquently about this [Lindquist 1951]; a synopsis of his comments can be found in [Koretz 2008, 37].)
4. Inflation. Test scores can be increased without increasing student learning. This assertion has been convincingly demonstrated, but it is widely ignored by many in the education establishment [Koretz 2008, chap. 10]. In fact, the assertion should not be surprising. Every teacher knows that providing strategies for test-taking can improve student performance and that narrowing the curriculum to conform precisely to the test (“teaching to the test”) can have an even greater effect. The evidence shows that these effects can be substantial: One can dramatically increase test scores while at the same time actually decreasing student learning. “Test scores” are not the same as “student achievement.”
This last problem plays a larger role as the stakes increase. This is often referred to as Campbell’s Law: “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to measure” [Campbell 1976]. In its simplest form, this can mean that high-stakes tests are likely to induce some people (students, teachers, or administrators) to cheat ... and they do [Gabriel 2010].
But the more common consequence of Campbell’s Law is a distortion of the education experience, ignoring things that are not tested (for example, student engagement and attitude) and concentrating on precisely those things that are.
Value-Added Models
In the past two decades, a group of statisticians has focused on addressing the first of these four problems. This was natural. Mathematicians routinely create models for complicated systems that are similar to a large collection of students and teachers with many factors affecting individual outcomes over time.
Here’s a typical, although simplified, example, called the “split-plot design.” You want to test fertilizer on a number of different varieties of some crop. You have many plots, each divided into subplots. After assigning particular varieties to each subplot and randomly assigning levels of fertilizer to each whole plot, you can then sit back and watch how the plants grow as you apply the fertilizer. The task is to determine the effect of the fertilizer on growth, distinguishing it from the effects from the different varieties. Statisticians have developed standard mathematical tools (mixed models) to do this.
Does this situation sound familiar? Varieties, plots, fertilizer ...students, classrooms, teachers?
Dozens of similar situations arise in many areas, from agriculture to MRI analysis, always with the same basic ingredients—a mixture of fixed and random effects—and it is therefore not surprising that statisticians suggested using mixed models to analyze test data and determine “teacher effects.”
This is often explained to the public by analogy. One cannot accurately measure the quality of a teacher merely by looking at the scores on a single test at the end of a school year. If one teacher starts with all poorly prepared students, while another starts with all excellent, we would be misled by scores from a single test given to each class.
To account for such differences, we might use two tests, comparing scores from the end of one year to the next. The focus is on how much the scores increase rather than the scores themselves. That’s the basic idea behind “value added.” But value-added models (VAMs) are much more than merely comparing successive test scores.
Given many scores (say, grades 3–8) for many students with many teachers at many schools, one creates a mixed model for this complicated situation. The model is supposed to take into account all the factors that might influence test results — past history of the student, socioeconomic status, and so forth. The aim is to predict, based on all these past factors, the growth in test scores for students taught by a particular teacher. The actual change represents this more sophisticated “value added”— good when it’s larger than expected; bad when it’s smaller.
The best-known VAM, devised by William Sanders, is a mixed model (actually, several models), which is based on Henderson’s mixed-model equations, although mixed models originate much earlier [Sanders 1997]. One calculates (a huge computational effort!) the best linear unbiased predictors for the effects of teachers on scores. The precise details are unimportant here, but the process is similar to all mathematical modeling, with underlying assumptions and a number of choices in the model’s construction.
History
When value-added models were first conceived, even their most ardent
supporters cautioned about their use [Sanders 1995, abstract]. They were
a new tool that allowed us to make sense of mountains of data, using
mathematics in the same way it was used to understand the growth of
crops or the effects of a drug. But that tool was based on a statistical
model, and inferences about individual teachers might not be valid,
either because of faulty assumptions or because of normal (and expected)
variation.
Such cautions were qualified, however, and one can see the roots of the modern embrace of VAMs in two juxtaposed quotes from William Sanders, the father of the value-added movement, which appeared in an article in Teacher Magazine in the year 2000. The article’s author reiterates the familiar cautions about VAMs, yet in the next paragraph seems to forget them:
Sanders has always said that scores for individual teachers should not be released publicly. “That would be totally inappropriate,” he says. “This is about trying to improve our schools, not embarrassing teachers. If their scores were made available, it would create chaos because most parents would be trying to get their kids into the same classroom.”
Still, Sanders says, it’s critical that ineffective teachers be identified. “The evidence is overwhelming,” he says, “that if any child catches two very weak teachers in a row, unless there is a major intervention, that kid never recovers from it. And that’s something that as a society we can’t ignore” [Hill 2000].
Over the past decade, such cautions about VAM slowly evaporated, especially in the popular press. A 2004 article in The School Administrator complains that there have not been ways to evaluate teachers in the past but excitedly touts value added as a solution:
“Fortunately, significant help is available in the form of a relatively new tool known as value-added assessment. Because value-added isolates the impact of instruction on student learning, it provides detailed information at the classroom level. Its rich diagnostic data can be used to improve teaching and student learning. It can be the basis for a needed improvement in the calculation of adequate yearly progress. In time, once teachers and administrators grow comfortable with its fairness, value-added also may serve as the foundation for an accountability system at the level of individual educators [Hershberg 2004, 1].”
And newspapers such as The Los Angeles Times get their hands on seven years of test scores for students in the L.A. schools and then publish a series of exposés about teachers, based on a value-added analysis of test data, which was performed under contract [Felch 2010]. The article explains its methodology:
“The Times used a statistical approach known as value-added analysis, which rates teachers based on their students’ progress on standardized tests from year to year. Each student’s performance is compared with his or her own in past years, which largely controls for outside influences often blamed for academic failure: poverty, prior learning and other factors.
Though controversial among teachers and others, the method has been increasingly embraced by education leaders and policymakers across the country, including the Obama administration.”
It goes on to draw many conclusions, including:
“Many of the factors commonly assumed to be important to teachers’ effectiveness were not. Although teachers are paid more for experience, education and training, none of this had much bearing on whether they improved their students’ performance.”
The writer adds the now-common dismissal of any concerns:
“No one suggests using value-added analysis as the sole measure of a teacher. Many experts recommend that it count for half or less of a teacher’s overall evaluation.
“Nevertheless, value-added analysis offers the closest thing available to an objective assessment of teachers. And it might help in resolving the greater mystery of what makes for effective teaching, and whether such skills can be taught.”
The article goes on to do exactly what it says “no one suggests” — it measures teachers solely on the basis of their value-added scores.
What Might Be Wrong with VAM?
As the popular press promoted value-added models with ever-increasing zeal, there was a parallel, much less visible scholarly conversation about the limitations of value-added models. In 2003 a book with the title Evaluating Value-Added Models for Teacher Accountability laid out some of the problems and concluded:
“The research base is currently insufficient to support the use of VAM for high-stakes decisions. We have identified numerous possible sources of error in teacher effects and any attempt to use VAM estimates for high-stakes decisions must be informed by an understanding of these potential errors [McCaffrey 2003, xx].”
In the next few years, a number of scholarly papers and reports raising concerns were published, including papers with such titles as “The Promise and Peril of Using Valued-Added Modeling to Measure Teacher Effectiveness” [RAND, 2004], “Re-Examining the Role of Teacher Quality in the Educational Production Function” [Koedel 2007], and “Methodological Concerns about the Education Value-Added Assessment System” [Amrein-Beardsley 2008].
What were the concerns in these papers? Here is a sample that hints at the complexity of issues.
• In the real world of schools, data is frequently missing or corrupt. What if students are missing past test data? What if past data was recorded incorrectly (not rare in schools)? What if students transferred into the school from outside the system?
• The modern classroom is more variable than people imagine. What if students are team-taught? How do you apportion credit or blame among various teachers? Do teachers in one class (say mathematics) affect the learning in another (say science)?
• Every mathematical model in sociology has to make rules, and they sometimes seem arbitrary. For example, what if students move into a class during the year? (Rule: Include them if they are in class for 150 or more days.) What if we only have a couple years of test data, or possibly more than five years? (Rule: The range three to five years is fixed for all models.) What’s the rationale for these kinds of rules?
• Class sizes differ in modern schools, and the nature of the model means there will be more variability for small classes. (Think of a class of one student.) Adjusting for this will necessarily drive teacher effects for small classes toward the mean. How does one adjust sensibly?
• While the basic idea underlying value-added models is the same, there are in fact many models. Do different models applied to the same data sets produce the same results? Are value-added models “robust”?
•Since models are applied to longitudinal data sequentially, it is essential to ask whether the results are consistent year to year. Are the computed teacher effects comparable over successive years for individual teachers? Are value-added models “consistent”?
These last two points were raised in a research paper [Lockwood 2007] and a recent policy brief from the Economic Policy Institute, “Problems with the Use of Student Test Scores to Evaluate Teachers”, which summarizes many of the open questions about VAM:
“For a variety of reasons, analyses of VAM results have led researchers to doubt whether the methodology can accurately identify more and less effective teachers. VAM estimates have proven to be unstable across statistical models, years, and classes that teachers teach. One study found that across five large urban districts, among teachers who were ranked in the top 20% of effectiveness in the first year, fewer than a third were in that top group the next year, and another third moved all the way down to the bottom 40%. Another found that teachers’ effectiveness ratings in one year could only predict from 4% to 16% of the variation in such ratings in the following year.
“Thus, a teacher who appears to be very ineffective in one year might have a dramatically different result the following year. The same dramatic fluctuations were found for teachers ranked at the bottom in the first year of analysis. This runs counter to most people’s notions that the true quality of a teacher is likely to change very little over time and raises questions about whether what is measured is largely a “teacher effect” or the effect of a wide variety of other factors [Baker 2010, 1].”
In addition to checking robustness and stability of a mathematical model, one needs to check validity. Are those teachers identified as superior (or inferior) by value-added models actually superior (or inferior)? This is perhaps the shakiest part of VAM. There has been surprisingly little effort to compare valued-added rankings to other measures of teacher quality, and to the extent that informal comparisons are made (as in the LA Times article), they sometimes don’t agree with common sense.
None of this means that value-added models are worthless—they are not. But like all mathematical models, they need to be used with care and a full understanding of their limitations.
How Is VAM Used?
Many studies by reputable scholarly groups call for caution in using VAMs for high-stakes decisions about teachers.
A RAND research report: The estimates from VAM modeling of achievement will often be too imprecise to support some of the desired inferences [McCaffrey 2004, 96].
A policy paper from the Educational Testing Service’s Policy Information Center: VAM results should not serve as the sole or principal basis for making consequential decisions about teachers. There are many pitfalls to making causal attributions of teacher effectiveness on the basis of the kinds of data available from typical school districts. We still lack sufficient understanding of how seriously the different technical problems threaten the validity of such interpretations [Braun 2005, 17].
A report from a workshop of the National Academy of Education: Value-added methods involve complex statistical models applied to test data of varying quality. Accordingly, there are many technical challenges to ascertaining the degree to which the output of these models provides the desired estimates [Braun 2010].
And yet here is the LA Times , publishing value-added scores for individual teachers by name and bragging that even teachers who were considered first-rate turn out to be “at the bottom”. In an episode reminiscent of the Cultural Revolution, the LA Times reporters confront a teacher who “was surprised and disappointed by her [value-added] results, adding that her students did well on periodic assessments and that parents seemed well-satisfied” [Felch 2010]. The teacher is made to think about why she did poorly and eventually, with the reporter’s help, she understands that she fails to challenge her students sufficiently. In spite of parents describing her as “amazing” and the principal calling her one of the “most effective” teachers in the school, she will have to change. She recants: “If my student test scores show I’m an ineffective teacher, I’d like to know what contributes to it. What do I need to do to bring my average up?”
Making policy decisions on the basis of value-added models has the potential to do even more harm than browbeating teachers. If we decide whether alternative certification is better than regular certification, whether nationally board certified teachers are better than randomly selected ones, whether small schools are better than large, or whether a new curriculum is better than an old by using a flawed measure of success, we almost surely will end up making bad decisions that affect education for decades to come.
This is insidious because, while people debate the use of value-added scores to judge teachers, almost no one questions the use of test scores and value-added models to judge policy. Even people who point out the limitations of VAM appear to be willing to use “student achievement” in the form of value-added scores to make such judgments. People recognize that tests are an imperfect measure of educational success, but when sophisticated mathematics is applied, they believe the imperfections go away by some mathematical magic. But this is not magic. What really happens is that the mathematics is used to disguise the problems and intimidate people into ignoring them—a modern, mathematical version of the Emperor’s New Clothes.
What Should Mathematicians Do?
The concerns raised about value-added models ought to give everyone pause, and ordinarily they would lead to a thoughtful conversation about the proper use of VAM. Unfortunately, VAM proponents and politicians have framed the discussion as a battle between teacher unions and the public.
Shouldn’t teachers be accountable? Shouldn’t we rid ourselves of those who are incompetent? Shouldn’t we put our students first and stop worrying about teacher sensibilities? And most importantly, shouldn’t we be driven by the data?
This line of reasoning is illustrated by a recent fatuous report from the Brookings Institute, “Evaluating Teachers: The Important Role of Value-Added” [Glazerman 2010], which dismisses the many cautions found in all the papers mentioned above, not by refuting them but by asserting their unimportance. The authors of the Brookings paper agree that value-added scores of teachers are unstable (that is, not highly correlated year to year) but go on to assert:
“The use of imprecise measures to make high-stakes decisions that place societal or institutional interests above those of individuals is widespread and accepted in fields outside of teaching [Glazerman 2010, 7].”
To illustrate this point, they use examples such as the correlation of SAT scores with college success or the year-by-year correlation of leaders in real estate sales. They conclude that “a performance measure needs to be good, not perfect”. (And as usual, on page 11 they caution not to use value-added measures alone when making decisions, while on page 9 they advocate doing precisely that.)
Why must we use value-added even with its imperfections? Aside from making the unsupported claim (in the very last sentence) that “it predicts more about what students will learn ... than any other source of information,” the only apparent reason for its superiority is that value-added is based on data. Here is mathematical intimidation in its purest form—in this case, in the hands of economists, sociologists, and education policy experts.
Of course we should hold teachers accountable, but this does not mean
we have to pretend that mathematical models can do something they
cannot. Of course we should rid our schools of incompetent teachers, but
value-added models are an exceedingly blunt tool for this purpose. In
any case, we ought to expect more from our teachers than what
value-added attempts to measure.
When we accept value-added as an “imperfect” substitute for all these things because it is conveniently at hand, we are not raising our expectations of teachers, we are lowering them. And if we drive away the best teachers by using a flawed process, are we really putting our students first?
Whether naïfs or experts, mathematicians need to confront people who misuse their subject to intimidate others into accepting conclusions simply because they are based on some mathematics. Unlike many policy makers, mathematicians are notbamboozled by the theory behind VAM, and they need to speak out forcefully. Mathematical models have limitations. They do not by themselves convey authority for their conclusions. They are tools, not magic. And using the mathematics to intimidate — to preempt debate about the goals of education and measures of success — is harmful not only to education but to mathematics itself.
References
Audrey Amrein-Beardsley, Methodological concerns about the education value-added assessment system, Educational Researcher 37 (2008), 65–75. http:// dx.doi.org/10.3102/0013189X08316420
Eva L. Baker, Paul E. Barton, Linda Darling-Hammond, Edward Haertel, Hellen F. Ladd, Robert L. Linn, Diane Ravitch, Richard Rothstein, Richard J. Shavelson, and Lorrie A. Shepard, Problems with the Use of Student Test Scores to Evaluate Teachers, Economic Policy Institute Briefing Paper#278, August 29, 2010, Washington, DC. http://www.epi.org/publications/entry/bp278
Henry Braun, Using Student Progress to Evaluate Teachers:
A Primer on Value-Added Models, Educational Testing Service Policy Perspective, Princeton, NJ, 2005. http://www.ets.org/Media/Research/pdf/PICVAM.pdf
Henry Braun, Naomi Chudowsky, and Judith Koenig, eds., Getting Value Out of Value-Added: Report of a Workshop, Committee on Value-Added Methodology for Instructional Improvement, Program Evaluation, and Accountability; National Research Council, Washington, DC, 2010.
http://www.nap.edu/catalog/12820.html
Donald T. Campbell, Assessing the Impact of Planned Social Change, Dartmouth College, Occasional Paper Series, #8, 1976. http://www.eric.ed.gov/PDFS/ED303512.pdf
Jason Felch, Jason Song, and Doug Smith, Who’s teaching L.A.’s kids?, Los Angeles Times, August 14, 2010. http://www.latimes.com/news/local/la-me-teachers-value-20100815,0,2695044.story
Trip Gabriel, Under pressure, teachers tamper with tests, New York Times, June 11, 2010.
http://www.nytimes.com/2010/06/11/education/11cheat.html
Steven Glazerman, Susanna Loeb, Dan Goldhaber, Douglas Staiger, Stephen Raudenbush, Grover Whitehurst, Evaluating Teachers: The Important Role of Value-Added, Brown Center on Education Policy at Brookings, 2010. http://www.brookings.edu/reports/2010/1117_evaluating_teachers.aspx
Ted Hershberg, Virginia Adams Simon and Barbara Lea Kruger, The revelations of value-added: An assessment model that measures student growth in ways that NCLB fails to do, The School Administrator, December 2004.
http://www.aasa.org/SchoolAdministratorArticle.aspx?id=9466
David Hill, He’s got your number, Teacher Magazine,
May 2000 11(8), 42–47.
http://www.edweek.org/tm/articles/2000/05/01/08sanders.h11.html
Cory Koedel and Julian R. Betts, Re-Examining the Role of Teacher Quality in the Educational Production Function, Working Paper #2007-03, National Center on Performance Initiatives, Nashville, TN, 2007.
http://economics.missouri.edu/working-papers/2007/wp0708_koedel.pdf
Daniel Koretz, Measuring Up: What Educational Testing Really Tells Us , Harvard University Press, Cambridge, Massachusetts, 2008.
E. F. Lindquist, Preliminary considerations in objective test construction, in Educational Measurement (E. F. Lindquist, ed.), American Council on Education, Washington DC, 1951.
J. R. Lockwood, Daniel McCaffrey, Laura S. Hamilton, Brian Stetcher, Vi-Nhuan Le, and Felipe Martinez, The sensitivity of value-added teacher effect estimates to different mathematics achievement measures, Journal of Educational Measurement 44(1) (2007), 47–67. http://dx.doi.org/10.1111/j.1745-3984.2007.00026.x
Daniel F. McCaffrey, Daniel Koretz, J. R. Lockwood, and Laura S. Hamilton, Evaluating Value-Added Models for Teacher Accountability, RAND Corporation, Santa Monica, CA, 2003.
http://www.rand.org/pubs/monographs/2004/RAND_MG158.pdf
Daniel F. McCaffrey, J. R. Lockwood, Daniel Koretz, Thomas A. Louis, and Laura Hamilton, Models for value-added modeling of teacher effects, Journal of Educational and Behavioral Statistics 29(1), Spring 2004, 67-101.
http://www.rand.org/pubs/reprints/2005/RAND_RP1165.pdf
RAND Research Brief, The Promise and Peril of Using Value-Added Modeling to Measure Teacher Effectiveness, Santa Monica, CA, 2004. http://www.rand.org/pubs/research_briefs/RB9050/RAND_RB9050.pdf
William L. Sanders and Sandra P. Horn, Educational Assessment Reassessed: The Usefulness of Standardized and Alternative Measures of Student Achievement as Indicators of the Assessment of Educational Outcomes, Education Policy Analysis Archives, March 3(6) (1995). http://epaa.asu.edu/ojs/article/view/649
W. Sanders, A. Saxton, and B. Horn, The Tennessee value-added assessment system: A quantitative outcomes-based approach to educational assessment, in Grading Teachers, Grading Schools: Is Student Achievement a Valid Evaluational Measure? (J. Millman, ed.), Corwin Press, Inc., Thousand Oaks, CA, 1997, pp 137–162.
-0-
Follow The Answer Sheet every day by bookmarking http://www.washingtonpost.com/blogs/answer-sheet. And for admissions advice, college news and links to campus papers, please check out our Higher Education page. Bookmark it!
Saturday, December 11, 2010
What Works in the Classroom? Ask the Students
Wow, $335M to associate two student test scores to a few characteristics. Someone is making a lot of money to help push a narrow teacher effectiveness model.
- Patricia
By SAM DILLON
Published: December 10, 2010
How useful are the views of public school students about their teachers?
Quite useful, according to preliminary results released on Friday from a $45 million research project that is intended to find new ways of distinguishing good teachers from bad.
Teachers whose students described them as skillful at maintaining classroom order, at focusing their instruction and at helping their charges learn from their mistakes are often the same teachers whose students learn the most in the course of a year, as measured by gains on standardized test scores, according to a progress report on the research.
Financed by the Bill and Melinda Gates Foundation, the two-year project involves scores of social scientists and some 3,000 teachers and their students in Charlotte, N.C.; Dallas; Denver; Hillsborough County, Fla., which includes Tampa; Memphis; New York; and Pittsburgh.
The research is part of the $335 million Gates Foundation effort to overhaul the personnel systems in those districts.
Statisticians began the effort last year by ranking all the teachers using a statistical method known as value-added modeling, which calculates how much each teacher has helped students learn based on changes in test scores from year to year.
Now researchers are looking for correlations between the value-added rankings and other measures of teacher effectiveness.
Research centering on surveys of students’ perceptions has produced some clear early results.
Thousands of students have filled out confidential questionnaires about the learning environment that their teachers create. After comparing the students’ ratings with teachers’ value-added scores, researchers have concluded that there is quite a bit of agreement.
Classrooms where a majority of students said they agreed with the statement, “Our class stays busy and doesn’t waste time,” tended to be led by teachers with high value-added scores, the report said.
The same was true for teachers whose students agreed with the statements, “In this class, we learn to correct our mistakes,” and, “My teacher has several good ways to explain each topic that we cover in this class.”
The questionnaires were developed by Ronald Ferguson, a Harvard researcher who has been refining student surveys for more than a decade.
Few of the nation’s 15,000 public school districts systematically question students about their classroom experiences, in contrast to American colleges, many of which collect annual student evaluations to improve instruction, Dr. Ferguson said.
“Kids know effective teaching when they experience it,” he said.
“As a nation, we’ve wasted what students know about their own classroom experiences instead of using that knowledge to inform school reform efforts.”
Until recently, teacher evaluations were little more than a formality in most school systems, with the vast majority of instructors getting top ratings, often based on a principal’s superficial impressions.
But now some 20 states are overhauling their evaluation systems, and many policymakers involved in those efforts have been asking the Gates Foundation for suggestions on what measures of teacher effectiveness to use, said Vicki L. Phillips, a director of education at the foundation.
One notable early finding, Ms. Phillips said, is that teachers who incessantly drill their students to prepare for standardized tests tend to have lower value-added learning gains than those who simply work their way methodically through the key concepts of literacy and mathematics.
Teachers whose students agreed with the statement, “We spend a lot of time in this class practicing for the state test,” tended to make smaller gains on those exams than other teachers.
“Teaching to the test makes your students do worse on the tests,” Ms. Phillips said. “It turns out all that ‘drill and kill’ isn’t helpful.”
- Patricia
By SAM DILLON
Published: December 10, 2010
How useful are the views of public school students about their teachers?
Quite useful, according to preliminary results released on Friday from a $45 million research project that is intended to find new ways of distinguishing good teachers from bad.
Teachers whose students described them as skillful at maintaining classroom order, at focusing their instruction and at helping their charges learn from their mistakes are often the same teachers whose students learn the most in the course of a year, as measured by gains on standardized test scores, according to a progress report on the research.
Financed by the Bill and Melinda Gates Foundation, the two-year project involves scores of social scientists and some 3,000 teachers and their students in Charlotte, N.C.; Dallas; Denver; Hillsborough County, Fla., which includes Tampa; Memphis; New York; and Pittsburgh.
The research is part of the $335 million Gates Foundation effort to overhaul the personnel systems in those districts.
Statisticians began the effort last year by ranking all the teachers using a statistical method known as value-added modeling, which calculates how much each teacher has helped students learn based on changes in test scores from year to year.
Now researchers are looking for correlations between the value-added rankings and other measures of teacher effectiveness.
Research centering on surveys of students’ perceptions has produced some clear early results.
Thousands of students have filled out confidential questionnaires about the learning environment that their teachers create. After comparing the students’ ratings with teachers’ value-added scores, researchers have concluded that there is quite a bit of agreement.
Classrooms where a majority of students said they agreed with the statement, “Our class stays busy and doesn’t waste time,” tended to be led by teachers with high value-added scores, the report said.
The same was true for teachers whose students agreed with the statements, “In this class, we learn to correct our mistakes,” and, “My teacher has several good ways to explain each topic that we cover in this class.”
The questionnaires were developed by Ronald Ferguson, a Harvard researcher who has been refining student surveys for more than a decade.
Few of the nation’s 15,000 public school districts systematically question students about their classroom experiences, in contrast to American colleges, many of which collect annual student evaluations to improve instruction, Dr. Ferguson said.
“Kids know effective teaching when they experience it,” he said.
“As a nation, we’ve wasted what students know about their own classroom experiences instead of using that knowledge to inform school reform efforts.”
Until recently, teacher evaluations were little more than a formality in most school systems, with the vast majority of instructors getting top ratings, often based on a principal’s superficial impressions.
But now some 20 states are overhauling their evaluation systems, and many policymakers involved in those efforts have been asking the Gates Foundation for suggestions on what measures of teacher effectiveness to use, said Vicki L. Phillips, a director of education at the foundation.
One notable early finding, Ms. Phillips said, is that teachers who incessantly drill their students to prepare for standardized tests tend to have lower value-added learning gains than those who simply work their way methodically through the key concepts of literacy and mathematics.
Teachers whose students agreed with the statement, “We spend a lot of time in this class practicing for the state test,” tended to make smaller gains on those exams than other teachers.
“Teaching to the test makes your students do worse on the tests,” Ms. Phillips said. “It turns out all that ‘drill and kill’ isn’t helpful.”
Monday, November 08, 2010
HISD looks at how to grade teachers
New criteria for evaluation will likely include kids' test scores
By ERICKA MELLON | HOUSTON CHRONICLE
Nov. 8, 2010
About 99 percent of teachers in the Houston school district receive satisfactory job evaluations, with their students' academic success barely playing a factor.
That's likely to change next year. Houston ISD Superintendent Terry Grier and the school board are developing a new appraisal system for the district's 13,000 teachers. The process will affect who's promoted and fired, as well as the training teachers receive.
The nitty-gritty of evaluations might seem incidental, but the topic has become a cornerstone of public education reform, with President Barack Obama calling on schools to hold teachers more accountable for their students' performance.
Ann Best, the Houston Independent School District's chief human resources officer, acknowledged that teachers are nervous about the changes, but she said they, along with parents and administrators, are represented on committees developing the new criteria.
"This effort is designed to ensure that we have effective teachers in all of our classrooms," Best said, "and part of having effective teachers is giving teachers really clear feedback about their performance. What we should expect coming out of this initiative is more effective teachers in our schools."
Still, HISD's largest teachers group is threatening a possible challenge at the Texas Education Agency.
Texas law says any teacher evaluation systems that veer from the state-approved model must be "developed by" school-based committees and a district-level committee.
Gayle Fallon, the president of the Houston Federation of Teachers, questioned whether HISD officials are letting the committees drive the process as required.
"What we're getting back from the teachers is they feel they're window dressing — that all the decisions have been made," Fallon said. "If (the new evaluation) isn't really developed by them and accepted by them, then we will promise a challenge at the state."
Best and Dan Weisberg of the New Teacher Project, a nonprofit helping HISD with the evaluation process, rejected the claim that teachers are being sidelined.
"There is no secret tool in any file that someone could say it's already been developed," Best said.
Sticking point
Committees have met over the last two months and, according to a presentation made to the school board, have agreed on three broad categories for the ratings: student performance, instructional practice and professional expectations.
The main sticking point, by all accounts, will be how to define student performance. The current appraisal tool focuses mostly on principals' observations of teachers — not on student test scores or other data.
The HISD board approved in February a policy ordering future evaluations to include an analysis of test score data called value-added, but trustees did not set any details, such as how much the scores should count in teacher ratings.
Value-added is meant to be a statistical measure of teacher effectiveness, looking at whether students' test scores met, exceeded or fell below expectations based on each child's past performance.
Fallon, whose union represents more than 7,000 educators, said it would be a "deal breaker" if the evaluations included the value-added data that HISD uses to determine teachers' bonuses. She and Chuck Robinson, who heads the district's second-largest teacher group, repeatedly have said the formula developed by North Carolina statistician Bill Sanders is confusing and unreliable.
HISD model criticized
HISD's general counsel, Elneita Hutchins-Taylor, declined to speculate what would happen if the committees decide against making value-added data part of the evaluation as the board policy mandates.
"It's premature for us to comment on what will or will not be in there," she said. "Obviously the board policy exists. I think it's a policy of an aspirational nature."
Randi Weingarten, the president of the American Federation of Teachers, met with members of the Houston chapter in October to encourage them to work with the district to adopt an "evaluation system that does look at both teacher observations and student learning."
Weingarten made headlines in January when she said she favored including student test scores and other measures of student performance in evaluations. But she has blasted HISD's value-added model.
"The more we look at it, the more concerned we are it's not ready for prime time," she said. "It's like a weather forecast."
A study of four years of teachers' evaluations in HISD found, as in other districts nationwide, that few teachers got poor marks. Of all the ratings, 62 percent were "exceeds expectations" and 37 percent were "proficient." The other 1 percent were "below expectations" or "unsatisfactory."
Best said she expects the new appraisal to better differentiate among teachers.
"To me it's impossible that nearly 100 percent of our staff are performing at the highest levels because there are varying levels of performance in any profession," she said.
Robinson, who leads the Congress of Houston Teachers, said it's more important to train principals or other evaluators than to change the appraisal document.
"It's the way it's applied and used and the kinds of instructional leadership you have in a building," he said, adding that teachers expect the new evaluation will lead to more firings.
"There's just a lot of fear, a lot of insecurity and a lot of mistrust - and not from people who are lazy, indifferent people who just want to collect a paycheck."
The focus of the evaluations, Best said, will be on helping teachers improve their skills.
"What we want to assure people of is the first step is not termination," she said. "The first step is having accurate data to inform a plan for your development. Then we have to assess the extent to which the implementation of that plan actually works."
By ERICKA MELLON | HOUSTON CHRONICLE
Nov. 8, 2010
About 99 percent of teachers in the Houston school district receive satisfactory job evaluations, with their students' academic success barely playing a factor.
That's likely to change next year. Houston ISD Superintendent Terry Grier and the school board are developing a new appraisal system for the district's 13,000 teachers. The process will affect who's promoted and fired, as well as the training teachers receive.
The nitty-gritty of evaluations might seem incidental, but the topic has become a cornerstone of public education reform, with President Barack Obama calling on schools to hold teachers more accountable for their students' performance.
Ann Best, the Houston Independent School District's chief human resources officer, acknowledged that teachers are nervous about the changes, but she said they, along with parents and administrators, are represented on committees developing the new criteria.
"This effort is designed to ensure that we have effective teachers in all of our classrooms," Best said, "and part of having effective teachers is giving teachers really clear feedback about their performance. What we should expect coming out of this initiative is more effective teachers in our schools."
Still, HISD's largest teachers group is threatening a possible challenge at the Texas Education Agency.
Texas law says any teacher evaluation systems that veer from the state-approved model must be "developed by" school-based committees and a district-level committee.
Gayle Fallon, the president of the Houston Federation of Teachers, questioned whether HISD officials are letting the committees drive the process as required.
"What we're getting back from the teachers is they feel they're window dressing — that all the decisions have been made," Fallon said. "If (the new evaluation) isn't really developed by them and accepted by them, then we will promise a challenge at the state."
Best and Dan Weisberg of the New Teacher Project, a nonprofit helping HISD with the evaluation process, rejected the claim that teachers are being sidelined.
"There is no secret tool in any file that someone could say it's already been developed," Best said.
Sticking point
Committees have met over the last two months and, according to a presentation made to the school board, have agreed on three broad categories for the ratings: student performance, instructional practice and professional expectations.
The main sticking point, by all accounts, will be how to define student performance. The current appraisal tool focuses mostly on principals' observations of teachers — not on student test scores or other data.
The HISD board approved in February a policy ordering future evaluations to include an analysis of test score data called value-added, but trustees did not set any details, such as how much the scores should count in teacher ratings.
Value-added is meant to be a statistical measure of teacher effectiveness, looking at whether students' test scores met, exceeded or fell below expectations based on each child's past performance.
Fallon, whose union represents more than 7,000 educators, said it would be a "deal breaker" if the evaluations included the value-added data that HISD uses to determine teachers' bonuses. She and Chuck Robinson, who heads the district's second-largest teacher group, repeatedly have said the formula developed by North Carolina statistician Bill Sanders is confusing and unreliable.
HISD model criticized
HISD's general counsel, Elneita Hutchins-Taylor, declined to speculate what would happen if the committees decide against making value-added data part of the evaluation as the board policy mandates.
"It's premature for us to comment on what will or will not be in there," she said. "Obviously the board policy exists. I think it's a policy of an aspirational nature."
Randi Weingarten, the president of the American Federation of Teachers, met with members of the Houston chapter in October to encourage them to work with the district to adopt an "evaluation system that does look at both teacher observations and student learning."
Weingarten made headlines in January when she said she favored including student test scores and other measures of student performance in evaluations. But she has blasted HISD's value-added model.
"The more we look at it, the more concerned we are it's not ready for prime time," she said. "It's like a weather forecast."
A study of four years of teachers' evaluations in HISD found, as in other districts nationwide, that few teachers got poor marks. Of all the ratings, 62 percent were "exceeds expectations" and 37 percent were "proficient." The other 1 percent were "below expectations" or "unsatisfactory."
Best said she expects the new appraisal to better differentiate among teachers.
"To me it's impossible that nearly 100 percent of our staff are performing at the highest levels because there are varying levels of performance in any profession," she said.
Robinson, who leads the Congress of Houston Teachers, said it's more important to train principals or other evaluators than to change the appraisal document.
"It's the way it's applied and used and the kinds of instructional leadership you have in a building," he said, adding that teachers expect the new evaluation will lead to more firings.
"There's just a lot of fear, a lot of insecurity and a lot of mistrust - and not from people who are lazy, indifferent people who just want to collect a paycheck."
The focus of the evaluations, Best said, will be on helping teachers improve their skills.
"What we want to assure people of is the first step is not termination," she said. "The first step is having accurate data to inform a plan for your development. Then we have to assess the extent to which the implementation of that plan actually works."
Monday, December 28, 2009
School-label method changing
State education officials looking for more-precise way to measure progress
by Pat Kossan | The Arizona Republic
Dec. 13, 2009
The way Arizona decides if a school is performing or failing is likely to change dramatically by 2011.
The new method would be more precise than any used previously and measure how successfully a school pushes its students - average or gifted, rich or poor - to learn more from year to year, state officials said.
It could mean additions and deletions on the list of best-performing schools in Arizona and would give parents better information about the quality of teaching at a school. It also could be used to determine which teachers a school retains and how much a teacher is paid.
"The way that (student academic) growth was measured in 2003 when I took office was sufficiently problematic that it would not be fair to judge teachers on that data," said Tom Horne, Arizona superintendent of public instruction. "Now the science and technology has developed to a point where I think you can do that fairly."
The State Board of Education will have the final say on when and how the state puts the new method to use. That vote is expected in the spring.
Students' progress impacts school labels
In 2002, the state began publicly labeling schools based on their student performance. As its chief measurement, the state uses year-to-year gains in the overall percentage of students at a school passing the AIMS exam.
If approved, beginning in 2011, measuring the success of an overall school's population would be replaced by measuring the year-to-year improvement of each of its student's AIMS scores, whether or not it's a passing score.
Individual student growth would become 50 percent of the way a school earns one of the state's six labels: excelling, highly performing, performing plus, performing, underperforming or failing. The state will use both the proposed new measurement and the old measurement to label schools in 2010 so educators and policy makers can see the difference in how their school would be labeled.
The new measurement is most commonly referred to as the "value-added" method and here's how it works:
• A student's AIMS score is measured on a scale of 200 at the bottom for Grade 3 and 900 at the top for the high-school exam.
• The state already has the ability to determine improvement in each student's AIMS scale score over previous years' scores. The new measurement allows the state to determine if a student's year-to-year progress matches progress made by other Arizona students who had similar scores last year and the year before.
• On a micro level, this new method can help teachers and parents determine if a student's learning is keeping pace or outpacing their true academic peers, whether the student is scoring in the 300s or in the 700s. On a macro level, it can determine if students in a classroom, a school or a district are outpacing similar students, keeping pace with them or falling behind.
New school search engine for parents of students
By 2011, parents would have access to a new search engine that uses the new measurement to compare schools within their district or neighborhood.
Horne said the visual charts that accompany the new method would make it easier for parents to find out what they really want to know: Are the teachers at a school capable of moving all students ahead in their learning and is their particular child moving up?
"It's a major breakthrough for parents to see exactly what is happening with their school and other schools in the neighborhood," Horne said.
Right now, the state is test-driving the new technology. Education officials hope to get individual student achievement graphs in the hands of every teacher, perhaps as soon as the start of school in January, to help them target their teaching to student needs.
"A teacher can also show them to a parent and the parent really understands that the kid has a lot of work to do to pass AIMS, which is very often the case, probably more than people realize," said Rebecca Gau, a researcher with the Arizona Charter School Association who helped to bring the new measurement to the state.
Teacher performance linked to student data
Colorado was the first state to use the newest growth measurement, but it's now being considered in many states, mostly as a way to measure a teacher's performance. The Obama administration is pushing states to find a fair way to link teacher pay to student test scores and there are big grants on the table for states willing to follow its lead.
A school or district, even the state, could use the data as part of each teacher's professional evaluation to help determine retention, training needs and pay, Horne said.
Horne likes the idea of using student growth to gauge teacher performance, and this method is the fairest he has seen.
"The teacher who made a lot of growth with poor kids would still show better than a teacher who made little growth with richer kids," Horne said.
It's a tool that researchers are continuing to refine and is worth exploring as a teacher-performance indicator, said Andrew Morrill, vice president of Arizona Education Association, the state's teachers union. Morrill cautions that teachers cannot be fairly evaluated using only one measurement.
It could be combined with other criteria, such as a teacher's willingness to continue pursuing additional education and to work with other teachers to develop effective lesson plans, he said.
"You still have the question of what and how many data indicators you're going to use in a fairly complex calculation," Morrill said.
The Arizona State Board of Education received an explanation of the new approach Dec. 7. The board is expected to convene a study session on the new method in the spring before it votes on the proposed change.
by Pat Kossan | The Arizona Republic
Dec. 13, 2009
The way Arizona decides if a school is performing or failing is likely to change dramatically by 2011.
The new method would be more precise than any used previously and measure how successfully a school pushes its students - average or gifted, rich or poor - to learn more from year to year, state officials said.
It could mean additions and deletions on the list of best-performing schools in Arizona and would give parents better information about the quality of teaching at a school. It also could be used to determine which teachers a school retains and how much a teacher is paid.
"The way that (student academic) growth was measured in 2003 when I took office was sufficiently problematic that it would not be fair to judge teachers on that data," said Tom Horne, Arizona superintendent of public instruction. "Now the science and technology has developed to a point where I think you can do that fairly."
The State Board of Education will have the final say on when and how the state puts the new method to use. That vote is expected in the spring.
Students' progress impacts school labels
In 2002, the state began publicly labeling schools based on their student performance. As its chief measurement, the state uses year-to-year gains in the overall percentage of students at a school passing the AIMS exam.
If approved, beginning in 2011, measuring the success of an overall school's population would be replaced by measuring the year-to-year improvement of each of its student's AIMS scores, whether or not it's a passing score.
Individual student growth would become 50 percent of the way a school earns one of the state's six labels: excelling, highly performing, performing plus, performing, underperforming or failing. The state will use both the proposed new measurement and the old measurement to label schools in 2010 so educators and policy makers can see the difference in how their school would be labeled.
The new measurement is most commonly referred to as the "value-added" method and here's how it works:
• A student's AIMS score is measured on a scale of 200 at the bottom for Grade 3 and 900 at the top for the high-school exam.
• The state already has the ability to determine improvement in each student's AIMS scale score over previous years' scores. The new measurement allows the state to determine if a student's year-to-year progress matches progress made by other Arizona students who had similar scores last year and the year before.
• On a micro level, this new method can help teachers and parents determine if a student's learning is keeping pace or outpacing their true academic peers, whether the student is scoring in the 300s or in the 700s. On a macro level, it can determine if students in a classroom, a school or a district are outpacing similar students, keeping pace with them or falling behind.
New school search engine for parents of students
By 2011, parents would have access to a new search engine that uses the new measurement to compare schools within their district or neighborhood.
Horne said the visual charts that accompany the new method would make it easier for parents to find out what they really want to know: Are the teachers at a school capable of moving all students ahead in their learning and is their particular child moving up?
"It's a major breakthrough for parents to see exactly what is happening with their school and other schools in the neighborhood," Horne said.
Right now, the state is test-driving the new technology. Education officials hope to get individual student achievement graphs in the hands of every teacher, perhaps as soon as the start of school in January, to help them target their teaching to student needs.
"A teacher can also show them to a parent and the parent really understands that the kid has a lot of work to do to pass AIMS, which is very often the case, probably more than people realize," said Rebecca Gau, a researcher with the Arizona Charter School Association who helped to bring the new measurement to the state.
Teacher performance linked to student data
Colorado was the first state to use the newest growth measurement, but it's now being considered in many states, mostly as a way to measure a teacher's performance. The Obama administration is pushing states to find a fair way to link teacher pay to student test scores and there are big grants on the table for states willing to follow its lead.
A school or district, even the state, could use the data as part of each teacher's professional evaluation to help determine retention, training needs and pay, Horne said.
Horne likes the idea of using student growth to gauge teacher performance, and this method is the fairest he has seen.
"The teacher who made a lot of growth with poor kids would still show better than a teacher who made little growth with richer kids," Horne said.
It's a tool that researchers are continuing to refine and is worth exploring as a teacher-performance indicator, said Andrew Morrill, vice president of Arizona Education Association, the state's teachers union. Morrill cautions that teachers cannot be fairly evaluated using only one measurement.
It could be combined with other criteria, such as a teacher's willingness to continue pursuing additional education and to work with other teachers to develop effective lesson plans, he said.
"You still have the question of what and how many data indicators you're going to use in a fairly complex calculation," Morrill said.
The Arizona State Board of Education received an explanation of the new approach Dec. 7. The board is expected to convene a study session on the new method in the spring before it votes on the proposed change.
Value-added education in the race to the top
David Davenport | SF Gate
Sunday, November 29, 2009
Bill Clinton may have invented triangulation - the art of finding a "third way" out of a policy dilemma - but U.S. Secretary of Education Arne Duncan is practicing it to make desperately needed improvements in K-12 education. Unfortunately, his promotion of value-added education through "Race to the Top" grants to states could be thrown under the bus by powerful teachers' unions that view reforms more for how they affect pay and job security than whether they improve student learning.
The traditional view of education holds that it is more process than product. Educators design a process, hire teachers and administrators to run it, put students through it and consider it a success. The focus is on the inputs - how much can we spend, what curriculum shall we use, what class size is best - with very little on measuring outputs, whether students actually learn. The popular surveys of America's best schools and colleges reinforce this, measuring resources and reputation, not results. As they say, Harvard University has good graduates because it admits strong applicants, not necessarily because of what happens in the educational process.
In the last decade, the federal No Child Left Behind program has ushered in a new era of testing and accountability, seeking to shift the focus to outcomes. But this more businesslike approach does not always fit a people-centered field such as education. Some students test well, and others do not. Some schools serve a disproportionately high number of students who are not well prepared. Even in good schools, a system driven by testing and accountability incentivizes teaching to the test, neglecting other important and interesting ways to engage and educate students. As a result, policymakers and educators have been ambivalent, at best, about the No Child Left Behind regime.
An interesting middle ground has emerged: value-added education. In this approach, already employed in several states, the focus is on how much learning each student does each year. Testing is still done, but it is for the purpose of measuring annual gains by individual students. After all, isn't that what any people-servicing profession should seek to do: improve the lot of the patient, client or student?
The controversy develops because the value-added approach reveals which schools and teachers are doing the best job of improving student learning. The unions are stirred up about this because, as Secretary Duncan has said, currently "zero percent (of teacher evaluation) is based on student achievement." But, of course, that is part of the problem, that we have tolerated a system that does not measure and reward teacher performance.
Education expert Eric Hanushek of the Hoover Institution has estimated that merely replacing the nation's worst 6 to 10 percent of teachers with average teachers would make a huge difference in the quality of American education. Value-added testing would begin to allow such judgments to be made, which is why Gov. Arnold Schwarzenegger recently signed two bills making value-added data available for teacher and school evaluation and qualifying California for some of the $4.35 billion in "Race to the Top" federal grants.
Let's face it - the time to do nothing about our declining schools has passed. Value-added education offers a path out of the old process versus testing and accountability dilemma. To paraphrase Yogi Berra: We've come to a fork in the road, let's take it.
Sunday, November 29, 2009
Bill Clinton may have invented triangulation - the art of finding a "third way" out of a policy dilemma - but U.S. Secretary of Education Arne Duncan is practicing it to make desperately needed improvements in K-12 education. Unfortunately, his promotion of value-added education through "Race to the Top" grants to states could be thrown under the bus by powerful teachers' unions that view reforms more for how they affect pay and job security than whether they improve student learning.
The traditional view of education holds that it is more process than product. Educators design a process, hire teachers and administrators to run it, put students through it and consider it a success. The focus is on the inputs - how much can we spend, what curriculum shall we use, what class size is best - with very little on measuring outputs, whether students actually learn. The popular surveys of America's best schools and colleges reinforce this, measuring resources and reputation, not results. As they say, Harvard University has good graduates because it admits strong applicants, not necessarily because of what happens in the educational process.
In the last decade, the federal No Child Left Behind program has ushered in a new era of testing and accountability, seeking to shift the focus to outcomes. But this more businesslike approach does not always fit a people-centered field such as education. Some students test well, and others do not. Some schools serve a disproportionately high number of students who are not well prepared. Even in good schools, a system driven by testing and accountability incentivizes teaching to the test, neglecting other important and interesting ways to engage and educate students. As a result, policymakers and educators have been ambivalent, at best, about the No Child Left Behind regime.
An interesting middle ground has emerged: value-added education. In this approach, already employed in several states, the focus is on how much learning each student does each year. Testing is still done, but it is for the purpose of measuring annual gains by individual students. After all, isn't that what any people-servicing profession should seek to do: improve the lot of the patient, client or student?
The controversy develops because the value-added approach reveals which schools and teachers are doing the best job of improving student learning. The unions are stirred up about this because, as Secretary Duncan has said, currently "zero percent (of teacher evaluation) is based on student achievement." But, of course, that is part of the problem, that we have tolerated a system that does not measure and reward teacher performance.
Education expert Eric Hanushek of the Hoover Institution has estimated that merely replacing the nation's worst 6 to 10 percent of teachers with average teachers would make a huge difference in the quality of American education. Value-added testing would begin to allow such judgments to be made, which is why Gov. Arnold Schwarzenegger recently signed two bills making value-added data available for teacher and school evaluation and qualifying California for some of the $4.35 billion in "Race to the Top" federal grants.
Let's face it - the time to do nothing about our declining schools has passed. Value-added education offers a path out of the old process versus testing and accountability dilemma. To paraphrase Yogi Berra: We've come to a fork in the road, let's take it.
Wednesday, December 23, 2009
Superintendent spreads the gospel of 'value-added' teacher evaluations
Interesting sidenote mentioned in this article: "FOR THE RECORD:
Teacher evaluations: An article in some editions of the Oct. 18 Section A about evaluating teacher performance reported that the San Diego teachers unions spent nearly $400,000 in this fall's school board elections. The elections were held in 2008"
What needs to really be critically examined with value-added assessments is how they serve schools, sometimes at the expense of students meeting their goals (ideally for being college ready).
-Patricia
In Tenn. and N.C., Terry Grier adopted and expanded a statistical method of tracking student progress. Union resistance scuttled more modest efforts in San Diego, mirroring a brewing national debate.
By Jason Felch and Jason Song | LA Times
October 18, 2009
When Terry Grier was hired to run the San Diego Unified School District in January 2008, he hoped to bring with him a revolutionary tool that had never been tried in a large California school system.
Its name -- "value-added" -- sounded innocuous enough. But this novel number-crunching approach threatened to upend many traditional notions of what worked and what didn't in the nation's classrooms.
Rather than using tests to take a snapshot of overall student achievement, it used scores to track each pupil's academic progress from year to year. What made it incendiary, however, was its potential to single out the best and the worst teachers in a nation that currently gives virtually all of them a passing grade.
In previous jobs in the South, Grier had used the method as a basis for removing underperforming principals, denying ineffective teachers tenure and rewarding the best educators with additional pay.
In California, where powerful teachers unions have been especially protective of tenure and resistant to merit pay, Grier had a more modest goal: to find out if students in the district's poorest schools had equal access to effective instructors.
Still, it proved radioactive to San Diego's teachers union. Like many unions across the country, it saw the approach as a flawed instrument, a Trojan horse for introducing merit pay and a threat to hard-won employment protections.
After nearly two years of grinding battles with the union and school board on this and other issues, Grier recently left for Houston, where the district uses value-added results as a basis for teacher bonuses.
The opposition in San Diego, Grier said mildly, was "more entrenched than I thought it would be."
His fight there offers a preview of a debate that is about to engulf the nation's schools.
The Obama administration has made value-added a pillar of its school-reform efforts, including the $4.35-billion federal grant program known as Race to the Top, which requires states to link student scores to teachers.
The administration's endorsement thrust to the fore a 20-year-old idea that has long been confined to academic circles and a handful of states and school districts, including Chicago's, where it was championed by now-Secretary of Education Arne Duncan.
Gov. Arnold Schwarzenegger this month signed into law two measures that put California on the path to measuring student growth with the value-added method.
But most districts thus far have just begun to talk about using the method in teacher evaluations.
Value-added analysis promises to address one of public education's central conundrums: On one hand, research shows that effective teachers are the single most important factor in improving student performance. On the other, most states use subjective evaluation systems -- based on occasional classroom visits by administrators -- that give nearly all teachers a satisfactory rating.
"Zero percent [of teacher evaluation] is based on student achievement," Duncan said in a recent interview. "That's a problem."
Protection of tenure
With no objective measure of success, firing a tenured teacher for incompetence is nearly impossible in many districts.
Earlier this year, a Times investigation found that jettisoning a tenured teacher solely because he or she can't teach is rare. In 80% of terminations upheld by the state in the last 15 years, classroom performance was not even a factor.
"Allowing ineffective teachers to remain in the classroom is literally dragging down the nation," said education researcher Eric Hanushek of Stanford University's Hoover Institution.
In a forthcoming paper, Hanushek estimates that replacing the nation's worst 6% to 10% of instructors with merely average teachers would propel the United States from its below-average level into the ranks of the world's top five educational systems.
So far, nobody is proposing even that. Most proponents hope to use value-added scores as one of several measures to identify teachers who need help.
Still, the potential of the value-added method has caught the eye of policy wonks, educational reformers and billionaire philanthropists such as Eli Broad, Bill Gates and brothers Lowell and Michael Milken.
But value-added has met with staunch opposition from teachers unions, which often argue that standardized tests are not a good measure of students' performance, let alone teachers'.
They also cite some experts' concerns that the statistical methods of the value-added approach may rest on ill-founded assumptions.
For example, some research shows it's easier to achieve gains with certain types of students. The result could be the unfair rewarding or punishing of teachers based on the students they are assigned.
"The danger is that too much weight will be put on this," said Helen Ladd, a Duke University professor, who urges more study. "I think the policymakers are ahead of the statisticians here."
But cracks in union opposition are beginning to show. Randi Weingarten, president of the 1.4-million-member American Federation of Teachers, once fought efforts to use the value-added approach in New York, and in August cautioned the Obama administration against "over-reliance on an unproven idea."
But this month, her union announced innovation awards for several local unions that plan to adopt value-added measures as part of a new teacher evaluation system even though Weingarten said she still has reservations.
"What we've realized is, if we don't do it by ourselves, no one is going to listen to us," she said.
An early skeptic
Grier was skeptical when he first encountered value-added results in the early 1990s as superintendent of schools in Williamson County, Tenn., just south of Nashville.
Though it's one of the wealthiest districts in the country, many students in its highest-achieving schools were not making much annual progress. Students in some of the poorer schools, meanwhile, were making remarkable gains, learning two years' worth of material in a single grade.
Then, at an academic conference, Grier was won over after seeing a presentation by William Sanders, then a University of Tennessee statistician.
Sanders described how he had developed a statistical tool that, in essence, took into account three years of a student's test scores to estimate the child's academic trajectory. Then he compared it with the student's actual performance in the fourth year. The difference between the expected growth and actual growth became known as "the teacher effect."
The approach overcame the Achilles' heel of traditional achievement tests, which critics say reflect socioeconomic status more than learning. They judge this year's students against last year's, ignoring potential differences between the two groups.
With the value-added method, students are compared to themselves from year to year, so the results are not skewed by income levels, parental involvement, race or gender.
"A teacher or principal has no responsibility for what's happened to the kids in the past," Sanders said. "What they do have responsibility for is the student's rate of progress. This is what was different about what we did."
Sanders' research found that there is a huge disparity in the effectiveness of teachers, and that good teachers could make an enormous difference in a student's learning, regardless of the child's background.
Subsequent research largely confirmed his findings, and in the process many orthodoxies of education policy were called into question. In particular, studies indicated that the traditional measures of a teacher's quality -- such as years of experience, credentials and education -- have little bearing on his or her effectiveness.
Grier was convinced. "I thought it could be a game-changer," he said.
To challenge students at his wealthy schools, Grier pushed principals to focus on growth rather than hitting fixed scores. He used the data to assign teachers to students who matched their strengths. He denied tenure to at least four teachers whose students did not show growth for several years, and reassigned some low-performing principals while encouraging others to retire.
At his next job in Guilford County, N.C., Grier hired Sanders to improve on the district's existing value-added program.
The goal of the program, dubbed Mission Possible, was to attract effective teachers to low-performing schools. To do this, Grier offered highly effective teachers up to $10,000 to relocate. They got up to $4,000 more if students' test scores rose.
Many educators protested the plan, but, with no collective bargaining in North Carolina, the board approved Mission Possible in 2006.
By 2007, teacher turnover at some of the area's lowest-performing schools had been cut by almost half, according to a recent independent study. But the program hasn't significantly changed student achievement -- at least, not yet.
"It's a mixed bag," said board member Alan Duncan. The real test, he said, will be a few years from now, when the board expects the changes to show up in test scores.
North Carolina has since adopted value-added statewide.
Teacher resistance
Grier said he knew that Mission Possible wasn't possible in San Diego. "The board made it clear they were not interested in a merit-pay program," he said.
Instead, Grier wanted to find out how the district's high-performing teachers were distributed. He suspected that many were clustered north of Interstate 8, in the wealthier part of the city.
He lured Sanders, now with a for-profit data analysis company, to visit San Diego with the promise of golf at Torrey Pines. The statistician's presentation persuaded the school board to approve a one-year, $80,000 contract with Sanders' company -- but limited its use to identifying students in need of extra help.
Union officials decried the move, saying the money would have been better spent hiring another teacher. They suspected Grier was trying to move the district toward merit pay. "We knew his history," said union President Camille Zombro.
She and other union members interrupted a school board meeting in June to present the board with a petition complaining about Grier's top-down management style and calling for his removal.
"He looked good in a suit, but he wasn't willing to collaborate," Zombro said.
Relations with the administrators union also soured when Grier tried to use value-added scores as a component of principal evaluation, a plan that was eventually rejected.
The board did not renew Sanders' contract this fall and never made public the results of his analysis.
Grier took over as Houston's superintendent in September. Reached by telephone recently, he seemed invigorated.
"They do things differently out here," he said. "It's a breath of fresh air."
Teacher evaluations: An article in some editions of the Oct. 18 Section A about evaluating teacher performance reported that the San Diego teachers unions spent nearly $400,000 in this fall's school board elections. The elections were held in 2008"
What needs to really be critically examined with value-added assessments is how they serve schools, sometimes at the expense of students meeting their goals (ideally for being college ready).
-Patricia
In Tenn. and N.C., Terry Grier adopted and expanded a statistical method of tracking student progress. Union resistance scuttled more modest efforts in San Diego, mirroring a brewing national debate.
By Jason Felch and Jason Song | LA Times
October 18, 2009
When Terry Grier was hired to run the San Diego Unified School District in January 2008, he hoped to bring with him a revolutionary tool that had never been tried in a large California school system.
Its name -- "value-added" -- sounded innocuous enough. But this novel number-crunching approach threatened to upend many traditional notions of what worked and what didn't in the nation's classrooms.
Rather than using tests to take a snapshot of overall student achievement, it used scores to track each pupil's academic progress from year to year. What made it incendiary, however, was its potential to single out the best and the worst teachers in a nation that currently gives virtually all of them a passing grade.
In previous jobs in the South, Grier had used the method as a basis for removing underperforming principals, denying ineffective teachers tenure and rewarding the best educators with additional pay.
In California, where powerful teachers unions have been especially protective of tenure and resistant to merit pay, Grier had a more modest goal: to find out if students in the district's poorest schools had equal access to effective instructors.
Still, it proved radioactive to San Diego's teachers union. Like many unions across the country, it saw the approach as a flawed instrument, a Trojan horse for introducing merit pay and a threat to hard-won employment protections.
After nearly two years of grinding battles with the union and school board on this and other issues, Grier recently left for Houston, where the district uses value-added results as a basis for teacher bonuses.
The opposition in San Diego, Grier said mildly, was "more entrenched than I thought it would be."
His fight there offers a preview of a debate that is about to engulf the nation's schools.
The Obama administration has made value-added a pillar of its school-reform efforts, including the $4.35-billion federal grant program known as Race to the Top, which requires states to link student scores to teachers.
The administration's endorsement thrust to the fore a 20-year-old idea that has long been confined to academic circles and a handful of states and school districts, including Chicago's, where it was championed by now-Secretary of Education Arne Duncan.
Gov. Arnold Schwarzenegger this month signed into law two measures that put California on the path to measuring student growth with the value-added method.
But most districts thus far have just begun to talk about using the method in teacher evaluations.
Value-added analysis promises to address one of public education's central conundrums: On one hand, research shows that effective teachers are the single most important factor in improving student performance. On the other, most states use subjective evaluation systems -- based on occasional classroom visits by administrators -- that give nearly all teachers a satisfactory rating.
"Zero percent [of teacher evaluation] is based on student achievement," Duncan said in a recent interview. "That's a problem."
Protection of tenure
With no objective measure of success, firing a tenured teacher for incompetence is nearly impossible in many districts.
Earlier this year, a Times investigation found that jettisoning a tenured teacher solely because he or she can't teach is rare. In 80% of terminations upheld by the state in the last 15 years, classroom performance was not even a factor.
"Allowing ineffective teachers to remain in the classroom is literally dragging down the nation," said education researcher Eric Hanushek of Stanford University's Hoover Institution.
In a forthcoming paper, Hanushek estimates that replacing the nation's worst 6% to 10% of instructors with merely average teachers would propel the United States from its below-average level into the ranks of the world's top five educational systems.
So far, nobody is proposing even that. Most proponents hope to use value-added scores as one of several measures to identify teachers who need help.
Still, the potential of the value-added method has caught the eye of policy wonks, educational reformers and billionaire philanthropists such as Eli Broad, Bill Gates and brothers Lowell and Michael Milken.
But value-added has met with staunch opposition from teachers unions, which often argue that standardized tests are not a good measure of students' performance, let alone teachers'.
They also cite some experts' concerns that the statistical methods of the value-added approach may rest on ill-founded assumptions.
For example, some research shows it's easier to achieve gains with certain types of students. The result could be the unfair rewarding or punishing of teachers based on the students they are assigned.
"The danger is that too much weight will be put on this," said Helen Ladd, a Duke University professor, who urges more study. "I think the policymakers are ahead of the statisticians here."
But cracks in union opposition are beginning to show. Randi Weingarten, president of the 1.4-million-member American Federation of Teachers, once fought efforts to use the value-added approach in New York, and in August cautioned the Obama administration against "over-reliance on an unproven idea."
But this month, her union announced innovation awards for several local unions that plan to adopt value-added measures as part of a new teacher evaluation system even though Weingarten said she still has reservations.
"What we've realized is, if we don't do it by ourselves, no one is going to listen to us," she said.
An early skeptic
Grier was skeptical when he first encountered value-added results in the early 1990s as superintendent of schools in Williamson County, Tenn., just south of Nashville.
Though it's one of the wealthiest districts in the country, many students in its highest-achieving schools were not making much annual progress. Students in some of the poorer schools, meanwhile, were making remarkable gains, learning two years' worth of material in a single grade.
Then, at an academic conference, Grier was won over after seeing a presentation by William Sanders, then a University of Tennessee statistician.
Sanders described how he had developed a statistical tool that, in essence, took into account three years of a student's test scores to estimate the child's academic trajectory. Then he compared it with the student's actual performance in the fourth year. The difference between the expected growth and actual growth became known as "the teacher effect."
The approach overcame the Achilles' heel of traditional achievement tests, which critics say reflect socioeconomic status more than learning. They judge this year's students against last year's, ignoring potential differences between the two groups.
With the value-added method, students are compared to themselves from year to year, so the results are not skewed by income levels, parental involvement, race or gender.
"A teacher or principal has no responsibility for what's happened to the kids in the past," Sanders said. "What they do have responsibility for is the student's rate of progress. This is what was different about what we did."
Sanders' research found that there is a huge disparity in the effectiveness of teachers, and that good teachers could make an enormous difference in a student's learning, regardless of the child's background.
Subsequent research largely confirmed his findings, and in the process many orthodoxies of education policy were called into question. In particular, studies indicated that the traditional measures of a teacher's quality -- such as years of experience, credentials and education -- have little bearing on his or her effectiveness.
Grier was convinced. "I thought it could be a game-changer," he said.
To challenge students at his wealthy schools, Grier pushed principals to focus on growth rather than hitting fixed scores. He used the data to assign teachers to students who matched their strengths. He denied tenure to at least four teachers whose students did not show growth for several years, and reassigned some low-performing principals while encouraging others to retire.
At his next job in Guilford County, N.C., Grier hired Sanders to improve on the district's existing value-added program.
The goal of the program, dubbed Mission Possible, was to attract effective teachers to low-performing schools. To do this, Grier offered highly effective teachers up to $10,000 to relocate. They got up to $4,000 more if students' test scores rose.
Many educators protested the plan, but, with no collective bargaining in North Carolina, the board approved Mission Possible in 2006.
By 2007, teacher turnover at some of the area's lowest-performing schools had been cut by almost half, according to a recent independent study. But the program hasn't significantly changed student achievement -- at least, not yet.
"It's a mixed bag," said board member Alan Duncan. The real test, he said, will be a few years from now, when the board expects the changes to show up in test scores.
North Carolina has since adopted value-added statewide.
Teacher resistance
Grier said he knew that Mission Possible wasn't possible in San Diego. "The board made it clear they were not interested in a merit-pay program," he said.
Instead, Grier wanted to find out how the district's high-performing teachers were distributed. He suspected that many were clustered north of Interstate 8, in the wealthier part of the city.
He lured Sanders, now with a for-profit data analysis company, to visit San Diego with the promise of golf at Torrey Pines. The statistician's presentation persuaded the school board to approve a one-year, $80,000 contract with Sanders' company -- but limited its use to identifying students in need of extra help.
Union officials decried the move, saying the money would have been better spent hiring another teacher. They suspected Grier was trying to move the district toward merit pay. "We knew his history," said union President Camille Zombro.
She and other union members interrupted a school board meeting in June to present the board with a petition complaining about Grier's top-down management style and calling for his removal.
"He looked good in a suit, but he wasn't willing to collaborate," Zombro said.
Relations with the administrators union also soured when Grier tried to use value-added scores as a component of principal evaluation, a plan that was eventually rejected.
The board did not renew Sanders' contract this fall and never made public the results of his analysis.
Grier took over as Houston's superintendent in September. Reached by telephone recently, he seemed invigorated.
"They do things differently out here," he said. "It's a breath of fresh air."
Wednesday, May 14, 2008
Educators want TAKS to count progress
By JENNIFER RADCLIFFE | Houston Chronicle
May 13, 2008
Texas students should be measured on gains they make throughout the school year, rather than facing punitive measures if they fail to clear the hurdles set by the state's standardized test, educators and community leaders told legislators Monday.
Making progress on the test, they argue, is a more important indicator that students and teachers are trying their hardest. It would also take pressure off students, who can currently be retained or kept from graduating if they don't pass certain parts of the exam, educators told members of the Select Committee on Public School Accountability during a public hearing in Aldine.
About 100 people attended the hearing, one in a series being held by the 15-member panel on how the state's Texas Assessment of Knowledge and Skills testing system should be changed.
"We are mindfully listening and internalizing all the testimony," Brownsville ISD deputy superintendent Beto Gonzales said.
Legislators already opted to replace high school-level TAKS tests with end-of-course exams, a move they hope will provide a more accurate read of what students are learning.
Houston mother Thelma de la Cruz, whose son is a fourth-grader at Harvard Elementary, said TAKS testing has left her family "exhausted, pressured and embattled."
"Students bear the weight of testing on their shoulders all year," said de la Cruz, who said her son is in therapy because of TAKS stress.
Jefferson Davis High School senior Jesus Santoya said state testing puts him in a constant state of worry.
"It is scary to know that if I don't meet some certain standards, it makes me, my teachers, my school and my family look bad," the 18-year-old said.
Aldine Superintendent Wanda Bamberg told the panel that they need to reduce the number of tests given, measure districts against those with similar demographics and better align the state's accountability system with the federal No Child Left Behind law.
She also supports the idea of moving to a statewide "value-added" system — or looking at the growth students make over the school year.
"That's the layer of the onion we need to peel," she said.
May 13, 2008
Texas students should be measured on gains they make throughout the school year, rather than facing punitive measures if they fail to clear the hurdles set by the state's standardized test, educators and community leaders told legislators Monday.
Making progress on the test, they argue, is a more important indicator that students and teachers are trying their hardest. It would also take pressure off students, who can currently be retained or kept from graduating if they don't pass certain parts of the exam, educators told members of the Select Committee on Public School Accountability during a public hearing in Aldine.
About 100 people attended the hearing, one in a series being held by the 15-member panel on how the state's Texas Assessment of Knowledge and Skills testing system should be changed.
"We are mindfully listening and internalizing all the testimony," Brownsville ISD deputy superintendent Beto Gonzales said.
Legislators already opted to replace high school-level TAKS tests with end-of-course exams, a move they hope will provide a more accurate read of what students are learning.
Houston mother Thelma de la Cruz, whose son is a fourth-grader at Harvard Elementary, said TAKS testing has left her family "exhausted, pressured and embattled."
"Students bear the weight of testing on their shoulders all year," said de la Cruz, who said her son is in therapy because of TAKS stress.
Jefferson Davis High School senior Jesus Santoya said state testing puts him in a constant state of worry.
"It is scary to know that if I don't meet some certain standards, it makes me, my teachers, my school and my family look bad," the 18-year-old said.
Aldine Superintendent Wanda Bamberg told the panel that they need to reduce the number of tests given, measure districts against those with similar demographics and better align the state's accountability system with the federal No Child Left Behind law.
She also supports the idea of moving to a statewide "value-added" system — or looking at the growth students make over the school year.
"That's the layer of the onion we need to peel," she said.
Tuesday, December 18, 2007
Gates funding gives boost to HISD program
Sounds like 4.5 million new reasons to promote teaching to the tests. -Patricia
$4.5 million to help train teachers new way to analyze student test scores
By ERICKA MELLON | Houston Chronicle
December 7, 2007
The Houston school district's push to grade teachers on their students' progress got a $4.5 million boost Thursday from the Bill & Melinda Gates Foundation.
The grant, the second multimillion-dollar award the district has received for this effort in recent months, will help fuel what school and foundation leaders call a major reform plan to improve teaching and ensure that all students are prepared for college.
"This project is about helping teachers help kids perform at their highest rates," said Steven Seleznow, an education director for the Seattle-based foundation started by the Microsoft Corp. chairman and his wife.
The district plans to use the grant mostly to train teachers on a different way of analyzing test scores. The "value-added" method evaluates how much progress individual students are making on standardized tests year after year — or whether teachers are adding value to students' education.
Value-added models
Under the performance-pay plan, unveiled amid controversy last year, teachers, principals and even Superintendent Abelardo Saavedra are eligible for bonuses based on this student growth.
This year, the district has revised its program and contracted with William Sanders, the pioneer of value-added models, to crunch the numbers. The school board approved paying his company no more than $473,000 this year.
Some teachers, however, say the bonus program still is unfair and too complicated. Various education researchers also question whether value-added models truly identify the best teachers and whether bonuses should be linked to them.
"I'm a skeptic because these are not perfect measures of teacher performance," said Dale Ballou, an associate professor of public policy and education at Vanderbilt University in Nashville. "It doesn't mean it's a bad idea to try this. The question is whether we get something out of this that is good enough that it motivates teachers."
Sanders developed his model in Tennessee, where it has been part of the school accountability system since the early 1990s.
The Houston district also has hired Battelle for Kids, an Ohio-based nonprofit, to train employees about the data and create user-friendly online charts showing which schools are making progress and which aren't moving students along fast enough.
Parents and the public will have access to some of the color-coded charts, which could be a helpful, if confusing, tool to evaluate a school's performance. The district plans to use some of its grant money to produce documents to help parents understand the system, which confounds even veteran educators.
Traditionally, parents could find out only the percentage of students at a given school who passed the Texas Assessment of Knowledge and Skills.
"Parents don't simply want their child to pass, and yet the existing systems — No Child Left Behind, the state accountability system — are all about passing," said school board vice president Harvin Moore. "When schools are judged on what percentage of the children passed, people focus on the 'bubble' kids, kids that are just under passing and the ones that are just over passing, to make sure they don't slip back. It's not that anybody intends to do this."
The board approved paying Battelle a maximum of $10.4 million over three years, with grants funding most of the tab.
The district has named its latest reform effort ASPIRE (Accelerating Student Progress, Increasing Results and Expectations).
Lisa Auerbach, who teaches at Herod Elementary, supports the new focus on student growth. She typically tries to calculate her students' progress on the national Stanford test.
"I have kids who may not yet be on grade level, but, by gosh, I've taken them further than a year's growth," she said. "I call that success because, if the next teacher can do the same thing, then over time that gap is going to narrow."
Complexity 'biggest flaw'?
Steve Antley, a social studies teacher at Marshall Middle School, said that even after attending training on the value-added system, he finds it too complicated to be helpful.
"It reminded me of being in a statistics class during graduate school," said Antley, who is president of the Congress of Houston Teachers. "I think the complex nature of the plan is its biggest flaw in terms of using it for performance pay. To create a performance-pay model, teachers should be able to clearly understand how the money's being awarded."
Sanders makes no apologies for his complex formula, though he wouldn't say whether he supports HISD's decision to award bonuses based on it.
"I'm not going to trade simplicity of calculations for the reliability of the information," he said. "Before groups of teachers, I often hold up a cell phone and I say, 'I don't have a clue what's inside this, but I have to have trust that when I punch the numbers, it's going to call the right number.' "
$4.5 million to help train teachers new way to analyze student test scores
By ERICKA MELLON | Houston Chronicle
December 7, 2007
The Houston school district's push to grade teachers on their students' progress got a $4.5 million boost Thursday from the Bill & Melinda Gates Foundation.
The grant, the second multimillion-dollar award the district has received for this effort in recent months, will help fuel what school and foundation leaders call a major reform plan to improve teaching and ensure that all students are prepared for college.
"This project is about helping teachers help kids perform at their highest rates," said Steven Seleznow, an education director for the Seattle-based foundation started by the Microsoft Corp. chairman and his wife.
The district plans to use the grant mostly to train teachers on a different way of analyzing test scores. The "value-added" method evaluates how much progress individual students are making on standardized tests year after year — or whether teachers are adding value to students' education.
Value-added models
Under the performance-pay plan, unveiled amid controversy last year, teachers, principals and even Superintendent Abelardo Saavedra are eligible for bonuses based on this student growth.
This year, the district has revised its program and contracted with William Sanders, the pioneer of value-added models, to crunch the numbers. The school board approved paying his company no more than $473,000 this year.
Some teachers, however, say the bonus program still is unfair and too complicated. Various education researchers also question whether value-added models truly identify the best teachers and whether bonuses should be linked to them.
"I'm a skeptic because these are not perfect measures of teacher performance," said Dale Ballou, an associate professor of public policy and education at Vanderbilt University in Nashville. "It doesn't mean it's a bad idea to try this. The question is whether we get something out of this that is good enough that it motivates teachers."
Sanders developed his model in Tennessee, where it has been part of the school accountability system since the early 1990s.
The Houston district also has hired Battelle for Kids, an Ohio-based nonprofit, to train employees about the data and create user-friendly online charts showing which schools are making progress and which aren't moving students along fast enough.
Parents and the public will have access to some of the color-coded charts, which could be a helpful, if confusing, tool to evaluate a school's performance. The district plans to use some of its grant money to produce documents to help parents understand the system, which confounds even veteran educators.
Traditionally, parents could find out only the percentage of students at a given school who passed the Texas Assessment of Knowledge and Skills.
"Parents don't simply want their child to pass, and yet the existing systems — No Child Left Behind, the state accountability system — are all about passing," said school board vice president Harvin Moore. "When schools are judged on what percentage of the children passed, people focus on the 'bubble' kids, kids that are just under passing and the ones that are just over passing, to make sure they don't slip back. It's not that anybody intends to do this."
The board approved paying Battelle a maximum of $10.4 million over three years, with grants funding most of the tab.
The district has named its latest reform effort ASPIRE (Accelerating Student Progress, Increasing Results and Expectations).
Lisa Auerbach, who teaches at Herod Elementary, supports the new focus on student growth. She typically tries to calculate her students' progress on the national Stanford test.
"I have kids who may not yet be on grade level, but, by gosh, I've taken them further than a year's growth," she said. "I call that success because, if the next teacher can do the same thing, then over time that gap is going to narrow."
Complexity 'biggest flaw'?
Steve Antley, a social studies teacher at Marshall Middle School, said that even after attending training on the value-added system, he finds it too complicated to be helpful.
"It reminded me of being in a statistics class during graduate school," said Antley, who is president of the Congress of Houston Teachers. "I think the complex nature of the plan is its biggest flaw in terms of using it for performance pay. To create a performance-pay model, teachers should be able to clearly understand how the money's being awarded."
Sanders makes no apologies for his complex formula, though he wouldn't say whether he supports HISD's decision to award bonuses based on it.
"I'm not going to trade simplicity of calculations for the reliability of the information," he said. "Before groups of teachers, I often hold up a cell phone and I say, 'I don't have a clue what's inside this, but I have to have trust that when I punch the numbers, it's going to call the right number.' "
Subscribe to:
Posts (Atom)