Translate

Showing posts with label multiple criteria/measures. Show all posts
Showing posts with label multiple criteria/measures. Show all posts

Thursday, December 04, 2025

What We Refuse to Learn About Standardized Testing: Dr. Gerald Bracey and the 89th Legislature, by Angela Valenzuela, Ph.D.

What We Refuse to Learn About Standardized Testing: Dr. Gerald Bracey and the 89th Legislature

by 

Angela Valenzuela, Ph.D.

December 4, 2025

Learn more about Dr.
Bracey here.


When I revisited Gerald W. Bracey’s 2009 commentary in Educational Leadership, I was struck by how uncannily prophetic it feels in this political moment—especially in light of the sweeping changes enacted by the 89th Texas Legislature in 2025. Best known for his famous "Bracey Reports," his critique of the nation’s “test mania” reads not as a relic of a bygone era but as a warning flare we failed to heed, one whose consequences unfolded vividly during last session’s battles over testing, accountability, and privatization (see Valenzuela, 2025, for discussion).

While public school advocates secured incremental gains—slightly less disruptive testing policies, quicker turnaround on results, the elimination of a handful of assessments, and improved transparency—the underlying system remains largely intact. Schools are still judged primarily by test scores, accountability frameworks remain narrowly defined, and the deeper structural reforms that advocates have long demanded are once again deferred (Valenzuela, 2025). In this vacuum, the push for charterization and privatization only accelerates. 

Even so, there was a brief moment of meaningful political possibility: Texas lawmakers introduced proposals to overhaul or even eliminate the STAAR exam altogether. Raise Your Hand Texas, in particular, deserves a shout-out for leading the charge. For the first time in decades, a major statewide assessment regime came under serious legislative challenge, and the House even passed a bill to move the state away from STAAR. Yet despite the symbolic significance of this breakthrough, the effort stalled—revealing just how entrenched the test-based accountability system remains.

Bracey began with an observation so commonsensical that many policymakers today still stumble past it: no standardized test can ever know a child better than the teacher who sees that child every day. Yet over a fifty-year span, he noted, the United States managed to devolve from viewing tests as occasionally useful tools, to treating them as compulsory, and finally elevating them into the dominant—almost exclusive—measure of educational quality.

Even though tests like NAEP, PISA, TIMSS, or STAAR were never designed to evaluate teaching, curriculum, or the complex, relational work of schooling, policymakers have repeatedly misappropriated them for exactly those purposes. The result, Bracey argued, has been a national preoccupation with numerical indicators that flatten the human reality of learning, obscure the deeper conditions shaping students’ educational lives, and induce undue stress on teachers and the teaching profession.

That critique feels especially urgent in Texas today. This year, the Legislature passed a statewide voucher program that redirects public funds into private schools, one that rests on the long-standing assumption that public schools are failing (Edison, 2025). This assumption gets reinforced year after year through simplistic interpretations of test scores. 

These debates echo Bracey’s core argument: when policy decisions hinge on flawed measures, those decisions inevitably warp the system they intend to improve. Vouchers, sold as a remedy for supposedly failing schools, rely on the very test-based narratives Bracey spent a lifetime challenging. The claim that low standardized test scores reflect poor teaching ignores the deep structural factors shaping learning in Texas—poverty, segregation, underfunding, and what we are by now discovering as the soaring costs (or "price") of privatization. 

Yet these structural realities rarely appear in the public conversation. Instead, test scores are brandished as evidence that public schools are beyond repair, clearing the political path for vouchers, school district takeovers, and the redirection of taxpayer dollars into private hands.

The fight over STAAR reveals a similar contradiction. Legislators across the spectrum have acknowledged that STAAR tells parents little about what their children actually know and does nothing to inform day-to-day instruction. These critiques echo Bracey almost word-for-word. Still, unless Texas reimagines assessment from the ground up—beginning with teaching rather than measurement—any replacement risks replicating the same distortions. Bracey insisted that the best assessments are those built by teachers and anchored in teacher-made curricula, not imposed from above. 

Bracey also challenged the persistent belief that national or international test scores determine economic success. He reminded readers that Japan continued to dominate global assessments even as its economy faltered, and that countries like Iceland maintained high scores while facing economic collapse. 

The notion that bubbling in answers on a fourth- or eighth-grade exam shapes global competitiveness is, to use Bracey’s word, “easily refuted.” Yet Texas, like much of the nation, continues to tie student outcomes to broader narratives about economic health and workforce readiness, despite overwhelming evidence that economic forces are far larger than any test score.

Bracey’s insistence on returning to the human, relational work of teaching feels especially vital. Public schools are not failing; they are absorbing the accumulated burdens of inequality, political interference, and relentless underfunding—burdens that privatization schemes will only intensify. Vouchers will not solve the challenges facing Texas students. Nor will another standardized test, no matter how politically appealing.

Bracey’s enduring message is that education is not a number. It is not a rank, a percentile, or a scaled score. It is the web of relationships, communities, and possibilities that unfold when schools are supported rather than scapegoated. If Texas is serious about creating a stronger and more equitable education system, then we must move beyond the illusions spun by test scores and invest once again in teachers, teacher-made curriculum, and public schools as public goods.

In this critical moment, Bracey’s voice reminds us that the greatest danger is not that our tests show too little—but that we have come to believe they show too much.

Reference

Bracey, G. W. (2009). Multiple measuresEducational Leadership, 67(3), 32–37.

Edison, J. (2025, May 3). Private school vouchers are now law in Texas. Here’s how they will work. The Texas Tribune. https://www.texastribune.org/2025/05/03/texas-school-vouchers-greg-abbott-signs/

Valenzuela, A. (2025, September 15). Accountability without justice: The continuing agenda to demonize K–12 public schools to set the stage for further privatization. Educational Equity, Politics & Policy in Texashttps://texasedequity.blogspot.com/2025/09/accountability-without-justice.html 



Wednesday, March 11, 2015

State suspends use of test scores to measure school quality


Big news out of California. Seems they may be hitting a "re-set button."  We need all kinds of "re-set buttons." :-) -Angela


 

The State Board of Education unanimously voted to suspend for a year the Academic Performance Index, which is based on standardized test scores and widely used to evaluate a school's performance in boosting academic achievement. Since the state is rolling out new tests this year, board members said they wanted at least two years of results to judge school progress.
Amid a national backlash against the overuse of test scores, board members also voted to shift from a school quality measure based solely on exam results to one that would include other factors. Possible additions include student attendance, dropout rates, suspensions, English proficiency, access to educational materials and performance in college-level classes.
"We have an opportunity to hit the reset button," board member Patricia A. Rucker said at the Sacramento meeting.
The two proposals drew no opposition from more than a dozen speakers from state organizations representing school board members, administrators, teachers and parents. Los Angeles Unified also supported the proposals.
Edgar Zazueta, the district's lobbyist, said students have not had enough time to master the computerized process for the new tests on which state performance measures will be based. In a dry run of 775 schools this month, one-third were unable to connect to the state testing website.
Zazueta also said the district participated in statewide practice runs of the new tests last year but could not diagnose problems with them because the state did not release results.
"Bottom line, we need more time to learn from the new testing platform and to continue to work with our students, parents and teachers as we transition to a new accountability system," he said in an email.
In comments at the board meeting, Brian Rivas of the Education Trust-West, an Oakland-based advocacy organization for educational equity, cautioned that any new system must focus on closing achievement gaps among different groups of students.
Sherry Griffith of the Assn. of California School Administrators stressed that district officials and principals would continue to push hard for student improvement, using "every bit of data" from local and state tests.
"This is not about suspending accountability," she said.
Various educational groups are working with the state to develop a broader new school quality measure. Specific proposals are expected later this year.
Twitter: @TeresaWatanabe

Saturday, December 04, 2010

8th-grade retention rate is same as before law required passage of TAKS

Check out an earlier post to this blog titled "HISD considers easing the path to the next grade" on the same issue of retention and promotion policies.

-Patricia


By TERRENCE STUTZ / The Dallas Morning News
Saturday, November 27, 2010

AUSTIN – An education law that was designed to cause more eighth-graders who can't pass the TAKS test to be held back is actually having little impact on the percentage of students who are flunking.

Although the retention rate for eighth-graders jumped the first year they were required to pass the TAKS to be promoted to high school, the rate has now dropped back to what it was before the tough new standard was implemented to help stamp out social promotion in middle schools, new data from the Texas Education Agency shows.

The percentage of students retained for the 2009-10 school year – 1.5 percent of all eighth-graders – was identical to the figure from two years earlier when there was no state requirement for those students to pass the TAKS. In the first year of the testing mandate, nearly 2 percent – 6,323 pupils – were held back in the fall of 2008.

Nearly 40,000 eighth-graders failed the TAKS in 2009, but fewer than one in 10 was held back. Most of those promoted were beneficiaries of a waiver provision that allows a student to move to the next grade if the teacher, parents and principal agree that's best for the child.

State education officials had warned of a possible spike in the number of students flunking. But they said that millions of dollars in additional funding and required remedial classes targeting students who fail the TAKS eliminated the worst-case scenario for the new requirement, called the Student Success Initiative.

Other educators said the low retention rate also reflects how local schools are reluctant to keep low-performing eighth-graders in middle school for an extra year, partly because officials fear they might become more likely to drop out.

House Public Education Committee Chairman Rob Eissler, R-The Woodlands, said the requirement may have to be studied to see whether it is accomplishing what was intended – curtailing the number of students automatically promoted to the next grade regardless of what they achieve.

"It's important to see how effective this program is and to modify it if it's not," Eissler said. "You don't want to just flow students through the eighth grade and then make it worse when they get to ninth grade."

Last year, Texas spent about $44 million on extra classes, tutoring, summer school and teacher training as part of the Student Success Initiative program.

"The money may be better spent on other programs," Eissler said, noting that lawmakers will be looking at ways to improve middle school education during their next session that begins in January.

Dallas-area rates

Overall, 4 percent of all students in Texas schools – 177,701 pupils – were retained for the 2009-10 school year based on their grades and test scores in the spring of 2009. The highest retention rate in the elementary grades was 5.6 percent in the first grade, and the highest in secondary grades was 12.3 percent in the ninth grade.

Among Dallas-area school districts, overall retention rates ranged from half a percent in Highland Park to 6.2 percent in Duncanville. In the Dallas district, the rate across all grade levels was 5.1 percent, and for eighth-graders 2.5 percent.

Much of the attention in Texas has been on fifth- and eighth-graders because they are the only students required to pass the TAKS to gain promotion under a 1999 law passed by the Legislature to curtail the widespread practice of social promotion.

The law was originally proposed by former Gov. George W. Bush, who persuaded lawmakers to launch the Student Success Initiative program shortly before he became president. Third-graders also came under the original social promotion law, but the Legislature decided to drop them from the testing requirement last year.

For fifth-graders, the retention rate in 2009-10 was 1.7 percent. The vast majority of students who failed the math or reading sections of the TAKS – 87 percent – were promoted using the waiver provision or by giving them an alternative exam.

Dropout concerns

Debbie Ratcliffe of the TEA gave credit to tutoring and other remedial programs that are helping students who fail the Texas Assessment of Knowledge and Skills, but she acknowledged that some students are being promoted because of concerns that retention causes some students to drop out of school.

"When you have a student who is not successful at that age, he or she begins to give serious thought to dropping out, and one of the characteristics of a dropout is that they are over-age for their grade level after being retained," she said.

"A lot of schools will go ahead and promote those children [who fail the TAKS], especially if their grades have been fair to good. The idea is to keep them moving forward in the classes where they do well, while the schools try to improve their skills in classes where they have having problems," she said. "One of the strengths of the law is that it causes schools to focus on these struggling students and find ways to help them succeed."

But many educators still question the concept of requiring students to pass a standardized test to be promoted. And minority rights groups criticize the requirement because it has a disproportionate impact on minority students.

Retention rates for black and Hispanic students in Texas in 2009 were more than twice that for white students, as more than three-fourths of all students retained were black or Hispanic, according to TEA figures.

Proponents of the testing requirement say it is unsound to promote students who lack the necessary skills for the next grade level. Groups such as the National Center for Fair and Open Testing argue that retention – based on a test – hurts rather than helps the student.

"Students that have been retained once have a 40 percent higher chance of dropping out and a 60 percent higher chance if retained twice," the center said in a 2007 report. "This happens largely because being over-age in-grade damages students' self-confidence and leads them to disengage from school."

Check out the Dallas-Area Retention Rate Figures

HISD considers easing the path to the next grade

What research shows is that the best thing you can do for a student is to promote them with resources. What we too often see is students retained with no resources, other than perhaps test-preparation.

Extending the multiple criteria assessment that was passed last session in House Bill 3 is one of the best things that HISD can do for its students. What's concerning though is that the policy to evaluate teachers based on student test scores is coming into opposition to this because (as the article shows) teachers are explicitly expressing their perception that promoted students are a liability and potential threat to their jobs.

It's clear that multiple criteria assessment for students is also necessary for teachers. This shouldn't be a problem since we now do it for schools under the new accountability system.

-Patricia


Draft proposal suggests reversing policy of promotion through tests, but Grier questions lowering standard

By ERICKA MELLON | HOUSTON CHRONICLE
Nov. 19, 2010

Houston school district officials are debating whether to make it easier for students to be promoted to the next grade level, reversing a decade-old policy touted as one of the toughest in Texas.

A draft proposal from Superintendent Terry Grier's administration, delivered to the school board Thursday, suggests that students should no longer be automatically retained if they fail the state TAKS test or the national Stanford exam.

The change would bring HISD in line with most other districts, which don't require the testing for promotion, but some board members — and Grier himself — questioned whether lowering the standards was the right move.

HISD implemented the stricter rules in the late 1990s under then-Superintendent Rod Paige in an effort to end social promotion - which is the practice of advancing children based on their age, not their academic ability.

"I'm not a fan of social promotion," Grier said. "I worry about the rigor that's in our district right now, and I certainly don't want it to be easier to promote students who don't have the skills they need to be successful at the next grade level."

But, Grier added, research shows that students who are at least two years older than their classmates - generally because they were retained - are more likely to drop out.

Grier said he plans to make a formal recommendation on new promotion standards in coming months, and it could differ from the proposal made by Chuck Morris, his chief academic officer, and Carla Stevens, the assistant superintendent of research.

Whatever HISD decides, fifth- and eighth-graders still would have to pass the Texas Assessment of Knowledge and Skills under state law.
Scores good, but . . .

Stevens said the tougher promotion standards have not pushed student performance above other districts that follow the state guidelines.

"If all of that had worked," she said, "then our test scores should have skyrocketed and our graduation rates should have skyrocketed. Our TAKS scores are good. Our SAT and ACT scores are good, but they're not really where we want them to be."

Former HISD board member Don McAdams, who supported the stricter promotion standards at the time, expressed concern.

"The whole point about promotion standards is to make sure that every kid meets the standard," said McAdams, who now works as a consultant to school boards. "If the promotion standard is what the teachers say, there is a tendency statistically for students to be promoted who aren't ready for the next grade level."

Paige, who left his job running HISD to become the U.S. secretary of education, said he didn't want to second-guess the current administration and board.

"In general, high standards are the way to go, but they have information that I don't have access to," he said.

Stevens suggested that holding students back a grade because they failed the TAKS or the Stanford in one subject doesn't always make sense.

"It may be one piece of something they're not getting, maybe a component of math," she said. "That doesn't mean they should repeat the entire grade and repeat science and social studies and everything. We want them to be receiving on-grade instruction and then receive interventions and assistance in areas where they are deficient."

Last year, nearly 26,000 students in HISD - about 13 percent of the district - didn't meet the promotion standards. But the majority didn't end up being retained because they passed their courses in summer school or a campus committee gave them a waiver.
Above-average retention

Still, HISD retains a higher percentage of students than the state average, according to data from the Texas Education Agency.

Gayle Fallon, the president of the Houston Federation of Teachers, said it wouldn't be fair to hold teachers accountable for the students' TAKS and Stanford scores if they are no longer part of the promotion requirements. The district gives bonuses to teachers based on student test scores and is working to incorporate them into teachers' formal job evaluations.

"You can't measure teachers on a standard that every child in the room knows doesn't count," she said.

Monday, December 28, 2009

Quality exams strike right balance, ensure quality

This is one of a series of articles in this month's Educational Leadership. Here's a link to preview the others here

-Patricia

Assessment quality and assessment balance—only these can ensure that multiple measures give stable estimates of student achievement.

Stephen Chappuis, Jan Chappuis and Rick Stiggins | Educational Leadership
November 2009 | Volume 67 | Number 3
Multiple Measures Pages 14-19

Long before No Child Left Behind (NCLB), high-stakes tests were common in schools. Cut scores on tests have dictated promotion from one grade level to the next, and teachers have used them to assign passing or failing grades. High school students continue to take course placement exams, subject-area finals, exit exams, and college entrance tests. Making decisions that affect individuals and groups of students on the basis of a single measure is part of our past and current practice.

In the past, few educators, policymakers, or parents would have considered questioning the accuracy of these tests. Most assumed that a low score or grade was probably justly assigned and that a decision made about a student as a result was as defensible as the evidence on which it was based.

But NCLB has exposed students to an unprecedented overflow of testing. In response to the accountability movement, schools have added new levels of testing that include benchmark, interim, and common assessments. Using data from these assessments, schools now make decisions about individual students, groups of students, instructional programs, resource allocation, and more. We're betting that the instructional hours sacrificed to testing will return dividends in the form of better instructional decisions and improved high-stakes test scores.

Given the rise in testing, especially in light of a heightened focus on using multiple measures, it's increasingly important to address two essential components of reliable assessments: quality and balance.
Keys to Quality

Although it may seem as though having more assessments will mean we are more accurately estimating student achievement, the use of multiple measures does not, by itself, translate into high-quality evidence. Using misinformation to triangulate on student needs defeats the purpose of bringing in more results to inform our decisions.

Five keys to assessment quality provide the larger picture into which our multiple measures must fit (Stiggins, Arter, Chappuis, & Chappuis, 2006). Only assessments that satisfy these standards—whether teachers' classroom assessments, department or grade-level common assessments, or benchmark or interim tests—will be capable of informing sound decisions.
Clear Purpose

The assessor must begin with a clear picture of why he or she is conducting the assessment. Who will use the results to inform what decisions? The assessor might use the assessment formatively—as practice or to inform students about their own progress—or summatively—to feed results into the grade book. In the case of summative tests, the reason for assessing is to document individual or group achievement or mastery of standards and measure achievement status at a point in time. The purpose is to inform others—policymakers, program planners, supervisors, teachers, parents, and the students themselves—about the overall level of students' performance.
Clear Learning Targets

The assessor needs to have a clear picture of what achievement he or she intends to measure. If we don't begin with clear statements of the intended learning—clear and understandable to everyone, including students—we won't end up with sound assessments.

For this key to quality, it's important to know the learning targets represented in the written curriculum. The four categories of learning targets are

* Knowledge targets, which are the facts and concepts we want students to know. In math, a knowledge target might be to recognize and describe patterns.
* Reasoning targets, which require students to use their knowledge to reason and problem solve. A reasoning target in math might be to use statistical methods to describe, analyze, and evaluate data.
* Performance skill targets, which ask students to use knowledge to perform or demonstrate a specific skill, such as reading aloud with fluency.
* Product targets, which specify that students will create something, such as a personal health-related fitness plan.

For each assessment, regardless of purpose, the assessor should organize the learning targets represented in the assessment into a written test plan that matches the learning targets represented in the curriculum.

For example, Figure 1 shows a 3rd grade math test plan. It defines what the test will cover, including such specific learning targets as being able to multiply by two (one of the learning targets in the curriculum). Creating a plan like this for each assessment helps assessors sync what they taught with what they're assessing. It also helps them assign the appropriate balance of points in relation to the importance of each target as well as the number of items for each assessed target.

Figure 1. A Sample 3rd Grade Math Test Plan

Sound Assessment Design

This key ensures that the assessor has translated the learning targets into assessments that will yield accurate results. It calls attention to the proper assessment method and to the importance of minimizing any bias that might distort estimates of student learning.

Teachers have choices in the assessment methods they use, including selected-response formats, extended written response, performance assessment, and personal communication. Selecting an assessment method that is incapable of reflecting the intended learning will compromise the accuracy of the results. For example, if the teacher wants to assess knowledge mastery of a certain item, both selected-response and extended written response methods are good matches, whereas performance assessment or personal communication may be less effective and too time-consuming. Figure 2 (page 18) clarifies which assessment methods are most likely to produce accurate results for different learning targets.

Figure 2. Choosing the Right Assessment

Bias can also creep into assessments and erode accurate results. Examples of bias include poorly printed test forms, noise distractions, vague directions, and cultural insensitivity. Teachers can minimize bias in a number of ways. For example, to ensure accuracy in selected-response assessment formats, they should keep wording simple and focused, aim for the lowest possible reading level, avoid providing clues or making the correct answer obvious, and highlight crucial words (for instance, most, least, except, not).
Effective Communication of Results

The assessor must plan to manage information from the assessment appropriately and report it in ways that will meet the needs of the intended users, keeping in mind the following: Are results communicated in time to inform the intended decisions? Will the users of the results understand them and see the connection to learning? Do the results provide clear direction for what to do next?

This key relates directly back to the purpose of the assessment. For instance, if students will be the users of the results because the assessment is formative, then teachers must provide the results in a way that helps students move forward. Specific, descriptive feedback linked to the targets of instruction and arising from the assessment items or rubrics communicates to students in ways that enable them to immediately take action, thereby promoting further learning.

For example, let's say the content standard you're teaching to is "Understands how to plan and conduct scientific investigations," and your assessment rubric states that a strong hypothesis includes a prediction with a cause-and-effect reason. Feedback to students can use the language of the rubric: "What you have written is a hypothesis because it is a prediction about what will happen. You can improve it by explaining why you think that will happen." Or, you can highlight the phrases on the rubric that describe the hypothesis's strengths and areas for improvement and return the rubric with the work.

A grade of D+, on the other hand, may be sufficient to inform a decision about a student's athletic eligibility, but it is not capable of informing the student about the next steps in learning.
Student Involvement in the Assessment Process

Students learn best when they monitor and take responsibility for their own learning. This means that teachers need to write learning targets in terms that students will understand.

For example, suppose we are preparing to teach 7th graders how to make inferences. After defining inference as "a conclusion drawn from the information available," we might put the learning target in student-friendly language: "I can make good inferences. This means I can use information from what I read to draw a reasonable conclusion." If we were working with 2nd graders, the student-friendly language might look like this: "I can make good inferences. This means I can make a guess that is based on clues."

Teachers should design the assessment so students can use the results to self-assess and set goals. A mechanism should be in place for students to track their own progress on learning targets and communicate their status to others. For example, a student might assess how strong his or her thesis statement is by using phrases from a rubric, such as "Focuses on one specific aspect of the subject" or "Makes an assertion that can be argued."
Keys to Balance

The goal of a balanced assessment system is to ensure that all assessment users have access to the data they want when they need it, which in turn directly serves the effective use of multiple measures.

One way to think about the various uses of assessment in a balanced system is by grouping the assessments into levels associated with the frequency of their administration. Ongoing classroom assessments serve both formative and summative purposes and meet students' as well as teachers' information needs. Periodic interim/benchmark assessments can also serve program evaluation purposes, as well as inform instructional improvement and identify struggling students and the areas in which they struggle. Annual state and local district standardized tests serve annual accountability purposes, provide comparable data, and serve functions related to student placement and selection, guidance, progress monitoring, and program evaluation.

Effectively planning for the use of multiple measures means providing assessment balance throughout these three levels, meeting student, teacher, and district information needs. This is done using both formative and summative assessments, large-group and individual testing, assessing a range of relevant learning targets using a range of appropriate assessment methods.

As a "big picture" beginning point in planning for the use of multiple measures, assessors need to consider each assessment level in light of four key questions, along with their formative and summative applications1 :
What decisions will the assessment inform?

At the level of ongoing classroom assessments, formative applications involve what students have mastered and what they still need to learn. At the level of periodic interim/benchmark assessments, they involve which standards students are not mastering and where teachers can improve instruction right away. At the level of annual state/district standardized assessments, they involve where and how teachers can improve instruction—next year.

Summative applications refer to grades students receive (classroom level); whether the program of instruction has delivered as promised and whether the school should continue to use it (periodic assessment level); and how many students have met standards (annual testing level).
Who is the decision maker?

This will vary. The decision makers might be students and teachers at the classroom level; instructional leaders, learning teams, and teachers at the periodic level; or curriculum and instructional leaders and school and community leaders at the annual testing level.
What information do the decision makers need?

From a formative point of view, decision makers at the classroom assessment level need evidence of where students are on the learning continuum toward each standard, whereas decision makers at the next two levels want to know which standards students are struggling to master.

From a summative point of view, users at the classroom and periodic assessment levels want evidence of mastery of particular standards; at the annual testing level, decision makers want the percentage of students meeting each standard.
What are the essential assessment conditions?

These conditions are most articulated at the classroom assessment level, through the use of clear curriculum maps for each standard, accurate assessment results, effective feedback, and results that point student and teacher clearly to next steps. Summative applications at this level include accurate summaries of evidence and grading symbols that carry clear and consistent meaning for all.

At the periodic level of assessment, essential assessment conditions include results that show mastery of program standards aggregated over students. At the annual testing level, accurate evidence of how each student did in mastering each standard aggregated over students is needed.
What Assessments Can—and Cannot—Tell Us

In such an intentionally designed and comprehensive system, a wealth of data emerges. Inherent in its design is the need for all assessors and users of assessment results to be assessment literate—to know what constitutes appropriate and inappropriate uses of assessment results—thereby reducing the risk of applying data to decisions for which they aren't suited.

For example, because they understand what is appropriate at each of the three levels of assessment—both formatively and summatively—assessment-literate teachers would not

* Use a reading score from a state accountability test as a diagnostic instrument for reading group placement.
* Use SAT scores to determine instructional effectiveness.
* Rely solely on performance assessments to test factual knowledge and recall.
* Assess learning targets requiring the "doing" of science with a multiple-choice test.

Assessment literacy is the foundation for a system that can take advantage of a wider use of multiple measures. At the classroom level, teachers can choose among the four assessment methods (selected-response, extended written response, performance assessment, and personal communication). Most assessments developed beyond the classroom rely largely on selected-response or short-answer formats and are not designed to meet the daily, ongoing information needs of teachers and students. As such, not only are they limited in key formative uses, but they also cannot measure more complex learning targets at the heart of instruction.

Because classroom teachers can effectively use all available assessment methods, including the more labor-intensive methods of performance assessment and personal communication, they can provide information about student progress not typically available from student information systems or standardized test results. The classroom is also a practical location to give students multiple opportunities to demonstrate what they know and can do, adding to the accuracy of the information available from that level of assessment.
A Solid Foundation for a Balanced System

Educators are more likely to attend to issues of quality and serve the best interests of students when we build balanced systems, with assessment-literate users. From that foundation we can develop coordinated plans for the use of multiple measures, taking advantage of dependable data generated at every level of assessment.
References

Chappuis, J. (2009). Seven strategies of assessment for learning. Portland, OR: Educational Testing Service.

Stiggins, R., Arter, J., Chappuis, J., & Chappuis, S. (2006). Classroom assessment for student learning—Doing it right, using it well. Portland, OR: Educational Testing Service.
Endnote

1 A detailed chart listing key issues and their formative and summative applications at each of the three assessment levels is available at www.ascd.org/ASCD/pdf/journals/ed_lead/el200911_chappius_table.pdf

Stephen Chappuis (schappuis@ets.org), Jan Chappuis (jchappuis@ets.org), and Rick Stiggins (rstiggins@ets.org) work with the ETS Assessment Training Institute in Portland, Oregon (www.ets.org/ati).

The Tests That Won't Go Away

Marge Scherer | Educational Leadership

November 2009 | Volume 67 | Number 3
Multiple Measures Pages 5-5

How many hours of classroom time do you typically spend administering standardized tests to students each school year? In my search for that statistic, I found one high school teacher estimating he spent 40 school days each year administering and prepping students for "bubble tests."

Perhaps an even more important question is, How many hours does a teacher spend preparing students for "multiple assessments"?

That answer depends on the interpretation of the term assessment—are you counting pop quizzes and spelling bees, essays and multimedia projects, teacher-made and standardized tests, entrance and exit tests, pre-tests and post-tests, interim and benchmark assessments, statewide and national tests, and preparation for the AP exam, SAT, and ACT? Are you adding in daily, minute-by-minute checks for understanding? If all answers apply, many teachers might answer that they spend all their time teaching, if not to the tests, then with the tests in mind.

There is no doubt that in the past 10 years, school culture has become a testing culture. As assessment experts Stephen Chappuis, Jan Chappuis, and Rick Stiggins write (p. 15), "NCLB has exposed students to an unprecedented overflow of testing."

But do all these "multiple measures" really lead us to achieve the three most often cited goals of testing: building proficiency in basic skills, closing achievement gaps, and fostering the top-notch knowledge and skills that students will need in a competitive global society? No, according to the testing experts. Neither using single tests nor incorporating multiple measures necessarily leads in these three directions. Our authors describe what makes it more likely that the instructional hours sacrificed to testing will return dividends in the form of more learning for students and better instructional decisions for teachers.

Here, in brief, is what they tell us:

Become assessment literate. If educators understand the different ways to define multiple measures, the various ways to combine measures, and when to use which (Susan Brookhart's chart on p. 10 certainly helps), they will be one step closer to identifying and choosing appropriate measures for various educational purposes. As assessment-literate professionals, we are in a much better position to educate those in the general public, media, and policymaking positions who blindly accept the validity of any test for any purpose.

Our authors also issue cautionary warnings about specific tests—from the NAEP to TIMSS (p. 32) to value-added assessments (p. 38). It is essential that more people understand the aims of these tests, whom they test and how, and their strengths and limitations in providing useful and valid information about students and schools. Although assessment literacy is no magic bullet—Jim Popham calls it a magic BB—it has a power of its own to transform assessment into a form of teaching.

Keep your eyes on the prize. If the goal is not just achieving higher scores, but furthering students' learning and understanding, assessment must be for learning, not just of learning. Kari Smith (p. 26) speaks from experience about how to make assessment into a teaching tool. Her students not only take tests but also make tests and learn from all the processes involved, from wording the questions to grading the answers. Douglas Fisher and Nancy Frey (p. 20) suggest a systematic approach that incorporates three components: feed-up, feed-back, and feed-forward. Together these three steps—not bubble tests—make up the strongest intervention available to increase student achievement.

Asked to name the biggest obstacle their school is facing this year, 41 percent of responding educators in a recent informal ASCD SmartBrief poll picked "pressure on students and teachers to improve test results." The general public has a different slant, however. In the recent annual PDK/Gallup poll, Americans by a two-to-one majority supported annual testing of students in grades 3 through 8.

"We will not soon be doing away with standardized tests, nor should we really want to," Kari Smith writes. "Tests illuminate an important aspect of students' learning, namely the ability to present factual knowledge within a given time limit." But that knowledge goes only so far. Standardized tests— even under a system of "multiple measures"— do not guarantee better teaching and learning and can have the opposite effect when used wrongly and excessively. It is time to shine a bright light on multiple measures and use them in a more sophisticated way.

Let's make testing serve teaching instead of the other way around.

Wednesday, December 23, 2009

The Big Tests: What Ends Do They Serve?

Nice writeup by the late Gerald Bracey.

-Patricia


Gerald Bracey | Educational Leadership
November 2009 | Volume 67 | Number 3
Multiple Measures Pages 32-37

To measure the quality of our schools, we need more instruction-sensitive measures than NAEP, PISA, or TIMSS.

I was recently interviewed by the editor of my local paper, the Port Townsend Leader, who expressed a pretty low opinion of tests. His wife teaches 3rd grade in a public school, and he can't imagine how anyone would think that a test could reveal more information about a child than a teacher collects as a matter of course. I agree.

In the last 50 years, the United States has descended from viewing tests first as a useful tool, then as a necessity, and finally as the sole instrument needed to evaluate teachers, schools, districts, states, and nations (Bracey, 2009). In a nation where test mania prevails, tests will occupy part of the education landscape until we can dig ourselves out of that 50-year hole. In the meantime, it's interesting to consider what some of the well-known testing programs measure and what their appropriate (and inappropriate) uses might be. Here I look at three testing programs—one domestic and two international.
National Assessment of Educational Progress (NAEP)

When U.S. Commissioner of Education Francis Keppel proposed the NAEP in the 1960s, he ran into a buzz saw of objections from virtually every education organization in the nation. "Local control" was sacrosanct then, and the groups feared that a national test would inevitably lead to a national curriculum. Opposition diminished only after Keppel agreed to house the program in a state policy institution, the Education Commission of the States, and to report results in no smaller unit than "region." (After fears subsided, both of these conditions were abandoned.)

Keppel and the NAEP's chief developer, Ralph W. Tyler, intended the assessment to be solely descriptive. Its purpose was to provide an indicator of the nation's general education health by determining what students knew and didn't know in the same way that a health survey determines what proportion of people have tuberculosis or low body fat.

In 1983, administration of the NAEP was put out for competitive bid and awarded to the Educational Testing Service. In 1988, Congress amended the NAEP law to permit state-by-state comparisons and to create the National Assessment Governing Board (NAGB), whose task was to decide what students of a certain age should know. The NAEP thus became prescriptive as well as descriptive.

In its prescriptive aspect, the NAEP reports the percentage of students reaching various achievement levels—Basic, Proficient, and Advanced. The achievement levels have been roundly criticized by many, including the U.S. Government Accounting Office (1993), the National Academy of Sciences (Pellegrino, Jones, & Mitchell, 1999); and the National Academy of Education (Shepard, 1993). These critiques point out that the methods for constructing the levels are flawed, that the levels demand unreasonably high performance, and that they yield results that are not corroborated by other measures.

In spite of the criticisms, the U.S. Department of Education permitted the flawed levels to be used until something better was developed. Unfortunately, no one has ever worked on developing anything better—perhaps because the apparently low student performance indicated by the small percentage of test-takers reaching Proficient has proven too politically useful to school critics.

For instance, education reformers and politicians have lamented that only about one-third of 8th graders read at the Proficient level. On the surface, this does seem awful. Yet, if students in other nations took the NAEP, only about one-third of them would also score Proficient—even in the nations scoring highest on international reading comparisons (Rothstein, Jacobsen, & Wilder, 2006).

Additional characteristics of the NAEP make it a poor accountability tool. First, because any given student would need hours to complete the whole test, no student ever takes the entire test, nor does any school have all its students participate. Neither districts, nor schools, nor individual students find out how they performed (although NAEP has conducted "trial" assessments in 11 large urban districts to explore the feasibility of reporting NAEP data at the district level). This can be taken as both a strength and a weakness. Students, especially older students, likely don't take the NAEP as seriously as they take the SAT, ACT, or high-stakes state tests, so their scores may underestimate their actual achievement. On the other hand, the fact that the NAEP is not a high-stakes test means that there are almost no test-gaming efforts to artificially increase scores. (This also applies to both of the international tests discussed later.)

Claims that recent gains in NAEP trends indicate the success of No Child Left Behind have been widely disputed—in fact, it appears that NAEP increases slowed after NCLB came into existence (Fuller, Wright, Gesicki, & Kang, 2007). Such claims would not be valid in any case, because the NAEP was not designed to measure the performance of schools. The assessment attempts to cover a broad range of knowledge and skills, but it doesn't rest on any specific curriculum or theory of learning. NAEP has nothing to say about education quality at the district or school level and little to say about the smallest reported unit, the state.
Program for International Student Assessment (PISA)

PISA has tested 15-year-olds in reading, mathematics, and science every three years since 2000. It always measures all three topics, but each administration emphasizes one. The Paris-based Organization for Economic Cooperation and Development (OECD) administers PISA to students in the 30 countries that comprise the OECD and to a similar number of partner nations. The next PISA report, which will emphasize reading, will be published in 2010.

The United States usually scores below average on PISA tests. U.S. politicians and media often uncritically accept the tests as valid and point to U.S. schools as being at fault. Such conclusions are wrong on a number of counts.

In the first place, we should question the good sense of comparing a diverse, 300-million-person nation like the United States with tiny homogeneous city-states like Hong Kong and Singapore. In addition to size, other factors complicate the issues. In Hong Kong, schools concentrate on English, Chinese, and mathematics. Proposals to introduce "liberal studies," which looks like critical thinking to me, have stirred great controversy. In Singapore, schools serve a relatively small proportion of low-income students because many low-paying jobs are done by thousands of Malaysians who enter the country each day and return home in the evening or by "guest workers," mostly from Indonesia and the Philippines, who cannot bring their spouses or families.

Those who cite PISA results to criticize the U.S. education system also ignore a number of characteristics that keep PISA from being useful for comparing the quality of schools in different nations. One problem is the fact that PISA is administered only to 15-year-olds. Because different nations start formal schooling at different ages and have different policies about students repeating a grade, such a limited snapshot can hardly tell us much about a nation's overall success in educating students.

Another problem is the design of the test items. As PISA officials write, "The assessment focuses on young people's ability to use their knowledge and skills to meet real-life challenges, rather than merely on the extent to which they have mastered a specific school curriculum" (OECD, 2005, p. 12). Because the test purportedly measures students' ability to incorporate information that they might not have learned in school, PISA's design would seem to bias it toward affluent students whose homes and families have more resources.

The University of Oslo's Svein Sjøberg (2007) points out that PISA's "requirement that the text should be more or less identical [in different countries] results in rather strange prose in many languages" (p. 14), and that the translations of at least one PISA item word-for-word from English to Norwegian rendered it nonsensical. He is quite skeptical, as am I, that questions can be rendered free of cultural bias and translated into the many languages of PISA countries and still be the "same" questions. And some of the passages for science and math questions are so long and discursive that they obviously measure reading skills as well.

PISA reports contain the nations' average score, rank, and proportion of students reaching various levels of achievement. Virtually all the media and political attention goes to the average scores and ranks. But as Hal Salzman of the Urban Institute and Lindsay Lowell of Georgetown University observe (2008), the students scoring average are not likely to become national leaders in their chosen fields. Future innovators and leaders are more likely to come from high scorers—and the United States produces more than twice as many of these as any other OECD nation does. The bad news is that the United States also produces more low scorers than any other nation except Mexico.
Trends in International Mathematics and Science Study (TIMSS)

TIMSS comes to us from the International Association for the Evaluation of Educational Achievement in The Netherlands, but most of the technical work is conducted at Boston College. It measures selected math and science skills in grades 4 and 8 using short, fact-oriented stems and mostly multiple-choice questions.

We have been through four rounds of TIMSS: 1995, 1999, 2003, and 2007. As with PISA, politicians and the public are quick to use TIMSS results to criticize the quality of U.S. schools. In his March 2009 speech to the Hispanic Chamber of Commerce, President Obama observed, "In 8th grade math, we've fallen to 9th place." Ninth place was indeed the U.S. rank (among 46 nations) for the 2007 TIMSS administration, but in 1995, the United States ranked 28th out of 41 countries. U.S. scores as well as ranks have actually risen for 8th graders, and they have been stable for 4th graders.

The TIMSS developers explicitly make a causal connection between high scores and a country's economic health and claim that "there is almost universal recognition that the effectiveness of a country's educational system is a key element in establishing competitive advantage in what is an increasingly global economy" (Mullis, Martin, & Foy, 2008). Even if this were true, the question would be, Does TIMSS measure that effectiveness? The answer is no. No test can do that, because no test can measure the many complexities of an "educational system," much less a test that measures only two subjects. To get some idea of the complexity of an "educational system," I suggest that readers glance through the 100-plus goals of public education in John Goodlad's 1979 classic, What Schools Are For.
The Education/Economy Fallacy

Both politicians and the media have relentlessly linked scores on national and international assessments to economic health. Release of the PISA results in 2004, for instance, led to headlines like "Economic Time Bomb" (Kronholz, 2004) and "Math + Test = Trouble for the U.S. Economy" (Chaddock, 2004).

This notion is easily refuted by the example of Japan, which led the world in test scores and economic growth in the 1980s but saw its economy sink into the Pacific in the 1990s. Throughout this period, Japanese students continued to ace tests, but Japan's economy sputtered into the new century and slipped back into recession in 2007.

It is doubtful that the ability of 4th and 8th graders to bubble in answer sheets has any connection to the economy. In fact, although educators might not want to recognize it, the current economic calamity should drive home the reality that the economic forces at play in the world dwarf the effects of education. Iceland scores high on international assessments, but in the global crisis of 2008–2009 it became an economic basket case with a national debt equal to 850 percent of its gross domestic product.

Education, by itself, does not produce jobs. There are regions of India, for example, where thousands of applicants show up for a single job requiring moderate education. The people who noticed this phenomenon worry that overeducation in the absence of job production could destabilize India (Jeffrey, Jeffery, & Jeffery, 2008). Similar worries no doubt afflict the government of China, where 33 percent of 2008 college graduates are still looking for jobs (Johnson, 2009).

Those who decry the United States' rankings on international tests should note that the Institute for Management Development (2009) and the World Economic Forum (Porter & Schwab, 2008), two organizations that rate nations on global competitiveness, rank the United States as the most competitive nation in the world—especially in the area of innovation.

In an interview, Singapore Minister of Education Tharman Shanmugaratnam acknowledged that Singapore students score well on tests but often don't fare as well as U.S. students 10 or 20 years down the road. He cited creativity, curiosity, and a sense of adventure as some of the qualities tests don't measure, adding, "These are the areas where Singapore must learn from America" (Zakaria, 2006). Sadly for American students, as Robert Sternberg (2006) observed, "The increasingly massive and far-reaching use of standardized tests is one of the most effective, if unintentional, vehicles this country has created for suppressing creativity" (p. 47).
Blunt Instruments

This brief look at several widely recognized assessments demonstrates that none of these tests are useful for comparing the quality of schools or teachers—especially in the United States, with its diverse population, high poverty rates (by far the highest among developed nations), and wide variety of pedagogical philosophies. As former Commissioner of Education Statistics Mark Schneider said, the tests are "blunt instruments. … A dozen factors could be behind a nation's test score" (Cavanagh & Manzo, 2009, p. 16).

Nations vary greatly in the extent of their efforts to motivate students to do well on the assessments. In Germany, where PISA has likely received more attention than in any other country, PISA-prep books can be found in airports. Observers at a school in Taiwan reported that on PISA testing day, parents gathered with their children on the school grounds urging them to do well. The students then marched into the school to the national anthem and heard a motivational speech from the principal (Sjøberg, 2007).

Can low-scoring and middle-scoring nations learn anything from the high scorers? Mostly, no. After A Nation at Risk appeared in 1983, Secretary of Education Terrel Bell dispatched a team to Japan. The effort came to naught, no doubt in part because Japanese schools and U.S. schools are embedded in vastly different cultures.

W. Norton Grubb and an OECD team observing schools in Finland, which ranks at the top on PISA, found some things the United States could likely adopt—for example, the interlocking system in which teachers and specialists work to head off learning problems early on. But they also noted some things we could not adopt without also adopting other large segments of the Finnish social system, such as comprehensive health care and public housing. Grubb (2007) pointed out, "The Finns take it as axiomatic that both high-quality schooling and nonschool programs are necessary for equity" (p. 109).
A Better Way

To be related to school quality, tests must be sensitive to instruction. Most of the tests used for accountability today aren't—in fact, the manner in which they are constructed prevents them from being sensitive to instruction. That means that schools under the gun to raise test scores increasingly rely on strategies that get immediate, but short-lived results. Evaluation based on instruction-insensitive tests cannot help but reduce the quality of teaching (and teacher morale).

The best assessment system, but a difficult one to bring off, begins with teachers rather than with external measures that are imposed on them. The state of Nebraska developed such a system—the School-based Teacher-led Assessment and Reporting System (STARS)—based on instruction-driven measurement as opposed to the dysfunctional, measurement-driven instruction that predominates elsewhere. (Alas, it appears to have been almost eclipsed by the statewide program installed to meet NCLB requirements.) It is that kind of system—not NAEP, TIMSS, PISA, or similar tests—that will tell us what we need to know about our schools.
References

Bracey. G. W. (2009). Education hell: Rhetoric vs. reality. Alexandria, VA: Educational Research Service.

Cavanagh, S., & Manzo, K. K. (2009, April 22). International exams yield less-than-clear lessons. Education Week, 28(29), 1, 16–17.

Chaddock, G. R. (2004, December 7). Math + test = trouble for U.S. economy. Christian Science Monitor. Available: www.csmonitor.com/2004/1207/p01s04-ussc.html

Fuller, B., Wright, J., Gesicki, K., & Kang, E. (2007). Gauging growth: How to judge No Child Left Behind? Educational Researcher, 36(5), 268–278.

Goodlad, J. (1979). What schools are for. Bloomington, IN: Phi Delta Kappa.

Grubb, W. N. (2007). Dynamic inequality and intervention: Lessons from a small country. Phi Delta Kappan, 89(2), 105–114.

Institute for Management Development. (2009). World competitiveness yearbook. Lausanne, Switzerland: Author.

Jeffrey, C., Jeffery, P., & Jeffery, R. (2008). Degrees without freedom? Masculinities and unemployment in northern India. Palo Alto, CA: Stanford University Press.

Johnson, I. (28 April, 2009). China faces a grad glut after boom at colleges. Wall Street Journal, p. A1.

Kronholz, J. (2004, December 7). Economic time bomb: U.S. teens are among the worst at math. Wall Street Journal, p. B1.

Mullis, I. V. S., Martin, M. O., & Foy, P. (2008). TIMSS 2007 international mathematics report. Chestnut Hill, MA: Boston College.

Obama, B. (2009, March 10). President Obama's remarks to the Hispanic Chamber of Commerce. New York Times. Available: www.nytimes.com/2009/03/10/us/politics/10text-obama.html

OECD. (2005). PISA 2003 data analysis manual. Paris: Author.

Pellegrino, J. W., Jones, L. R., & Mitchell, K. J. (Eds). (1999). Grading the nation's report card: Evaluating NAEP and transforming the assessment of educational progress. Washington, DC: National Academy of Sciences.

Porter, M. E., & Schwab, K. (2008). The global competitiveness report 2008–2009. Geneva, Switzerland: World Economic Forum.

Rothstein, R., Jacobsen, R., & Wilder, T. (2006, November 29). Proficiency for all is an oxymoron. Education Week, 26(13), 32, 44.

Salzman, H., & Lowell, L. (2008). Making the grade. Nature, 453, 28–30.

Shepard, L. (1993). Setting performance standards for student achievement. Stanford, CA: National Academy of Education, Stanford University.

Sjøberg, S. (2007). PISA and "real life challenges": Mission impossible? Available: http://folk.uio.no/sveinsj/Sjoberg-PISA-book-2007.pdf

Sternberg, R. J. (2006, February 22). Creativity is a habit (Commentary). Education Week, p. 47.

U.S. Government Accounting Office. (1993). Educational achievement standards: NAGB's approach yields misleading interpretations (GAO/PEMD-93-12). Washington, DC: Author.

Zakaria, F. (2006, January 9). We all have a lot to learn. Newsweek. Available: www.fareedzakaria.com/ARTICLES/newsweek/010906.html

Saturday, June 13, 2009

Alternative Testing on the Rise

Va. Expands Use of 'Portfolio' to Measure Learning of Challenged Students

By Michael Alison Chandler
Washington Post Staff Writer
Monday, June 8, 2009

For eight days this spring, the Dulles Expo Center was transformed into an industrial grading complex. Boxes of tests were trucked in from schools, unloaded onto pallets into a warehouse and distributed to folding tables, where more than 1,500 Fairfax County teachers and staff worked with sharpened pencils and bar code scanners. These were not multiple-choice tests that computers grade in seconds. They were thick "portfolio" tests representing a year's worth of student worksheets, quizzes and activities. The time-intensive evaluations have proliferated in recent years in response to the testing requirements of the federal No Child Left Behind law.

The District and many states, including Maryland and Virginia, use portfolios for students with serious cognitive disabilities. But Virginia has gone much further, expanding their use for students with learning disabilities or beginning English skills. Statewide, the number of math and reading portfolios submitted for such students nearly doubled in a year, from 15,400 in 2006-07 to more than 30,000 in 2007-08, and state officials predict another jump this school year.

Portfolios have long been used for in-depth evaluations because they can gauge more skills and higher-order thinking. Many educators say the year-long portfolios are a fairer way to measure what some students know than a one-day snapshot.

"We all learn differently," said Patrick K. Murphy, assistant superintendent for accountability in Fairfax schools and Arlington County's incoming superintendent. "We also have to recognize there are different ways people can show proficiency beyond a multiple-choice test."

ad_icon

Pass rates for portfolio tests are relatively high, which helps educators meet academic benchmarks but raises questions about the tests' value in rating schools. Portfolios also are expensive, costing Fairfax more than $500,000 for training and scoring this year alone, not to mention thousands of teacher hours spent compiling them.

The Virginia Grade Level Alternative, one of the state's portfolio tests, is available to some special education students from third to eighth grade who are learning grade-level material but struggle with multiple-choice tests. Someone with severe test anxiety or an information processing disability might be eligible, officials said. A student who might not correctly choose "Answer B: The third U.S. president" for a question about Thomas Jefferson but who could describe Monticello and a president who promoted ideals of freedom yet owned slaves could also be eligible.

The federal government approved Virginia's reading portfolio for beginning English learners in 2007 after protests by local school boards that the regular grade-level test was unfair. The 2002 federal law requires public schools to test students in reading and math in grades three through eight and once in high school. Schools that fail to reach target pass rates for all groups of students, including those with disabilities and English learners, face possible sanctions.

In Northern Virginia, portfolio testing has expanded significantly over two years. About 8,600 math and reading portfolios were compiled in Fairfax this school year, up from 5,900 in 2007-08 and 600 in 2006-07. Similar trends are playing out in Arlington and in Prince William and Loudoun counties.

Pass rates have increased in part because school systems have grown more comfortable compiling portfolios. Last school year, 94 percent of Fairfax students evaluated through portfolios passed in reading, and 84 percent passed in math, up from 79 percent and 70 percent, respectively, in 2006-07. Statewide, 87 percent of such students passed in reading and math last school year, up from 81 percent and 84 percent the year before.

Assembling the portfolios is a feat. Throughout the year, teachers compile worksheets, quizzes, audio or video clips and other examples of what students have learned. Third-grade math teachers documented that students understood 93 concepts, including some at first- or second-grade levels. One six-inch binder was filled with activity sheets that showed how a student used Goldfish crackers to count by twos and compared piles of cubes to demonstrate the concepts "more than" and "less than."

It took a full day for two scorers at the Expo Center to evaluate two or three third-grade math binders, with both independently scoring each piece of evidence. Scores were checked by "bubblers," who transferred the results onto Scantron sheets, which were then double-checked by "double-bubblers" before the binders were loaded back into boxes and shipped out again.

Many parents and teachers say portfolios have improved instruction and ensure that special education students are exposed to an entire year's curriculum, not a shortened version.

That is reassuring to Tia Marsili of Vienna, whose eighth-grade daughter has Down syndrome. "If you cannot show me what you have been doing," she said, "I'm afraid she has not learned the content."

Leonard Bumbaca, president of the Fairfax Education Association, said portfolios are also better in evaluating a teacher's effectiveness than a standardized test. But he warned that the process is "overloading teachers." Some teachers are also feeling pressure to use portfolios more often to achieve high pass rates, Bumbaca said. "Teacher time is a scarce resource," he said. "We don't think it's the best decision . . . to apply this test broadly."

State and local officials say they are monitoring portfolio testing to ensure that it is not overused or misused.

Andrea Rosenthal of Oak Hill, the mother of a Fairfax special education student, said high pass rates on portfolio tests are often misleading because many children who score well on them are far below grade level on other measures. "It benefits the state, not the child, to say they are at grade level when they are not," Rosenthal said.

Thursday, January 17, 2008

Identifying Successful Schools for Low Income and Minority Group Students

Interesting study and policy recommendations though each of the schools examined in the study are very small, between 300 - 500 students. Surely not the case for the greater majority of high schools in California's USDs. To read the policy brief and full report click here . -Patricia

The National Center for Fair and Open Testing
Issue: 01/2008

Most research that looks at successful schools or tracks improvement in education uses standardized test scores as the sole criterion for measuring progress. Such research may produce strong evidence about the practices and policies that raise test scores, but it provides little useful information about what goes into high quality education. Using test results to identify successful schools and then determine "effective practices" for raising those scores ignores the important but untested learning and characteristics that students, parents and the public seek from schools.


A recently released study on equitable high schools intentionally took a very different approach. High Schools for Equity: Policy Supports for Student Learning in Communities of Color examined five nonselective California public high schools. While the researchers explicitly recognized the limits and biases of standardized exams, they first used test scores to establish a pool of schools because scores are among the few indicators systematically collected. They then evaluated a richer array of evidence to identify five high-quality institutions. Justice Matters and the School Redesign Network at Stanford University sponsored the study; the research team was led by Diane Friedlaender and Linda Darling-Hammond.


The study defined school success as providing an education that is academically rigorous while being relevant, responsive and connected to students' cultures. Each profiled school constructs successful learning experiences for low-income students of color. The environment is characterized by caring, respectful relationships with students and families, and the school offers a range of supports tailored to bolster learning. Success also included the schools' ability to retain students through to graduation.


According to the report, "We sought evidence that students in the schools learn to demonstrate their knowledge in rigorous and authentic ways that ensure they are able to investigate and evaluate ideas, communicate and defend their thoughts orally and in writing, and develop intellectual and practical products that meet high standards of evidence and performance."


The schools make extensive use of performance assessments, portfolios and exhibitions. Work on the performance tasks is more engaging for students and provides opportunity for regular feedback by teachers. These assessments are a fundamental part of the instructional process while providing more comprehensive evidence of student achievement. Teachers use them in their own extensive professional development.


High Schools for Equity also studied how district and state policies hinder or enable the schools' efforts to carry out successful practices. It found that high-stakes standardized tests are a serious impediment to the schools' ability to engage in high quality instruction.


One policy recommendation is to redesign assessment systems at the state and local levels to better represent applications of knowledge and skills through performance assessment. The report's recommendations are based in the California context, but many are nationally relevant.

Monday, December 17, 2007

Calls Grow for a Broader Yardstick for Schools

This is an important, but actually very long, unending battle. -Angela

Calls Grow for a Broader Yardstick for Schools
By Maria Glod
Washington Post Staff Writer
Sunday, December 16, 2007; A11

For nearly six years, the federal government has defined school success mainly by how many students pass state reading and math tests. But a growing number of educators and lawmakers are pushing to give more weight to graduation rates, achievement in science and history and even physical education.

The debate over the formula for rating the nation's public schools has stalled efforts in Congress to revise the No Child Left Behind law. At issue: What's the best way to measure whether schools are doing their job?

Unlike questions on the state math and reading tests taken by millions of children, this one has no clear answer. Reaching consensus in the coming election year is expected to be difficult. Without congressional action, the 2002 law will stay as it is.

"Lots of stakeholders have different answers to this question," said Michael Casserly, executive director of the Council of the Great City Schools, a D.C.-based coalition of urban school systems. "The tug of war is over, if not state assessments, then what? You're ultimately going to get as many answers to the question as there are people to answer it."

The American Society of Civil Engineers wants science tests added to the mix. The NAACP and other groups say schools should get credit for achievement in subjects other than reading and math, as well as for improvement in graduation and college admission rates. Some want to give schools points for progress on locally developed tests and for increasing the number of students who excel in Advanced Placement classes.

Reps. Zach Wamp (R-Tenn.), Ron Kind (D-Wis.) and Jay Inslee (D-Wash.) say the law should push children to exercise more than their brains. They introduced a bill to give schools points if students spend more time on physical education.

Advocates for "multiple measures" say that learning is too complex to be judged by annual tests and argue that spontaneity and creativity in classrooms are being lost to test preparation and drills.

"There ought to be more in determining students' success than just one test score," said Reg Weaver, president of the National Education Association, the largest teachers union. "Preparing a child for the 21st century means reading and math. But it also means science; it means civics; it means art."

But the Bush administration and some civil rights, education and business groups say that too many tweaks would weaken a law credited with revealing pockets of struggling students, especially among poor children, minorities and those with disabilities. In their view, an overly complex rating system would mask problems in schools with many students who haven't mastered basic reading and math, skills they call the building blocks to success.

"Proponents of multiple measures say it will give a richer, fuller view of a school, but this isn't about a rich view of a school. It's about failures in fundamental gate-keeping subject areas," said Amy Wilkins, a vice president of the Education Trust, a D.C.-based advocate of better schools for the disadvantaged. "Parents know, 'My school is in trouble because it's not teaching reading and math.' "

At Charles H. Flowers High School in Prince George's County, where test scores in reading and math have usually met state benchmarks, the principal, Helena Nobles-Jones, said: "It's okay to add other factors, but they can't replace reading and math." The two subjects, she said, "are so very critical to any career a child would choose."

The law requires annual reading and math tests in third through eighth grades and once in high school. Schools and subsets of students -- including ethnic minorities and students from poor families -- must make gains over time. High schools also must reach target graduation rates, but the state goals have been criticized as weak and inconsistent.

Certain schools that don't meet standards are required to allow students to transfer or face other sanctions. The law aims to have 100 percent of children proficient in reading and math by 2014. But the ratings are more about identifying struggling schools than rewarding excellence.

George Miller (D-Calif.), chairman of the House education committee, has been trying to craft a definition of school success that goes beyond standardized tests. In a draft bill he circulated last summer, math and reading scores would remain the biggest factors in rating schools. But schools also could gain points for raising science, history, civics or writing scores or increasing the number of students who succeed in college preparatory courses. The proposal would establish a national system to measure graduation rates, and high schools could be rewarded for progress.

Miller said such changes would encourage schools to lower dropout rates, broaden the curriculum and encourage more disadvantaged students to enroll in challenging classes. He said he aims to provide a "better, fairer picture of what's happening in schools."

But some GOP leaders say the picture would only become murkier. Under the law, parents can see how their school stacks up against others across town or across the state on the same exams. If some groups of students struggle, it shows.

Rep. Howard P. "Buck" McKeon (Calif.), the top Republican on the education committee, said the law lets parents "cut through the clutter and see clearly how their children's schools are performing." Add too many measures, he said, and accountability would be lost.

Many advocates for children with disabilities agree. Ricki Sabia, associate director of the National Down Syndrome Society's Policy Center, said the law has forced schools to focus more on children with special needs. "What we've seen in the past five years is kids with disabilities are doing better than anyone expected," Sabia said. "We are very wary of seeing things roll back."

Likewise, the Mexican American Legal Defense and Educational Fund applauds the law for drawing attention to Latino student achievement. The organization isn't opposed to adding measures, said Peter Zamora, a regional counsel for the group, as long as the system "can't be easily gamed."

But many educators report increasing pressure to tailor lessons to annual state exams, leading students to miss out on other educational opportunities.

"The fear is you have this narrowing of the breadth and depth" of the curriculum, said Elizabeth Burmaster, Wisconsin's state superintendent. Burmaster, president of the Council of Chief State School Officers and a former music and drama teacher, supports using local assessments together with state tests. "It's much more complicated," she said. "But it's more accurate."

Tom Loveless, director of the Brown Center on Education Policy at the Brookings Institution, said the challenge is creating a rating system that includes a range of measures and provides a clear picture of a school's effectiveness.

"Most schools people -- and a lot of people who think about schools -- think school is about a bunch of different things, not just reading and math," Loveless said. "The problem is . . . as you list all those things, suddenly it's not as clear-cut what's a successful school and what's a failing school."


http://www.washingtonpost.com/wp-dyn/content/article/2007/12/15/AR2007121501747.html