May 16, 2008
Today's Video
This is a good video to point your friends to the next time one of those new brain-based education fads comes knocking on your school district's door.
As Willingham makes clear, most of the "programs" out there trying to capitalize on important-sounding "brain-based" terminology are mostly bunk. He also explains why.
If you haven't done so yet, go read every one of Willingham's cognitive science articles from AFT's American Educator. You'll be the better person for it. (or at least you'll stop leaving me silly comments showing you haven't read them.)
May 14, 2008
Decodable vs. Predictable Texts
Someone requested an example, and before I had a chance to provide one, a teacher who goes by the handle PalisadesK provided not only a great example, but further elaborated on my point. Here's what she wrote.
Here are two examples from the Reading A-Z site, which offers both “leveled” books (balanced literacy approach – sight words and “predicting” from pictures and the first letter) and “decodable” stories, based on letter-sound correspondences that have been taught. In decodable stories, children should be able to work out the words even if they have not seen that exact word before, providing they have learned the letter/sound association, and learned the skill of blending sounds into words.
Both these are from the Kindergarten level, roughly partway through K:
Leveled Book:
(Fountas and Pinnell Guided Reading Level A)
Maria Counts Pumpkins
(each sentence is on a page by itself, with a detailed illustration)
Maria has one pumpkin. Maria has two pumpkins…. (etc.) up to Maria has seven pumpkins. Maria has too many pumpkins!
Note that most of these words are not “decodable” by the beginning kindergarten student, who would not have learned diphthongs (ou), or how to sound out two-syllable and three-syllable words like “Maria” and pumpkins” and “many.” The number words would have been taught as sight words, and can also be inferred from the pictures. “has” can be decoded, but might also have been taught as a sight word. Many K students learn letter names, but are not taught to decode words, left-to-right, saying the sounds. They are told to use the first letter sound as a “cue” to guess the word.
Now here’s a decodable story which would be appropriate for a child who had learned most of the single consonant sounds and the short vowel sounds (plus the long vowel digraph ay and final y and i-consonant-e as a long I sound)), and been taught to blend them together. Programs like “Jolly Phonics” teach these to four and five year olds in one school term, with a lot of practice in blending new words and spelling by sounds.
Decodable Book:
My Pug Has Fun
(format is the same, but there are several sentences on each page and a detailed drawing. The drawing would help the child confirm that he decoded correctly )
My pug Bud and I like to play. We like to tug on the rug. We like to get a bug. We like to sit on a rug in the sun. (New page, shows pug digging in a sandbox, holding a coffee mug in his jaws, while child takes a nap on a blanket or mat) My pug likes to tug my mug. He dug a pit. He put my mug in a pit.
There are a couple more pages about things the pug and boy do together. It’s not deathless literature, but it’s cute. The child can read it independently.
Many of the “leveled books” at the early stages (you can see some free samples on readinga-z.com) are quite contrived and boring. It’s hard to imagine any kid staying up with a flashlight to read these things under the covers! Although balanced literacy proponents make a big deal about exposing children to “authentic literature,” it is a sad fact that most of the “literature” the students read is contrived and far from authentic – or interesting.
Decodable stories are no prose masterpieces, either, but they do provide children with a chance to consolidate an important skill. Moreover, children who master decoding early and fluently will soon be reading that “authentic literature” that “whole language” enthusiasts love.
May 13, 2008
Looking Beyond the Reading First Controversy
Shep does a good job describing the scandal (or lack thereof) aspect of Reading First and does some solid reporting on the data coming back from the states that tried to implement the program with fidelity.
Full Disclosure: I helped Shep analyze the state data to make sure it was being reported accurately.
May 12, 2008
The Layman's Guide to Reading First
So this is my attempt at a brief and simple primer on what actually went on in Reading First. To understand what's going on you need to understand the reading wars, basic economics, and a little history. Then all you need to do is to follow the money. I'll keep this as short and sweet as is humanly possible.
The diagram below shows you the kinds of reading programs in use.

There are two kinds of reading programs on the market: whole language/balanced literacy and phonics programs.
Whole language programs and balanced literacy programs are similar. Balanced literacy programs merely add a phonics component. Both programs use predictable texts which favors the guessing of text from context clues rather than the decoding of words via phonics skills.
Phonics based programs use highly decodable texts that facilitate the decoding of text via the use of phonics skills and discourages the use of context clues during the initial stages of reading instruction.
Oddly enough, the presence of a phonics component does not distinguish between between balanced literacy programs and phonics based programs. Both programs claim to have a phonics component. In fact, both will argue that they contain the five essential elements of reading instruction (ECRI), i.e., phonemic awareness, phonics, vocabulary development, reading fluency, and reading comprehension strategies. These are the five components that the National Reading Panel's meta-analysis found were present in all highly-effective reading programs (though the presence of these components does not guarantee effectiveness).
Rather, the most easily ascertainable difference between the two is in the texts that are used in the initial stages of reading. Look at the texts and you'll easily see which pedagogical lineage the program belongs to.
Most of the reading programs in use today have not been validated by scientifically based reading research (SBRR). At the time Reading First was adopted, only three reading programs have program specific SBRR -- Success for All (SfA), Direct Instruction's Reading Mastery (DI), and a prior edition of Open Court (OC) (though not the current edition). All of these programs are phonics based programs using highly decodable texts.
The publishers of these reading programs are capitalists and will dutifully provide whatever product the market calls for. The market has been predominantly calling for balanced literacy programs. This is because balanced literacy ideologues are firmly entrenched in most school districts throughout the country and at all levels of the education hierarchy, including at schools of education and at state level departments of education.
The poor track record of balanced literacy programs with 'at-risk" students and the mounting research base of phonics based programs has not served to dislodge the entrenched balanced literacy programs which remain in widespread use.
But, phonics advocates had the Bush administration's ear starting in 2000. Apparently, someone had the idea to let the federal government do what it does best -- bribe schools with federal dollars to adopt scientifically based reading programs for their "at-risk" students. And, thus, Reading First was born as part of NCLB and the effort to force more accountability and obtain better results from the federal dollars being sent to schools to increase the achievement of "at-risk: children. NCLB was enacted with overwhelming bi-partisan support.
Reading First was designed to cause change. And change is going to cause tension between those who want change (and federal dollars) and those that want to maintain the status quo (while getting federal dollars). Given the state of the reading wars let's see how all this breaks down.
At the state level, the motivation is to continue the status quo, i.e., balanced literacy, while grabbing as much Reading First funding as possible. This is called rent seeking and it is what monopolies do. (It's also why your utility companies and the DMV treat you like crap.) As a result, these folks were going to do whatever they could to get their existing reading programs funded with little or no modification. If you read the IG's reports on Reading First, you can plainly see that this activity was going on behind the scenes and DoE was trying to prevent it.
Reading program publishers were trying to increase their market share caused by the disruption caused by the injection of Reading First funding. To accomplish this goal they published phonics-based programs that were capable of being funded under Reading First, but yet were palatable to those balanced-literacy lovin' educators who would ultimately be selecting those programs at the state level. (Publishers are, by and large, capitalists who will provide the market with whatever it is asking for.) Thus, publishers were also trying to game the system by skirting the line between phonics-based programs and balanced-literacy programs in order to grab the biggest market share they could. Similarly, publishers were simultaneously trying to re-market their existing balanced-literacy programs as complying with the Reading First statute.
The U.S. Department of Education would be charged with determining which reading programs qualified for Reading First funding and which ones didn't. The DoE would be constrained by two primary factors: the statute itself and the department's prohibition against directly mandating curricula.
The relevant section of the Reading First statute is vague, requiring only that funding go to a school that selects and implements a reading program that is "based on scientifically based reading research" and includes "the essential components of reading instruction." Both of these terms are further defined in the statute, but not in a meaningful way that would have informed DoE where to draw the funding line.
But, let's cut right to the chase. There are only three meaningful places where the funding line could be drawn based on the statute.
1. Funding limited to reading programs having program-specific reading research. Basically, this would have limited funding to Success for All, Direct instruction, and Open Court. But, the statute clearly intends broader funding than this since it reads "based on" SBRR and not something more restrictive like "having its own program-specific" [SBRR]. Now Bob Slavin thinks the line should have been drawn here, but Slavin is an interested party that would reaped a financial windfall had the line been drawn here. I've been going back and forth with Dean Millot, both in private correspondance and on his blog, about this issue. Dean thinks this is where the line should have been drawn, but has been unable to convince me that his interpretation is proper. I will go so far to say that this is where I wish the line were drawn, but the statutory language precludes such a strict interpretation. Moreover, no one involved in the drafting of the law has gone on record and said that this is the proper interpretation of the statute. If anything, those members of Congress who have made statements appear to be of the opinion that the funding should have been widely distributed.
2. Funding limited to phonics-based programs having decodable texts. In my opinion, this is the most likely place that the funding line should have been drawn based on the statutory language. In fact, this is where DoE thought the line was drawn and was the distinguishing point at which they excluded programs from funding. All the major publishers must have thought that this was where the line was drawn as well sicne they all developed new phonics-based reading programs. Publishers were also trying no make their balanced literacy programs appear to be phonics-based by including a phonics component. these programs were rightfully sent packing by DoE since the research shows that a phonics component won't do the trick; it has to be a phonics component plus complementary decodable texts. In any event, this is where the main battle line was drawn and where much of the behind-the-scenes fighting took place. A few questionable reading programs appear to have been funded, but it appears that those programs snuck by with the help of ideologues at the state level that did whatever they could to obscure the reading programs they intended to fund on their Reading First applications. It also appears that DoE was fearful about violating the prohibition against mandating curricula and capitulated to the demands of some of the more aggressive states that wanted to fund balanced literacy programs. The facts also indicate that DoE did what they could to make sure as few non-phonics-based reading programs as possible.
3. Funding not limited; all reading programs funded. This is not a reasonable reading of the statute, but some members of Congress and the Inspector General believe that some reading programs were unfairly excluded from funding. Given that DoE pretty much funded all the phonics-based reading programs, this would mean that all reading programs were fundable for some to be unfairly excluded. This is a ridiculous interpretation, by any standard, but it appears to be the standard that the IG used in finding in its reports. And, those reports were parroted by various Democratic congressmen for partisan ends -- just in case you think your congressman is more concerned with making sure "at-risk" are properly taught with federal dollars than throwing them under the bus when political gains could be made.
With this background in mind, you can read any Reading First media story, any congressional press release, or any of the IG reports and make a fair evaluation of what actually went on. Did DoE violate the law?
We don't know for sure. The IG didn't find any actual violations in which a program was improperly excluded, was improperly include, or in which anyone in a position to make a funding decision made one that benefited him or herself.
At most, DoE played fast and loose with the law and regulations in an effort to maintain as much control over the process as possible. That activity was technically improper by an reasonable standard even if DoE thought they needed the control to fend off hostile publishers and states that were trying to get funded improperly. The conditions appear to be in place for violations to have occurred, but oddly, the IG failed to find any actual violations.
But that's for you to decide. If you find an actual violation let me know.
Slow Output
This time it's not for wont of other work or laziness.
I have three partially completed big-picture posts in varying states of completedness. I'm not quite done wrapping my head around them quite yet.
Hopefully, they will be worth the read when they do come.
patience.
May 9, 2008
Today's Question (redux)
NCLB is vague in its requirements with respect to measuring reading ability. NCLB is concerned with "proficiency in ... reading or language arts." That is the extent of the guidance given to us.
We know that reading is a complex skill comprising various subskills and content knowledge. But, what does it mean to be a proficient reader? What standardized test or battery of tests exist that accurately measure the "reading" ability of children and whether they are proficient?
Further, under NCLB it is the educators whose performance is being measured, even though the students are the ones taking the test. So the testing instrument must not allow educator subjectivity and must not be capable of being gamed by the educator. For example, Elizabeth's example in the post below describes a test that can be gamed by an educators since students can be taught to memorize the words appearing on the test and, thus, the test is not a true reflection of reading ability.
So, pretend you are a new superintendent of a school district who wants to accurately determine the reading ability of the children attending the schools in your district and how well they are being taught. So, for example, you want to know that your third graders are reading at a third grade level and will be capable of reading at a fourth grade level next year. You get to pick the standardized test(s) to be used. You will have non-reading-specialists monitoring the administration of the test(s). The monitors can identify outright cheating by teachers and/or students but nothing more subtle than that, i.e, they are incapable of making substantive determinations related to reading of any kind Otherwise, the administration of the tests is out of your control. Only the results of the test(s) will be reported to you.
What assessments do you select and why?
May 6, 2008
Interim Reading First Study
Reading First was a product of political comprise. Instead of limiting grants to research validated programs reading curricula, the Reading First was watered down to permit reading programs "based on" scientifically based reading research. In reality, all this meant was that publishers needed to provide curricula appearing to have "explicit and systematic instruction in phonemic awareness, phonics, vocabulary development, reading fluency, and reading comprehension strategies.
Reading Recovery is the antithesis of this approach to teaching reading. Yet look at what those clowns managed to show (pdf).
Designing effective reading instruction targeted to at-risk kids requires an orchestration of minute details and variables.
Throwing a few disjointed phonics exercises in your previous whole language program is not effective instruction, though you'd likely be able to make a case that it falls within the statutory language of "explicit systematic instruction in ... phonics" since that undefined term is all but meaningless. Oh sure, such instruction is going to be successful with many "advantaged" kids, but a broken clock is correct twice a day as well yet we don't say that this clock tells good time.
Let's be honest, many "real" phonics programs only perform marginally better than phony phonics programs. Phonics is not a magic wand that can be waved over a reading program and make it effective instruction. Phonics is a tool. A tool that can be wielded many ways, only some of which are effective with "at-risk" children. And phonics is only one smallish part of an effective reading program.
Prior to Reading First the major education publishers were not exactly cranking out high quality instructional programs--nor were they known for their ability to design effective instruction. Then an opportunity came along to grab a larger share of the reading curricula market by putting out a product that would be selected by all those schools with Reading First grants. All that needed to be done was to redesign your program so that it appeared to comply with the undefined statutory language of Reading First. Not exactly a Herculean task.
There are an infinite number of ways to design a reading program that complies with the Reading First statute. The probability of any of the major publishers stumbling upon one of the few effective combinations is pretty slim, especially considering their previous track record. And, the probability of a school selecting one of these newly-cranked-out reading programs from one of the major publishers and seeing improved results is similarly slim.
Then there are other problems.
Even if a publisher did stumble upon an effective program, the chances that a school would actually implement it with fidelity is slim as well. There's a reason why the few reading programs that have been validated by research tend to be scripted: without the scripts, schools would screw them up.
And selecting a generic reading comprehension test as your measure of achievement is going to be mostly testing student IQ/SES and the amount of background knowledge they've acquired which arguably has little to do with reading ability.
Add all this up and the only conclusion you'd expect from evaluating Reading First schools as a whole is going to be a null set. That's what the interim study appears to have found. Lest you forget, that's what Project Follow Through found as well. Most of the Follow Through schools failed to achieve positive results. In fact, almost all of the Follow Through schools showed negative results. Reading First is only somewhat less of a failure than Project Follow Through. But that's only if you look at the programs as a whole.
It's a fair bet that when you look at individual reading programs, some of the Reading First programs will show significant positive results. One of the Project Follow Through programs showed positive results.
In education we expect a preponderance of losers. Education is not yet a mature profession. It's not even a profession. We're not going to see improvement until we make a concerted effort to separate the winners from the losers, scrap the losers, fund the winners, and find effective means of identifying and developing new winners. The Federal Government has already attempted to go down this route twice and failed both times. The Democrats tried it with Project Follow Through and the Republicans tried it with Reading First. In both cases political forces overwhelmed and weakened the attempts, returning us back to the status quo.
I expect the same outcome with Reading First. Some winner might be identified in the final report, but that outcome will be overwhelmed by the overall failure of the program as a whole. History will no doubt repeat itself again the next time we spend lots of money on a fancy grant program. You can count on that.
You're kidding yourself if you think there will be a governmental/political solution to our education woes. That's not the way the world works. It works the other way.
April 19, 2008
Reasoning and Writing Cont'd

The students have been taught a procedure for testing the hypothesis set forth in the passage and writing a two paragraph essay analyzing the hypothesis. It's not surprising that you don't see anything like this coming out ed schools.
April 18, 2008
Reasoning and Writing
Levels A and B get students ready for real writing. Level C concentrates on narrative writing. Level D focuses on expository writing. Levels D through F are no walk in the park, even for higher performing kids. In fact, I'll go so far to say that most people are never taught and never learn the skills taught in the latter levels of RW.
Most people are simply unable to critically example text or a presentation of information and construct a coherent argument based thereon. Reading blogs and the comments section makes this fact abundantly clear. And the demographic that engages in such activities is highly skewed at the top of the cognitive curve.
What I like about RW is that it teaches grammar and logic in the context of writing. As students learn grammar they immediately incorporate what they learn into their writing assignments. Every lesson has a writing assignment.
For example, by level C, lesson 89, students are learning how to properly use pronouns in writing (jack and Jill = they), how to punctuate a series of nouns (bat, ball and glove), and how to construct paragraphs with multiple people speaking (a paragraph per speaker). The writing assignment for this lesson is based on the following series of pictures.

The assignment is for the students to draft a multiple paragraph narrative based on the sequence of events that take place in pictures 1, 2, and 4, including what must have happened in missing picture 3.
The students start off by setting the scene by writing where the people were, what they were doing, and what was happening in the background in the first picture.
Ann, Maria, and Tony were sitting on a bench at a bus stop. A house was burning behind them and smoke poured out of the window. Tony said, "I smell smoke."
Next, the students write about what happened in the second picture.
They stood up and faced the building. Ann pointed to the window and said, "It's coming from over there."
Tony said, "There's a dog in that window."
Finally, the students write about what must have happened in the missing third picture and what happened in the last picture. The students have learned to identify the differences between the second and fourth picture to determine what must have happened in the missing third picture. In this example some differences are the location of the children, what they were doing, and the presence of Mrs. Wilson.
They went over to the window. Ann climbed onto Tony's shoulders, reached into the window, and grabbed the dog. Ann held the dog in her arms. Mrs. Wilson arrived home holding a bag of groceries. She said, "You saved King."
As she ran to a telephone booth, Maria said, "I'll call the fire department."
After the students complete the writing exercise, the class discusses some examples from the students' writing. The teacher notes some areas that the students should have included in their writing and some common errors. The students fix-up their writing and turn it i tothe teacher for review. The teacher reviews the writing by the next lesson and reviews the previous assignment in the beginning of the next lesson.
By the end of level C (lesson 110), students will be able to construct simple multi-paragraph narratives based on a sequence of events like this example.
As you can hopefully see, this lesson has far more instructional value that the typical journaling exercises elementary school students typically engage in.
And, yes, I wrote those paragraphs all by myself. The question is will my second grader do a better job than me when he does this exercise tonight.
Update: Here's what the boy wrote for this lesson. I only corrected a few spelling errors.
Ann, Maria and Tony were sitting on a bench at a bus stop. There was a house on fire behind them. "I smell smoke," said Tony.
Maria, Ann, and Tony stood up and looked behind them and saw a house on fire. "It's coming from over there", said Ann.
"There's a dog in that window," said Tony.
They ran to the house and Ann got on Tony's shoulders. Then Ann grabbed King while old lady Wilson came with groceries saying, "You saved King."
Maria ran to the telephone booth saying, "I'll call the fire department."
April 17, 2008
More Gering Data
I know the in DI, they hit language instruction hard. Think of it as the "content" course you need before you get to the other content courses. Language is content.
Not unsurprisingly Gering is posting better language results.
Take for example example these scores from the Terra Nova Language test which show Gering fifth graders who received 3 yrs of DI language arts instructions compared to Gering seventh graders who did not.

That's even better than the reading scores.
How about some scores from Gering second graders on the Gates-MacGinitie Word Knowledge test.

The 2007 cohort received three years of DI instruction (grades K-2), the 2006 cohort received two years (grades 1-2), the 2005 cohort received 1 year (grade 2). Either the 2007 cohort had a lot of smart kids compared to the other cohorts or they learned a lot more words.
There's also been a lot of talk about NCLB causing schools to neglect the non-reading and non-math classes, such as social studies and science. Certainly, in DI a strong focus is placed on English Language Arts and Reading, but Gering doesn't appear to be suffering in these areas.

These are Terra Nova Science and Social Studies for 3rd and 4th grades. Remember you need to be able to read the tests in order to correctly answer questions. That point seems lost on many edu-pundits.
Finally, we turn to writing scores. The following scores show the improvement that Gering Hispanic students made in comparison to other Hispanic students in Nebraska as each cohort received more DI instruction in writing.

Gering went from significantly performing below the state average to performing better than the state average.
April 16, 2008
$24,606
Andrew Coulson of the Cato Blog has crunched the numbers for D.C. and came up with a per pupil spending number of $24,606. That's obscene.
There are quite a few people that still think that you can improve education outcomes by just throwing a few more dollars at the problem. I don't think these people realize just how much money they'll need to throw at the problem because $24,606 doesn't appear to be enough.
Today's chart courtesy of data from the U.S. Census should give you an idea just how much money we continue to throw at the problem year in and year out.

These are total expenditures in constant dollars (i.e., adjusted for inflation) for public schools. In other words, this is what we actually pay per student based on daily attendance numbers. Basically, the amount we spend on public education has doubled in real dollars since the early seventies.
Has any other product or service doubled in price since then?
Not gasoline.
April 15, 2008
New Features
I also added my first feed link to bloglines, an excellent reader and the one I use. Here's a fully functioning copy. Go ahead click it, I dare you.
If you have a bloglines account, clicking on the button should add the highly coveted d-ed reckoning feed to your bloglines feeds.
Mind you, the feeds have existed for quite some time. Now, they will hopefully be easier to subscribe to.
I'll add more buttons as time permits.
Update: Done and Done.
Tests
I know that a lot of very knowledgeable people read this blog. And, just because I disagree with some readers/commentors doesn't mean that I don't think they aren't knowledgeable. (How's that for a triple negative?)In any event, I'd like to know what you think are decent standardized tests for academic subjects such as reading and math. I'm especially interested in hearing your opinions on testing the various aspects of reading ability.
What tests are good for measuring the mechanics of reading, such as decoding ability, vocabulary knowledge and fluency? Are there any reliable tests. how about the end-product of reading instruction -- reading comprehension?
What about math? What's a good test for skills students should possess at the end of elementary school and/or to see if they are ready for algebra?
What about tests for history, geography, science?
Be sure to list the tests' weaknesses along with the advantages.
More Results Out of Gering
Thanks to the DI implementation and lots of hard work by the district, Gering Public Schools has managed to close some of the achievement gaps between Hispanic students and white students. In Gering, Hispanics represent a substantial minority of students, nearly 30%.
Here are Kindergarten scores from the DIBELS phoneme segment fluency test for Hispanic and white students before and after the DI implementation.

As you can see, the percentage of Hispanic students passing the test as increased by an amount sufficient to close the achievement gap with the white students whose pass rate has also increased.
Here are second grade scores from the DIBELS oral fluency test for Hispanic and white.

Again, note that white students have made significant gains compared to previous cohorts, but Hispanics have made even greater gains--gains sufficient to create a reverse achievement gap.
Bear in mind that closing the gap with respect to the number of students passing the benchmark is not the same as closing the achievement gap with respect to absolute achievement scores. It could be that the scores of white students are still higher than the scores of Hispanic students. I don't have data to report either way, at least not yet. But, keep in mind that NCLB is concerned with the closing the gap between the percentage of students meeting benchmarks, not with absolute scores.
This is how NCLB is supposed to work. Schools are supposed to be improving instruction such that student achievement is improved for all groups with the effect that more students from lagging groups will pass the benchmark and close he gap. This is how it is working in Gering.
(Continued)
Cost of Public Education Rising Faster Than the Cost of Gasoline

The chart shows that the retail price of public education per pupil has risen faster than the retail price of a gallon of gasoline.
I like the chart because it puts the steep rise in gasoline prices in contrast to the even steeper price of public education. And, of course, I can choose to forgo gassing up my gas-guzzlin' SUV, but I'm paying the price of public education whether I want to or not.
PS: You can always spot an effective post by the presence and strength of a Stephen Downes' argument in the comment thread. Stephen seems to be compelled to answer every blog post he disagrees with, a Sisyphean task if ever there was one. Unfortunately, he doesn't always have a good/persuasive argument to respond with.
April 14, 2008
Some Results Out of Gering

The above graph shows DIBELS Nonsense Word Fluency (NWF) and Oral Reading Fluency (ORF) test data for grades K-6 from the spring of 2004 (before DI) and in the spring of 2007 (after 3 years of DI). These DIBELS tests are good predictors of the risk of student reading failure in subsequent grades. The graph shows the percentage of students meeting the benchmark goals. Meeting the benchmark indicates a low risk of reading failure.
The students in K-2 have been in DI since they began school. The students in gardes 3-5 have received three years of DI, i.e., for example the fifth grade students received DI in grades 3-5, but did not receive DI in grades K-2. The students in grade 6 only received a single year of DI in sixth grade.
As you can see from the graph, all grades, despite many students not receiving DI in each gradehave made substantial gains and their risk of reading failure has been significantly reduced. The kindergarten class of 2007, for example, is performing above the top 1 percentile based on these scores.
Here is a graph of grades 1-5 showing the performance of economically disadvantaged students on the same tests.
Not unsurprisingly, in 2007, the low-SES students in grades 1-5 outperformed the low-SES students in 2004. What is surprising is that the low-SES students in 2007 also outperformed all 2004 students in each grade, not just the low-SES students. That's impressive. So much for socio-economic status being a determining factor of academic success. With effective instruction, the predictive value of socio-economic status is diminished.
In case you were wondering how this performance translates to reading performance. Below is a graph of Terra Nova Reading scores for the fifth grade class of 2007 (which received only three years of DI instruction) with the performance of three cohorts of seventh graders who did not receive any DI instruction. Terra Nova is a nationally normed standardized test.
As you can see from the graph, the fifth graders who received three years of DI outperformed the seventh graders who did not.
Continued in third post.
April 10, 2008
Gering Public Schools: The School District to Watch

Gering, a smallish 2,000 student district with 30% ethnic minorities (mostly non-ELL Hispanics) and with 43% of students on free/reduced lunch, was an underperforming school district back in 2002 when new district administrator, Don Hague, decided to do something drastic (and smart) with the district's new Reading First grant. Hague didn't just adopt a new research-based basal reading program for Gering, he didn't even adopt a research-validated reading program, he went whole hog and adopted a district-wide research-validated whole-school reform. Like I said, a smart move because most administrators would only have done the bare minimum needed, having the least amount of changes, to give the appearance they're doing something to solve the problem. Real reform requires more serious effort.
Hague also wisely chose, Direct Instruction (DI), as his whole-school reform and implemented the reform with the assistance of the National Institute of Direct Instruction (NIFDI). It's a wise choice because DI has a proven track record of success in grades K-3 along with some longitudinal data for grades 4 and 5. By adopting the whole-school version of DI, there is strong likelihood that Gering students will acquire all the fundamental skills they need for learning content area sbject matter starting in sixth grade.
After three years, Gering is already starting to see results.
Before implementing DI, there was a 23 point gap between Hispanic and white students in fluency benchmarks in second grade in the Gering. Last year, not only was the gap closed, a greater percentage of Hispanics met the fluency benchmark than did white students -- a -2% gap.
The district, not content waiting for the elementary students receiving DI to reach junior high, adopted the remedial DI reading program for its junior high students to improve their chances of succeeding in high school. After one year of remediation, Terra Nova scores went from a 39% pass rate to a 55% pass rate. That's an effect size of about 0.4 standard deviation (σ). Let's put that in perspective.
A 0.4σ improvement is about 60% better than the gains made in the Project Star class-size reduction study. And, in order to get similar gains via improving teacher effectiveness, you'd have to replace teachers with an average effectiveness at the 50th percentile with super teachers performing at the 94th percentile.
But, the best part about the intervention in Gering is that it will be producing a great deal of longitudinal data. Gering has four elementary schools. Each elementary school, and only those schools, feeds into a single junior high school. That junior high school, and only that junior high school, feeds into the sole high school. Hopefully, Gering will stick with the intervention for the next ten years or so, so the longitudinal data can be collected.
I've been in contact with Gering and NIFDI trying to get some additional data for the intervention and plan on reporting any results and answers I can get. In the meantime, take a look at the short documentary on Gering that was produced. Note in particular, the interviews with the teachers and the responses they give, especially with respect to the children's reactions to the intervention.
Continued in Second Post.
April 9, 2008
April 7, 2008
Your Pet Reform is Suckier Than You Think
We are not happy with student performance in the U.S. Which is to say, we are not happy with our schools' ability to educate.So back in 1965 we, as a nation, passed the Elementary and Secondary Education Act (ESEA) to improve the state of education. The ESEA established the Department of education to distribute funding to schools and school districts with a high percentage of students from low-income families. These became known as Title I schools. The ESEA lacked a real accountability provision and so schools were not accountable for achieving any results with the federal funds. Not unsurprisingly, those results were not forthcoming.
In 2001 we decided to ameliorate that deficiency by reauthorizing ESEA to include an accountability provision. We renamed ESEA to No Child Left Behind (NCLB) and increased funding to cover the new accountability provisions (i.e., standards setting and yearly testing in grades 3-8 and 11 in math and reading (and now science)).
Let's put the problem in perspective.

This graph represents our baseline student performance back in 2001. Let's set our goal such that half the students met the goal back in 2001. (For those of you familiar with NAEP, this goal falls between the proficient and basic levels of performance). The blue shaded area under the curve represents the percentage of students who met the standard.
(Another way of looking at this is to pretend that all students took a standardized test back in 2001, we normed the test to get a normal distribution of performance with 50% of the students passing and 50% failing, then we froze that test. Subsequent improvement would result in more than 50% of students passing the test. If you think this is a Lake Woebegone effect, you don't understand the effect.)
The goal of NCLB was to have virtually all students meet this standard. This would have required an improvement of over 2 standard deviations (σ). And, thus, the chase began trying to find an intervention that would improve student performance, hopefully by at least 2σ. Seven years later this remains the national focus, at least among those who haven't given up yet.
But, here's the problem. Most people don't seem to understand how much improvement is actually needed to comply with NCLB. How much is a 2σ improvement? A lot more than most people think.
In fact, some, perhaps many, think that a real 2σ improvement is impossible. Let's accept that premise for the time being and set our goal a bit lower. What would it take for a typical Title I classroom to perform as well as an average classroom? Here is the performance of a typical Title I classroom.

In the typical Title I classroom, only 20% (blue shaded area) of the students meet the standard. In order to have this Title I classroom perform as well as the typical classroom another 30% (yellow shaded area) of students would have to meet the standards. This represents an improvement of about 0.84σ, otherwise known as a large effect size.
What does this mean? It means that if we improved all schools across the board by about 0.84σ, then only about 20% of students nationwide wouldn't meet the standards. In other words, 80% of students would meet the standards. Now a 80% pass rate doesn't comply with NCLB; but, let me let you in on a little secret. States have found ways to cheat by lowering their standards and their cut scores such that if we loosened the NCLB requirements say to a 90% to 93% pass rate, we'd be within spitting distance of meeting NCLB requirements. But, the problem remains of how to squeeze out about 0.84σ of real school improvement in the first place.
There is, of course, no shortage of opinions as to how to improve schools. It seems that everyone has their pet reform that they think is going to be some sort of magic educational bullet. The problem is that most of these educational bullets are being shot out of pop guns.
For example, let's take the favored reform of most edu-commentators: class-size reduction. The theory is that by reducing class sizes down to ridiculously small (13 to 17 students per class) and ridiculously expensive levels than student achievement will improve. In fact, student achievement will tend to improve, just not very much. Certainly less than these commentators think. In most reasonable rigorous experiments, such as Project Star, gains from class size reduction were found to be almost 0.25σ. Not much. Here's a graph to show you how little of an improvement that really is.

See that red sliver? That's the amount of improvement you can expect to see from class-size reduction. Not much. By reducing our typical Title I classroom down to Project Star levels we can expect to raise student achievement by a whopping 8%, from a 20% to 28%. Break out the champagne, kids, it's time to celebrate!
We have a name for interventions that achieve effect sizes of less than 0.25σ -- not educationally significant. This is a realization that in the real world, such interventions will likely have little or no effect in student achievement. For example, Project Star was plagued with many methodological flaws that would serve to inflate the already small effect size it achieved under experimental conditions.)
But let's not pick on class-size reduction reforms too much. The sad fact is that about 95% of all education reforms fail to achieve even the small effect sizes achieved in Project Star. This means that most education reforms fail to achieve educationally significant effects. Now go back and look at the graph again. See the read sliver which represents the smallest educationally significant effect size (0.25σ)? The red sliver for almost all educational reforms is even smaller than that red sliver shown on the graph. Wrap your head around that. And, make sure you keep that in mind the next time you tout your pet education reform. It sucks. Now you know it; stop pretending that you don't.
Now let's briefly leave the world of reality and entering the realm of fantasy. A fantasy world where statistical correlation is the same as causation. This is the land of Kozol. This is where everyone who thinks that raising student socio-economic status (SES )will lead to student achievement. It's also the land where those who think that improving teacher effectiveness is the be-all-and-end-all of education reform. It isn't. Statistical correlations aren't reality, no matter how much you want them to be.
Let's pretend for the sake of argument there's some magic potion that could increase teacher effectiveness by 2σ. To put this in perspective. This would raise the effectiveness rating of an average teacher (50%) to a super teacher (90% effectiveness rating) and would raise a 25% teacher to a 75% teacher. Using data from this study, you can see what kind of improvement we might expect from these new magical super teachers.

See the slightly larger red sliver? That sliver represents a 0.35σ effect size. An effect size that is educationally significant. By, an effect size that still misses the goal by 19 percentage points, i.e., 41% of students failed to improve sufficiently in response to the super teachers. Achieving a 2σ increase in teacher effectiveness is a pipe dream. Even achieving a 1σ improvement is probably a pipe dream, especially when you consider that the study that looked into this question failed to find a correlation between any of the typical things (credentials, experience, etc.) thought to be associated with teacher effectiveness and increased student performance. With only a 1σ improvement, however, the effect size (about 0.26σ) becomes educationally insignificant.
Bear in mind that many pet reforms relate to increasing teacher effectiveness. Paying teachers more is an attempt to increase teacher effectiveness. Raising teacher prestige is an attempt to increase teacher effectiveness. requiring greater credentials is an attempt to increase teacher effectiveness.
Which finally brings us to the reason why improving student achievement by 0.84σ (a large effect size) is within the realm of possibility. That would be the little heard of Project Follow Through, the largest education experiment in U.S. education history in which one intervention, the Direct Instruction (DI) intervention actually achieved gains of at least 0.84σ, often more.

Notice the large red slice and the lack of a yellow slice indicating no shortfall. If only your pet education reform worked as well as this one.
Update: Teach Effectively has a related post and a link to an analysis of some of the few interventions, including effect sizes, that work. Go check them out.
Update II: Brett from the DeHavilland Blog has outed himself as a closet Vanilla Ice fan. I'm sure this was a difficult and painful decision for Brett and his family. The world needs more true heroes like Brett who aren't afraid to speak truth to power.