February 24, 2009

Rebranding our way to better schools


Our new Education Secretary, Arne Duncan, wants to fix our education woes by, wait for it:

“Let’s rebrand it,” he said in an interview. “Give it a new name.”


How about the We Still Don't Know How to Spend the Money Effectively Act?

That gets right at the heart of the matter. It also implies a solution: find out how to spend the money effectively.

This seems to me like the logical starting point. So, who knows of any federal programs with the goal of finding out what predictably works backed by evidence?

February 18, 2009

Today's Video

Classroom Management presentation by Saul Axelrod.



Lots of good tips. Long.

Direct Link
.

February 17, 2009

Alfie Kohn and the Murray Gell-Mann Amnesia effect Part II

Continuing on from Part I.

We're now getting into the last prong of Kohn's main argument, after which he appears to take a kitchen sink approach and throws in all the remaining negative information he could find.

Kohn's last argument is based on a logical fallacy:

Finally, outside evaluators of the project – as well as an official review by the U. S. General Accounting Office – determined that there were still other problems in its design and analysis that undermined the basic findings. Their overall conclusion, published in the Harvard Educational Review, was that, “because of misclassification of the models, inadequate measurement of results, and flawed statistical analysis,” the study simply “does not demonstrate that models emphasizing basic skills are superior to other models.”


As a preliminary matter, Kohn fails to mention that the "outside evaluators" he's referring to, House et al. were funded by the Ford Foundation which had also funded a few of the losing models in PFT. As such, Kohn's source has "potential bias" issues which Kohn fails to alert his readers to. Kohn also fails to alert his readers to all the other similar "outside evaluators" which analyzed both the PFT data and House's analysis and came to a different conclusion. These other outside evaluators are no more biased than House, so its curious to see why Kohn would fail to mention them.

Next comes the first of Kohn's logical fallacies. Kohn commits the fallacy of division (or whole-to-part fallacy) when Kohn claims PFT "simply 'does not demonstrate that models emphasizing basic skills are superior to other models.'" Even if the basic skills models as a whole weren't superior to the other models doesn't mean that the DI model alone wasn't. The data certainly shows that the DI model was the superior performer.

The point here is that Kohn is criticizing DI yet has already started to veer off and is trying to mislead readers by dragging in information pertaining to other programs or the more general classification of basic skills programs.

Next Kohn makes an appeal to authority, another logical fallacy, when he mentions that the results were "published in the Harvard Educational Review." I'd be willing to cut Kohn some slack had he mentioned all the other journals of the research he cites. But this is the only one he cites. Coincidence? What I do know is that he failed to mention another study on PFT that was also published in the Harvard Educational Review that affirmed the findings of PFT. Another convenient oversight.

Kohn again buries the weakest the parts of his argument in a footnote. Kohn parrots the House study's findings which is based on a reanalysis of the PFT data. This reanalysis was not without it's own problems as set forth in another study by Bereiter et al which reanalyzed the House reanalysis (It should be mentioned that this study has the same potential bias problems as the House study since Bereiter had professional ties to DI before PFT).

Let us therefore consider carefully what the House committee did in their reanalysis. First, they used site means rather than individual scores as the unit of analysis. This decision automatically reduced the Follow Through planned variation experiment from a very large one, with an N of thousands, to a rather small one, with an N in the neighborhood of one hundred. As previously indicated, we endorse this decision. However, it seems to us that when one has opted to convert a large experiment into a small one, it is important to make certain adjustments in strategy. This the House committee failed to do. If an experiment is very large, one can afford to be cavalier about problems of power, since the large N will presumably make it possible to detect true effects against considerable background noise. In a small experiment, one must be watchful and try to control as much random error as possible in order to avoid masking a true effect.

However, instead of trying to perform the most powerful analysis possible in the circumstances, the House committee weakened their analysis in a number of ways that seem to have no warrant. First, they chose to compare Follow Through models on the basis of Follow Through/Non-Follow Through differences, thus unnecessarily adding error variance associated with the Non-Follow Through groups. Next, they chose to use adjusted differences based on the "local" analysis, thus maximizing error due to mismatch. Next, they based their analysis on only a part of the available data. They excluded data from the second kindergarten-entering cohort, one of the largest cohorts, even though these data formed part of the basis for the conclusions they were criticizing. This puzzling exclusion reduced the number of sites considered, thus reducing the likelihood of finding significant differences. Finally, they divided each effect-size score by the standard deviation of test scores in the particular cohort in which the effect was observed. This manipulation served no apparent purpose. And minor though its effects may be, such as they are would be in the direction of adding further error variance to the analysis.

The upshot of all these methodological choices was that, while the House group's reanalysis largely confirmed the ranking of models arrived at by Abt Associates, it showed the differences to be small and insignificant. Given the House committee's methodology, this result is not surprising. The procedures they adopted were not biased in the sense of favoring one Follow Through model over another; hence it was to be expected that their analysis, using the same effect measures as Abt, would replicate the rankings obtained by Abt. (The rank differences shown in Table 7 of the House report are probably mostly the result of the House committee's exclusion of data from one of the cohorts on which the Abt rankings were based.) On the other hand, the procedures adopted by the House committee all tended in the direction of maximizing random error, thus tending to make differences appear small and insignificant.


It is one thing to fail to mention that your source is potentially biased due to its financial ties to both some of the losing programs in PFT and the "outside" evaluators. But failing to mention that your source has been itself criticized as systematically adopting procedures which all have the effect of minimizing the direction of error, many with dubious or no scientific validity, in favor of the very programs with the corporate ties is quite another. The prudent advocate would at least attempt to explain these criticisms away, but in any event, readers should be made aware of these significant infirmities in your underlying studies.

Kohn concludes his main argument with a gratuitous swipe:

Furthermore, even if Direct Instruction really was better than other models at the time of the study, to cite that result today as proof of its superiority is to assume that educators have learned nothing in the intervening three decades about effective ways of teaching young children. The value of newer approaches – including Whole Language, as we’ll see -- means that comparative data from the 1960s now have a sharply limited relevance.


I bet Kohn wishes he could take that crack about whole language (an educational philosophy so bad it's advocates had to re-brand it to disassociate it from the lengthy trail of negative data it amassed) back now.

And, what evidence is there that today's educators have learned anything since the mid 1970's (not the 60s as Kohn claims)? The longitudinal NAEP data tells a much different story.

Next Kohn gets into the kitchen sink part of his argument. Let's take each in turn.

  1. First Kohn cites some newspaper accounts on DI. Kohn admits these accounts are anecdotal. I agree with Kohn, but I'm wondering why he included them in his argument anyway. I'm guessing lurid innuendo. I could go into detail refuting the points Kohn recounts, but it's not worth the effort for anecdotes such as these.
  2. Next, Kohn makes another fallacy of division when he states "it’s common knowledge among many inner-city educators that children often make little if any meaningful progress with skills-based instruction." Of course, there's lots of data from inner-city schools pertaining to DI that show that this statement isn't true with respect to DI.
  3. Last, Kohn claims that there is "a lot more research dating back to the same era as the Follow Through project supports a very different conclusion" and then goes about citing various longitudinal studies which purport to show better long-term outcomes for various child-centered P-3 programs (which Kohn prefers) to DI programs. Apparently, Kohn hasn't heard of "confounding variables." And, chooses to ignre th efact that many of these studies were conducted by the sponsors of these programs. And, that some had serious methodological flaws and fails to mention that the High Scope study was also sponsored by the same people responsible for one of the worst performers in PFT.


Here's Kohn's big conclusion:

Still, with the single exception of the Follow-Through study (where a skills-oriented model produced gains on a skills-oriented test, and even then, only at some sites), the results are striking for their consistent message that a tightly structured, traditionally academic model for young children provides virtually no lasting benefits and proves to be potentially harmful in many respects.


This is only true for the studies Kohn has chosen to cite which are only the "negative" ones. Kohn has basically cherry-picked the studies and excluded all the ones showing positive effects, such as Gary Adams' meta analysis and the underlying studies. Kohn also cites his research uncritically and has ignored all the criticism directed at the studies he cites. He doesn't even offer an explanation, he simply ignores them. He also ignores all the potential bias problems and methodological flaws in his cited studies. Kohn seems to have an unrealistically high standard for DI studies and a very low one for research with conclusions he likes. I find it hard to believe that any study cited by KOhn in any of his books could withstand the constraints imposed by House et al. on the PFT data.

In short, Kohn's "hard evidence" against DI appears to be almost exclusively opinion, rather than fact. Little of this opinion is supported by data, though Kohn's use of selective quotes from "research" attempts to convey that impression. And, completely ignoring all the contrary evidence and presenting such a one-sided evaluation based on that cherry-picked "evidence" is reprehensible for anyone claiming to to be dispassionate.

In this analysis of DI, Kohn has shown himself to be an untrustworthy advocate and the same pattern of scholarly malfeasance is evident in all his writings I've read.

February 13, 2009

Alfie Kohn and the Murray Gell-Mann Amnesia effect

(the introduction can be found here)

In the comments of the recent Willingham-Kohn dust-up, edu-blogger Stuart Buck brought up DI and Kohn responded by citing this article by him which immediately reminded me of the Murray Gell-Mann Amnesia effect.

The late Michael Crichton once gave a speech describing what he termed the Murray Gell-Mann Amnesia effect.

Briefly stated, the Gell-Mann Amnesia effect is as follows. You open the newspaper to an article on some subject you know well... You read the article and see the journalist has absolutely no understanding of either the facts or the issues. Often, the article is so wrong it actually presents the story backward—reversing cause and effect. I call these the "wet streets cause rain" stories. Paper's full of them.

In any case, you read with exasperation or amusement the multiple errors in a story, and then turn the page to national or international affairs, and read as if the rest of the newspaper was somehow more accurate about Palestine than the baloney you just read. You turn the page, and forget what you know.


I know enough about the research on DI to know that Kohn's description of the DI research qualifies as one of the worst hatchet jobs in education policy advocacy. As such, it should serve as evidence that Alfie Kohn might not be a trustworthy source on education policy and that his analysis of education research should be closely scrutinized in order to stave off the Murray Gell-Mann Amnesia effect.

But don't take my word for it, let's review Kohn's description of the DI research.

After we get past an initial paragraph of over-heated inflammatory language, Kohn's first argument is related to the results of Project Follow Through (PFT) is:

Of course, even if these results could be taken at face value, we don’t have any basis for assuming that the model would work for anyone other than disadvantaged children of primary school age.


But, PFT involved some schools in middle class neighborhoods and middle-class kids took part and were evaluated as part of the research. In fact, a very diverse set of students were evaluated.

The DI model was the best performing model for "disadvantaged children" as Kohn acknowledges, but it was also the best performing model for high-performing students, White children, Native Americans, African-American students, Hispanic students, English language learners, urban children, rural children, and the very lowest of the disadvantaged children. So there is a basis, a good basis, for assuming that the DI model works for most primary school aged children. And, in fact, these results have been replicated numerous times in subsequent studies which Kohn also fails to acknowledge.

Kohn's next argument is a repetition of the "variability" argument that others have levied against the PFT results:

To begin with, the primary research analysts wrote that the “clearest finding” of Follow Through was not the superiority of any one style of teaching but the fact that “each model’s performance varies widely from site to site.”[1] In fact, the variation in results from one location to the next of a given model of instruction was greater than the variation between one model and the next. That means the site that kids happened to attend was a better predictor of how well they learned than was the style of teaching (skills-based, child-centered, or whatever).


This is spurious conclusion because most of the variability is attributed to the inclusion of two cohorts from the Grand Rapids site in the analysis which had severed ties with the DI sponsor well before the end of the study. This is well documented. The data from the Grand rapids site was the only "DI site" with low performance (a half standard deviation below the mean of the other DI sites). The Grand Rapids site is the only "DI site" that consistently fell below national norms. In fact, most of the variability in the remaining DI sites is above National norms.

In addition, subsequent research, involving researchers with at least some professional/reputational ties to DI, has shown that the variability between sites is mostly attributable to demographic factors and experimental error, and not to the DI program.

We disagree with both Abt and House et al. in that we do not find variability among sites to be so great that it overshadows variability among models. It appears that a large part of the variability observed by Abt and House et al. was due to demographic factors and experimental error. Once this variability is brought under control, it becomes evident that differences between models are quite large in relation to the unexplained variability within models.


In any event, even if the variability finding was characterized as the main finding by the primary researchers, this in no way diminishes the finding that DI was the superior performing program across the board for all measures tested for all groups tests. It's still a valid finding and has been upheld by numerous researchers examining the findings since the initial evaluation.

So, best performing program and the only program whose variability was mostly above national norms does not a valid criticism make. Strike two for Kohn.

Next Kohn, attacks the testing instruments used in PFT:

Second, the primary measure of success used in the study was a standardized multiple-choice test of basic skills called the Metropolitan Achievement Test. [(MAT)]


The MAT is not just a test of basic skills (such as Listening for Sound (sound-symbol relationships), Word Knowledge (vocabulary words), Word Analysis (word identification), Mathematic Computation (math calculations), Spelling, and Language (punctuation, capitalization, and word usage)).

It is also a test of cognitive skills as well. Several Metropolitan subtests measure indirect cognitive consequences of learning, such as the Reading subtest (which is, in effect, paragraph comprehension), the Mathematics Problem-Solving subtest, and the Mathematics Concepts test (knowledge of math principles and relationships).

This is important because Kohn goes on to claim:

While children were also given other cognitive and psychological assessments, these measures were so poorly chosen as to be virtually worthless.


Even if the other cognitive and psychological assessments were "poorly chosen" it does not diminish the fact that the MAT is a well respected test of both basic and cognitive/conceptual skills, as ackowledged by subsequent researchers. The other cognitive/conceptual skills test used was the Raven's Colored Progressive Matrices, but it did not prove to discriminate between models or show change in scores over time.

Also, the affective skills were assessed using two instruments: the Intellectual Achievement Responsibility Scale (to assess whether children attribute their success (+) or failures (-) to themselves or external forces) and the Coopersmith Self-Esteem Inventory (to assess how children feel about themselves, the way they think other people feel about them, and their feelings about school).

Kohn buries the reason why he believes the cognitive and affective skills tests were "poorly chosen" in a footnote.

There is strong reason to doubt whether tests billed as measuring complex “cognitive, conceptual skills” really did so. Even the primary analysts conceded that “the measures on the cognitive and affective domains are much less appropriate” than is the main skills test (Stebbins et al., 35). A group of experts on experimental design commissioned to review the study went even further, stating that the project “amounts essentially to a comparative study of the effects of Follow Through models on the mechanics of reading, writing, and arithmetic” (House et al., 1978, p. 145). (This raises the interesting question of whether it is even possible to measure the conceptual understanding or cognitive sophistication of young children with a standardized test.)


Let's take these "strong reasons" in order. Even if the primary analysts believed that these tests were "much less appropriate" this doesn't they didn't measure what they purported to measure. There is no evidence that they didn't. It could also be that the primary researchers believed that the basic skills tests were best suited for measuring what K-3 students are typically expected to know. In any event, the opinion that the tests were "much less appropriate" does not lead one to conclude that the tests didn't measure "complex 'cognitive, conceptual skills.'” This is an empirical question and Kohn provides no empirical support for his conclusion.

Next, Kohn relies on the opinions of the "experts" commissioned and funded by the Ford Foundation (which I'll get to later) for the proposition that PFT only measured the affects of "the mechanics of reading, writing, and arithmetic." Apparently, what these experts were getting at was that students who hadn't learned the mechanics of reading, writing, and doing arithmetic might not be able to demonstrate their cognitive skills. This is also an empirical question. But, these experts provided no empirical support for their conclusion. Other researchers, however, did look into the question once it was raised.

Conceivably, certain models-let us say those that avowedly emphasize "cognitive" objectives-are doing a superior job of teaching the more cognitive aspects of reading and mathematics, but the effects are being obscured by the fact that performance on the appropriate subtests depends on mechanical proficiency as well as on higher-level cognitive capabilities. If so, these hidden effects might be revealed by using performance on the more "mechanical" subtests as covariates.

This we did. Model differences in Reading (comprehension) performance were examined, including Word Knowledge as a covariate. Differences in Mathematics Problem Solving were examined, including Mathematics Computation among the covariates. In both cases the analyses of covariance revealed no significant differences among models. This is not a surprising result, given the high correlation among Metropolitan subtests. Taking out the variance due to one subtest leaves little variance in another. Yet it was not a forgone conclusion that the results would be negative. If the models that proclaimed cognitive objectives actually achieved those objectives, it would be reasonable to expect those achievements to show up in our analyses.


So, again we have no valid reason for discounting the results of the non-basic skills tests. Unsupported opinion is not a valid reason last I checked.

Lastly, Kohn raises a question:

This raises the interesting question of whether it is even possible to measure the conceptual understanding or cognitive sophistication of young children with a standardized test.


And then conspicuously fails to answer it. This is the poor man's version of debate. Moreover, nothing that precedes this "interesting question" is capable of actually raising it.Also , I'm not sure if Kohn is trying to claim that you can't measure these skills or that you can't measure these skills with a standardized test. Though, Kohn provides no support for either question.

Kohn's innuendo is that the students might have had conceptual understanding or cognitive sophistication that we must accept on faith despite the evidence that more students in the non-DI models were incapable of demonstrating these magical immeasurable skills on simple tests of comprehension of written paragraphs, mathematical problem-solving and knowledge of math principles and relationships, not to mention all the other "basic skills" that were measured.

Unfortunately, Kohn fails on all three counts to provide evidence that would compel a reader to follow him down his opinionated path that PFT only measured basic skills and that the other measures were "virtually worthless." Maybe this is why he buried this one in a footnote.

Next Kohn claims:

Some of the nontraditional educators involved in the study weren’t informed that their programs were going to end up being judged on this basis.


First of all, the DI educators were also non-traditional in as much as the other models' educators were. DI is about as far removed from traditional pedagogy as the other models.

Also, even if some of the other educators claimed that they were never initially told that their models weren't going to be judged on reading comprehension, math problem solving, and the like, they would have quickly learned what was coming down the pike since the PFT students were extensively tested throughout the study. And, it was the third and fourth cohorts that formed the cohorts of the evaluation. Whoever claims to not have known initially would certainly have found out during the time time the first two cohorts passed through.

Next Kohn claims:

The Direct Instruction teachers methodically prepared their students to succeed on a skills test and, to some extent at least, it worked.


Actually, the DI model systematically prepared their students to read, understand the conventions of language, to spell, and to do arithmetic with an emphasis "placed on the children's learning intelligent behavior rather than specific pieces of information by rote memorization." And, the students outperformed the other students on tests of sound-symbol relationships, vocabulary words, word identification, math calculations, spelling, punctuation, capitalization, word usage, paragraph comprehension,mathematics problem-solving , knowledge of math principles and relationships, and the affective measures. There was no evidence that the DI students engaged in test preparation as alluded to by Kohn.

PFT demonstrated, once again, that teaching these skills directly was more effective than teaching them obliquely which was what the other models believed would lead to superior performance. It turns out they were wrong and they continue to be wrong to this day.

For those keeping track at home, Kohn has now failed to establish the first two prongs of his argument. He has one more prong left which I'll take up in my next post.

February 12, 2009

The Intellectual Dishonesty of Alfie Kohn

In case you missed it, cognitive scientist Daniel Willingham recently criticised author Alfie Kohn for making factual errors, misinterpreting and oversimplifying the research, and making logical errors.

Willingham was too kind to Kohn.

Kohn responded and denied the allegations, casting most of the disagreements as merely a difference of opinion. Hopefully, Willingham will respond to Kohn's response and give him the smack-down he so rightly deserves because I, like Willingham, believe that Kohn is butchering the fair-reading of the research to lend credence to his crack-pot opinions and agenda.

Kohn is not the dispassionate advocate he pretends to be. He is a intellectually dishonest muck-raker with an agenda. A dangerous agenda for at-risk children.

You see, what Kohn does is prey on the sorry state of the quality of instruction and education research as a springboard for his opinions. For example, we know that praising students to increase motivation is difficult to do properly. It is difficult to get it right and easy to get it wrong and it is even more difficult to show positive academic results because those results are also dependent upon the quality of the delivered instruction which is often ineffective with at-risk kids, i.e., the ones who need the motivational praise. You see the problem --because Kohn doesn't. To Kohn, all praise or positive reinforcement is detrimental, unless you want to count Kohn's carefully worded weasel language he includes at the end of a long diatribe for plausibly deniability. Here's the weasel language that comes at the end of a long article informing the reader of how bad positive reinforcement is:

It’s not a matter of memorizing a new script, but of keeping in mind our long-term goals for our children and watching for the effects of what we say. The bad news is that the use of positive reinforcement really isn’t so positive. The good news is that you don’t have to evaluate in order to encourage.


Kohn ignores the large body of research in which the proper use of positive reinforcement was found to be effective in getting disruptive students to stop being disruptive so they can learn. What I haven't seen is a teacher of a classroom of disruptive kids following Kohn's advice and being able to get the classroom under control and then teach them effectively.

And, ultimately that's Kohn's main problem. He has lots of opinions on education, but no evidence of his opinions being put into practice and being effective. In fact, I'll go so far as saying that to the extent that his condoned practices have been actually been used, they've been failures. Miserable failures.

Look how poorly the Open education model and the other child-centered models fared in Follow-Through. That's some very inconvenient evidence for Kohn which he realizes and attacks. And, it's that hatchet job which I'll deconstruct in my next post.

February 10, 2009

The Cheese Stands Alone

Stephen Downes has finally seen the light as to the benefits of worked problems examples for inducing learning:

Today's newsletter is delayed a bit because I could not tear myself away from this wonderfully detailed set of instructions on how to make cheese. Makes me just want to go out and get myself some rennet. You'd probably have to practice a bit to really learn how to make cheese, but these instructions really look like all you'd need to get going. It's also important to have previously seen and tasted cheese, so you know what success looks like. (emphasis added)


At least when it comes to novice cheesemakers learning how to make cheese.

But apparently not for learning academic content which seems to require a more constructivist approach, according to Stephen.

Wouldn't the budding cheesemaker learn more from being handed all the necessary cheesemaking ingredients and provided the wonderfully engaging opportunity of floundering around making cheese on their own with some minimal guidance provided by the instructor.

I guess not. It doesn't work for chick-sexing apprentices? Why should it work for novice cheesemakers?

And for that matter why should it work for novice students of algebra?

Here's the analog in the algebra world. Behold Algebra: Structure and Method, Book 1, Dolciani (1981 Ed.). (Click to enlarge)



The classic worked product example for teaching how to solve simple equations using the multiplication property of equality (a page I selected randomly).

The "lesson" is followed by a few more example and then the student is provided the opportunity to practice what has been taught by working various oral, written, and open-ended problems, relevant to the lesson so they can " practice a bit to really learn" it.

This is the traditional way algebra is taught. Apparently, it's only "tedious lecture" and "rote learning." Of course if it were rote learning the student would only be able to solve 4x = 52 and would have to be taught 5x = 50 and 3x = 36. But as any good connectivist will tell you, the student should be able to generalize a solution for any similar problem fitting the pattern of the worked problem example after sufficient practice.

It seems to me that the primary difference between this method of learning and the constructivist method of learning is that the "wonderfully detailed set of instruction instructions" isn't provided to the student beforehand. The student is supposed to figure them out (i.e., construct) this knowledge for himself. At least that's the theory.

Of course, in the real world, even the constructivists would rather see the instructions beforehand.

Outliers

In his new book, Outliers, Malcolm Gladwell makes the case that opportunities and practice determine success. No doubt luck plays a part in success. And you'll get no argument from me regarding the need for lots of practice. But, Gladwell, greatly underestimates the role that innate talent plays in success. Mozart surely did practice a lot, but he was also very talented. The Beatles practiced quite a but during their Hamburg days, but John Lennon and Paul McCartney were also very talented songwriters. Talent is an important component. The talented make better use of their practice time and will be more successful, and hence more motivated, in their practice. Just ask this kid.

February 7, 2009

From the Department of Huh?

Comes the conclusion of this study out of Ohio State.

A study of college freshmen in the United States and in China found that Chinese students know more science facts than their American counterparts -- but both groups are nearly identical when it comes to their ability to do scientific reasoning.


But when you look at the researchers' description of the underlying study, you see that this conclusion isn't supported and leads me to question the researcher's own ability to reason scientifically at least in the domain of education.

The researchers administered three tests to incoming college freshmen from China and America who had just enrolled in a calculus-based introductory physics course.

The first test, the Force Concept Inventory, measures students’ basic knowledge of mechanics and the student's understanding of mechanics and forces. "The Force Concept Inventory is not 'just another physics test.' It assesses a student’s overall grasp of the Newtonian concept of force. Without this concept the rest of mechanics is useless, if not meaningless." (Force Concept Inventory, Hestenes, Wells, and Swackhamer, The Physics Teacher, Vol. 30, March 1992, 141-158).

The second test, the Brief Electricity and Magnetism Assessment, measures students’ understanding of electric forces, circuits, and magnetism.

The third test, the Lawson Classroom Test of Scientific Reasoning, measures generic science reasoning skills. You can see the kinds of questions on the exam of the appendix of this study.

The tests were given to Chinese students and American students. According to the researcher, in China, "every student in every school follows exactly the same curriculum, which includes five years of continuous physics classes from grades 8 through 12" and "schools emphasize a very extensive learning of STEM content knowledge" In the United States, "only one-third of students take a year-long physics course before they graduate from high school. The rest only study physics within general science courses. Curricula vary widely from school to school, and students can choose among elective courses" and "science courses are more flexible, with simpler content but with a high emphasis on scientific methods."

Keep those descriptions in mind because they'll be important for the conclusions drawn by the researchers.

Now let's turn to the results.

On the FCI, "[m]ost Chinese students scored close to 90 percent, while the American scores varied widely from 25-75 percent, with an average of 50." Clearly the Chinese students understand mechanics better than their American counterparts. One the BEMA, "Chinese students averaged close to 70 percent while American students averaged around 25 percent -- a little better than if they had simply picked their multiple-choice answers randomly." I guess all those Physics course helped the Chinese students understand physics, whereas all that emphasis on scientific methods at the expense of content didn't pan out so well for the Americans. These results are hardly surprising. Knowledge is domain specific and transference between domains is generally minimal.

On the Lawson Classroom Test of Scientific Reasoning "[b]oth American and Chinese students averaged a 75 percent score." So, the Chinese students were just as capable as the American students even though their course supposedly didn't emphasize "scientific methods" like the American students did.

The researchers, however, concluded:

Lei Bao, associate professor of physics at Ohio State University and lead author of the study, said that the finding defies conventional wisdom, which holds that teaching science facts will improve students’ reasoning ability.

“Our study shows that, contrary to what many people would expect, even when students are rigorously taught the facts, they don’t necessarily develop the reasoning skills they need to succeed,” Bao said. “Because students need both knowledge and reasoning, we need to explore teaching methods that target both.”


What? This isn't the conventional wisdom. The conventional wisdom is that that learning facts in a domain will improve the ability to reason in that domain. This wasn't tested in the study. What was tested in the study, via the FCI and the BEMA, was the students' understanding in the domain (physics) which was significantly higher for the Chinese students compared to the American students. Not surprisingly, the American students didn't understand much physics since they didn't learn many physics facts and their "scientific methods" instruction failed to fill the void. Constructivists take heed.

What the study also showed is that learning facts in one domain will not necessarily lead to transference to a different domain and an improvement in reasoning skills in general, whatever they may be (assuming they exist). Again, not a surprising outcome. But, the researchers' spin obscures this conclusion.

And here's the kicker.

Bao explained that STEM students need to excel at scientific reasoning in order to handle open-ended real-world tasks in their future careers in science and engineering.

Ohio State graduate student and study co-author Jing Han echoed that sentiment. “To do my own research, I need to be able to plan what I’m going to investigate and how to do it. I can’t just ask my professor or look up the answer in a book,” she said
.

The irony is that this physicist didn't do a very good job conducting an investigation in a foreign domain (education). If he wanted to know who was more capable of "handl[ing] open-ended real-world tasks" he should have tested this in a domain specific way. He should have given both groups open-ended real world physics problems and determined which group handled them better. I'm thinking it would have been the Chinese students.

And then we have the most unsupported conclusion of the study:

“The general public also needs good reasoning skills in order to correctly interpret scientific findings and think rationally,” he said.
How to boost scientific reasoning? Bao points to inquiry-based learning, where students work in groups, question teachers and design their own investigations. This teaching technique is growing in popularity worldwide.


The American students who presumably were instructed in inquiry-based techniques fared no better than the Chinese students in general reasoning ability. Inquiry-based teaching once again failed to show results. And it certainly did the students no favors when it came to the students understanding of physics in which they performed poorly.

I see nothing in this study that shows any benefits for inquiry learning. If anything, the study supports the notion that you can't teach general reasoning directly, both methods of teaching failed. What the study also clearly shows is the continuing importance of learning content if you want to understand something.

February 3, 2009

Whitmire Phones One In

Richard Whitmire riffs off the recent, and soon to be disappointing, Obama effect in his latest USA News editorial as a base to point out why he believes the "pathways to success" are not in place for NAMs to find academic success. The problem, however, is that Whitmire's claim that the pathways are "closing up" is all wrong.

First he gives the Grey Lady's education reporting way too much credit vis-a-vis the Obama Effect.

Even the New York Times weighed in with a story that made the Obama effect appear based on science (relying on a single study; am I alone in thinking that was sub-NYT standards?) by writing up a study claiming that black test takers upped their scores post-Inauguration Day, apparently the dividend of a "Yes we can" self-esteem movement.


Sadly, this kind of education reporting for the NYT is very much the rule and not the exception.

But on to the Whitmire's closing pathways.

First he claims that College is not sufficiently accessible. I'm not sure that's really the problem. State colleges already admit many students who aren't sufficiently prepared for college level work. Most of these ill-prepared students aren't going to make it out of college anyway, so I don't see accessibility as the problem; lack of preparation is the problem.

Whitmire recognizes this lack of preparation as a problem, but unfortunately blames the wrong culprit:

The stimulus bill proposed by the House would bump up Pell grants for poor students to make college more affordable, but that does not solve the biggest problem faced by these students: As a result of attending subpar high schools, they are not ready for college work.


The achievement gaps are present long before high-school. It is debatable whether elementary schools have really improved, as Whitmire claims, but one thing is clear they haven't improve enough yet. Middle-schoolers remain woefully unprepared for high-school level work, so why are we blaming high-schools for being unable to deal with all these ill-prepared children?

Next Whitmire claims that "[n]ational education reforms have pushed curriculum demands lower into the grades, handing kindergartners the verbal tasks that two decades ago confronted second graders." This is only partially true. Today's kindergartners are still doing the same stuff that many kindergartners of twenty years ago did. Teh only real difference is that back then we allowed the struggling students to wait until they were "developmentally ready" which has been proven to be a large waste of valuable academic time. Yet the problem remains that we are still often not too successful in teaching these at-risk kids. In this respect Whitmire has a valid point and literacy rates will have to soar for there to be an improvement.

WHitmire's next point that black boys need to be rescued is also a valid point, as long as if by rescued he means to provide them with the effective commercially available curricula that has existed for decades.

Last, Whitmire jumps on the teacher quality bandwagon with both feet:

Knowing what we know about the value of a high quality teacher, we should be on the verge of delivering those teachers to inner-city students.


Really? I though the "research" pretty much indicated that we don't have the foggiest idea how to make average teachers into superstar teachers. Actually do know how to improve the effectiveness of all teachers: hand them a effective curricula and teach them how to use it, but this isn't what most people mean when they talk about teacher quality.

Ultimately I disagree with Whitmire's major premise that the lack of pathways are what's holding back students. The pathways have been in place for all children of a certain ability level and family stability to take advantage of. It is the access to those pathways that need to be improved to accommodate a level of student ability that has never been able (or had the opportunity) to take advantage of them up until recently.

Today's Quote

Practice makes you good at learning, but being smart makes you good at practice.*

-- me (this morning)



*A less pithy, but more accurate quote might go: Practice makes you good at learning, but being smart increases the likelihood of initial success which increases your motivation to practice, and, thus, your willingness to practice.

January 30, 2009

Pennsylvania's High Remediation Rates

The Pennsylvania Department of Education released some data last Wednesday indicating that one in three Pennsylvania high school graduates who enrolls in a state-owned university or community college must enroll in remedial math and/or English courses before they are capable of taking college-level courses.

As usual, Pennsylvania has released its data in a way that makes it all but unusable for analysis purposes. basically, they just tell us the percentage of students in each district that took remedial courses, the number of courses taken, and the amount it cost. Not very helpful. But not to worry, fearless readers I did the heavy lifting of putting the data into my all-purpose Pennsylvania database of education statistics and managed run a few regressions.

It should be noted that there are selection bias issues out the wazoo because PA failed to disaggregate any data. So take that as a large warning in interpreting the data.

Anyway, off we go.

The first regression I ran compared the percentage of adults with bachelor degrees in the district and the percentage of students needing remediation. Adult education level is a proxy for student socio-economic status (SES) (and parental/student IQ). The thought is that students with higher SES levels should be better prepared for college level work. Let's see if that hypothesis holds up. (Note that I've indicated the Philadelphia School District as the alrge red triangle. And, note that Philadelphia generally appears above the regression line indicating that its actual remediation rate is generally higher than its predicted rate.)



No, the hypothesis doesn't hold up. Less than 1% of the variance in parental education level is associated with the percentage of students in need of remediation. We would probably have gotten a better result if we knew the actual parental education level of the students who were in need of remediation since the parental education levels in my database are for every adult in the district (not just parents and not just parents of remedial kids).

Now let's check if the percentage of students receiving free and reduced lunches (another proxy for poverty and low-SES) correlates with the percentage of remedial students.

Nope. Once again the variance is less than 1%. I was somewhat surprised at this outcome. Low-SES students are generally underperformers so I was expecting to see districts with high numbers of low-SES students have higher remediation rates. But that wasn't the case. Maybe the selection bias effects are showing up here in that the low-SES students may be applying to college at lower rates. This analysis would have benefited from the demographic data of the actual remedial students. I'm certain PA has his data. Release the data, PA!

Now let'see if school expenditures make a difference on remediation rates.

Not really. Again the variance is low -- 3.6%. If anything, school expenditures are negatively correlated with remediation rates. The more a district spends the higher the percentage of remedial students it gets. Instead of dreaming up various explanations for this outcome, let's just keep it simple and say that school expenditures don't seem to have much of an effect on remediation rates.

Now let's look at whether there is a correlation between the percentage of students failing PA's NCLB test (the easy minimum-skills PSSA test) and the remediation rate.

(Note: the x-axis caption is incorrect. It should read: % failing state test)

Again, the variance is somewhat low at 8.7%. Schools with higher PSSA pass rates produce slightly lower remediation rates. I guess that's not too surprising. Students who pass the PSSA probably know more and as a result are less likely to require remediation. This is another analysis that would have benefited greatly from knowing the percentage of remedial students that passed or failed the PSSA exam.

These results aren't terribly interesting since they basically are just showing low correlations. But the following regression is a bit of a doozie.

Let's take a look at the percentage of non-white students and the remediation rates. Note that in PA, the number of Asian and Native American students is very low, so you can read non-white as mostly black and Hispanic.

Hello! Finally a decent correlation. 25% of the variance in remediation rates is associated with race. The more black and Hispanic students in the district the higher the remediation rates.

I'm wondering why the PA Dept. of Ed didn't highlight this inconvenient piece of data. PA schools continue to do a poor job preparing black and Hispanic students for college. PA schools with high percentages of black and Hispanic students fared poorly with remediation rates. Yet, oddly, schools with high numbers of low-SES, regardless of race, didn't fare much worse than schools with low percentages. And, spending does not seem to affect the remediation rates, if anything the more schools spend the higher the remediation rates.

I'm not sure what to make of all this, but PA could have done a much better job analysing and presenting the data. But, i suppose it's not in its best interest to do so.

January 27, 2009

Gates Foundation still has a lot to learn about education

The Gates Foundation has squandered a tiny bit of Bill Gate's personal fortune on some very silly ideas on education.

Based on Bill Gate's 2009 Annual Letter it's apparent that trend is going to continue in the near future.

First, Bill sets a silly goal:

Our goal as a nation should be to ensure that 80 percent of our students graduate from high school fully ready to attend college by 2025.


The country's drop-out rate is presently higher than 20%. This means that at least every student who now completes high school must be capable of doing college level work.

Now consider this: Pennsylvania just found out that a third of its students who enrolled in a state school or community college were not prepared for college level work. Those are the kids who thought they knew enough to go to college. How about the ones who knew better to apply in the first place? About half the students know better than to even try.

Bill, just upped the ante big time on NCLB's lofty goals. It's one thing to claim that you'll get 100% of students to a loosely defined standard (i.e., "proficiency") based on the standard made up by each state, measured with a testing instrument devised by each state (the fox guarding the hen house), and with cushy safe-harbor provisions. But, it's another to set a high and easily measurable standard that can easily be verified (Hello remediation rates).

And one more thing, Bill. At least have the good sense to give us a cut point that will now result in a large achievement gap between the races/classes. An 80% cut point will leave us with a 23% achievement gap absent we finding the educational magic that allows the low group to gain relative to the high group. Good luck with that. The cut point has to be at least at 95% to get the achievement gap down to a politically acceptable level.

Next, Bill asks an easily answered rhetorical question:

Unlike scientists developing a vaccine, it is hard to test with scientific certainty what works in schools. If one school’s students do better than another school’s, how do you determine the exact cause?


There are basically two effective ways to do this:

Type 1 approach. The most obvious way would be to ask and check. Because successful applications have been created by design, ask the designer of the successful program what the variables are that the design controls and why. The answers imply tidy, controlled experiments that involve systematic investigation of the designer’s assertions.

Type 2 approach. A related approach would involve fully implementing a successful instructional program and then systematically altering the details of it (one at a time, while trying to maintain the others as they were in the original program). The changes would be correlated with changes in student performance. If no difference results from a change, the dimension that was changed does not function as a variable (at least within the range of variation observed). A manipulation that results in improved student performance identifies a variable that was not well designed by the original program. A change that results in inferior student performance identifies a variable that was designed better by the original program than it was in the modified program. This approach would need clear descriptions of what constituted improved performance. Efficiency is an important variable. If the change resulted in improved performance but required three times the instruction of the original, the rubric for judging efficiency would have to compute the ratio of improved performance over the time to arrive at a reasonable overall judgment of the net “desirability” of the change.


of course, this isn't the way we currently evaluate research. But that's a whole 'nother problem.

Lastly, Bill tells us how his foundation is all ready to jump to the next poorly researched education fad: teacher effectiveness.

It is amazing how big a difference a great teacher makes versus an ineffective one. Research shows that there is only half as much variation in student achievement between schools as there is among classrooms in the same school. If you want your child to get the best education possible, it is actually more important to get him assigned to a great teacher than to a great school.

Whenever I talk to teachers, it is clear that they want to be great, but they need better tools so they can measure their progress and keep improving. So our new strategy focuses on learning why some teachers are so much more effective than others and how best practices can be spread throughout the education system so that the average quality goes up.


Bill's rocking a dead baby with this new teacher effectiveness initiative of his. Five minutes on Google Live Search would have told him that.

here's some free advice, Bill. Next time some education expert claims they know what they are talking about, you need to follow Ronald Reagan's advice -- trust, but verify.

This brings us full circle back to the need for ascertaining what works in education research:

Type 1 approach. The most obvious way would be to ask and check. Because successful applications have been created by design, ask the designer of the successful program what the variables are that the design controls and why. The answers imply tidy, controlled experiments that involve systematic investigation of the designer’s assertions.


Never trust what anyone who in education tells you.

January 26, 2009

Creativity and Little Content

The Pittsburgh Regional Future City Competition is a charming example of what you get when you try to instill creativity in students without first teaching them the relevant underlying content knowledge.

The competition challenges middle school students to design a city of the future with a focus on water conservation, reuse, and renewable energy. The students use the game SimCity (Deluxe 4) to help them build their three-dimensional models to scale. They have a semester to dream up and then construct their miniature cities entirely out of recycled materials. Supposedly, this inspires them to consider engineering as a profession.

Let's see what we get after all that creativity.

When it comes to the perfect place to live, [L.U.R.E., a sprawling metropolis set in southern New Mexico] would seem to have all bases covered.

School is free for everyone, brought into individual homes via a holographic teacher. Nearly everyone in town is gainfully employed as an engineer.

Mountain goat racing and sand surfing satisfy a yen for sports and leisure. And if, for no apparent reason, you need a getaway, there's the Space Shuttle Gilligan to whisk you on a four-month vacation to the moon...

L.U.R.E's coffee shop was housed in a Starbucks Frappuccino cup; office buildings were fashioned from paper towel rolls.


I'm sure there was some creativity is selecting a frappuccino cup as the coffee house building, but I'm wondering how that creativity generalizes to a domain outside of coffee cup miniature modeling. And try not to think about what the citiy's brothel was constructed from.

And what's up with the animal abuse? Enslaving our mountain goat brethen for our personal amusement seems a bit cruel. I'm surprised PETA didn't protest this event. Now that might have shown some real-world problem solving and creativity in how to defuse a PR nightmare without resorting to the firehoses.

I'm also wondering how many vacation goers would be willing to fly in the Space Shuttle. I'm guessing that the students hadn't heard of the Space Shuttle's propensity for blowing-up
unexpectedly. Hope they got insurance for that one.

At least the teacher hologram initiative shows creativity. Hologram's of a person trying to convey information or say, plans, to others is something I've never seen before. I wonder where they got that idea.

Look, I'm sure the kids had a lot of fun. But where's the educational value?

Here's what the kids were supposed to learn about:

The National Engineers Week Future City Competition offers students a resourceful way to learn about engineering.

Students will:

  • Learn how engineers turn ideas into reality.
  • Develop a project plan to guide team activities.
  • Use SimCity™ software to design their city.
  • Build a city model using recycled materials.
  • Work as a team under the guidance of an engineer and a teacher.
  • Demonstrate writing skills by composing an essay on an engineering design problem.
  • Enhance communications skills through a team presentation.

Did the students really learn how engineer's turn ideas into reality? Is a miniature model reality? If so, then isn't my kindergartner's pictures (media: crayons and construction paper) teaching her the same thing, just without the fancy labels?

This is not how engineer's turn an idea into reality. It doesn't seem to me that the students needed to know any actual engineering or any engineering constraints to construct their models. So, this is how a non-engineer turns ideas into reality. And, I'm not sure this exercise , in any way, generalizes to any real-world situation.

I suppose the kids did learn how to play SimCity. Videogames 101. That's what kids need -- more time playing videogames. I'm sure SimCity is a neat program, but it's not exactly a precursor to AutoCAD or other real-world construction/drafing programs.

And how does building a model out of recycled mterials generalize to building real stuff with recylced materials? Someone explain that to me.

The rest of it can be summarized as "learning how to work in a group." Something that our educators think students need a lot of practice doing for the real world. Apparently, lazy students need to refine their shirking skills from a young age and the more capable students need to understand the hell that awaits them in the real-world as they are expected to carry the load of the shirkers and share the credit.

All kidding aside, what does participation in this project actually teach that generalizes to anything else in a different domain? To the extent the students learned any generic creativity, explain how this creativity might generalize to a domain that requires knowledge in that domain without the student knowing that knowledge?

January 24, 2009

Nature vs nurture

Brian Caplan of Econlog has a good article on the value of parenting in the Chronicle of Higher Education. The article contains a good discussion of the nature/nurture argument and the adoption/twin studies.

There are two kinds of special families: those with twins and those with adoptees. If you want to disentangle the effects of nature and nurture, one approach is to compare identical twins, who share all of their genes, to fraternal twins, who share only half. Another approach is to compare adoptees to members of their adoptive families. If identical twins are more similar than fraternal twins, we have strong reason to believe that the cause is nature. If adoptees resemble members of the families they grew up with, we have strong reason to believe that the cause is nurture.

By using — and refining — these twin and adoption methods, behavioral geneticists have produced credible answers to the nature-nurture controversy. To put it simply, nature wins. Heredity alone can account for almost all shared traits among siblings. "Environment" broadly defined has to matter, because even genetically identical twins are never literally identical. But the specific effects of family environment ("nurture") are small to nonexistent. As Steven Pinker, a professor of psychology at Harvard University, summarizes the evidence:

"First, adult siblings are equally similar whether they grew up together or apart. Second, adoptive siblings are no more similar than two people plucked off the street at random. And third, identical twins are no more similar than one would expect from the effects of their shared genes."

The punch line is that, at least within the normal range of parenting styles, how you raise your children has little effect on how your children turn out...

Recent scholarship does highlight some exceptions [but] the fact remains that people tend to greatly overestimate the power of nurture.

If family environment has little effect, why does almost everyone think the opposite? Behavior geneticists have a plausible explanation for our confusion: Family environment has substantial effects on children. Casual observers are right to think that parents can change their kids; the catch is that the effect of family environment largely fades out by adulthood. For example, one prominent study found that when adoptees are 3 to 4 years old, their IQ has a .20 correlation with the IQ of their adopting parents; but by the time adoptees are 12 years old, that correlation falls to 0. The lesson: Children are not like lumps of clay that parents mold for life; they are more like pieces of flexible plastic that respond to pressure, but pop back to their original shape when that pressure is released.


Bleak news indeed for the SES warriors.

This is why we see substantial IQ gains for some preschool programs whose effects fade by the end of elementary school. This is also why time should not be wasted in elementary school doing "developmentally appropriate" nonsense. There is a brief window of opportunity that needs to be taken advantage of.

January 23, 2009

The real menace of 21st century schools


I think we all agree that schools should be using current technology and teaching students how to use current technology. Kid should know how to use computers, and the internet, and how to use various software packages, etc. I won't go on ad nauseum here; there are lots of education blogs that specialize in using technology.


Actually, what they do often do is oversell the academic benefits of this technology. I love the irony of all these smart and knowledgeable people writing entries for other smart and knowledgeable people to read telling how the only thing holding back dumb people from educational attainment is better access to technology. Like the ditch-digger down the street who can barely put together a coherent though is capable of getting much out of the internet as they do. There's a large gulf between those with only a little bit of knowledge and those with a lot and google isn't capable of bridging it.



But let's not fight about that. There'll always be other posts for that. Let's find something I think we all agree upon: that schools are abusing the call for 21st century skills to draw attention away from the bad job they are currently doing.

Take for example today's top story in the Massachusetts' Standard Times about how Fairhaven High School's new plan to make itself into a 21st Century school which is promised to "equip students with the necessary knowledge and skills to be successful in today's world." As usual the devil is in the details.

Schools seem to follow the same script. First they wipe the slate clean and give the impression that they were doing a marvelous job back in the 20th century, but that simply isn't good enough anymore.

"Even though we're facing these difficult economic times, our commitment to kids has to stay strong," said Fairhaven High School Principal Tara Kohler, who presented the plan to the School Committee last week.

"What we've always done isn't good enough anymore."


Was it ever good enough? I seriously doubt that. We call this change for the sake of change or, more specifically, change for the sake of not establishing a longitudinal track record of bad data. It's much harder to hit a moving target. And, what's the real difference between yesterday's fad and today's. Very little.

After the slate is wiped clean, the new plan is presented in incomprehensible edu-jagon designed to sound much more impressive than it actually is.

The district's plan is centered on five goals:

* Ensuring all students have access to a quality education.
* Preparing students for a 21st-century job market and a global economy.
* Improving students' transitions to high school, thus increasing graduation rates.
* Increasing opportunities for students to take college courses and participate in internships or other school-to-career activities.
* Improving access to technology.


What the hell does this even mean. How is the school going to transform itself into a 21st Century school. Here's how.

While the high school has already made some progress toward these goals — a new computer lab was installed earlier this year and a new transition program for incoming freshmen was implemented — there is still a lot that can be done, according to Ms. Kohler.



They're installing a new computer lab.

That is so 1980's.

There's other assorted nonsense in the school's plan (foreign languages, assorted "green" nonsense, "Virtual High School" online courses, and creating an iPod mobile lab), but none of it is on par with the stuff the edu-tech bloggers get all giddy about.

But don't worry about that the school is now a 21st century school.

Obama effect for reals according to NYT


Sam Dillon of the NY Times breathlessly reports some real educational magic today:

[The]performance gap between African-Americans and whites on a 20-question test administered before Mr. Obama’s nomination all but disappeared when the exam was administered after his acceptance speech and again after the presidential election.

The inspiring role model that Mr. Obama projected helped blacks overcome anxieties about racial stereotypes that had been shown, in earlier research, to lower the test-taking proficiency of African-Americans, the researchers conclude in a report summarizing their results.


This one was so hot the Times couldn't wait for peer review. No need to waste time with that. The Obama effect is for real and it must be reported right this second.

In fairness, Dillion does mention that the study had not undergone peer review, but only provides an incomplete and misleading explanation of prior research on stereotype threat which might have tipped readers off as to the dubiousness of this latest study.

Here are the money grafs from the Wikipedia entry on stereotype threat:

Furthermore, while Sackett et al. do not dispute the fact that stereotype threat has a real, measurable effect on test scores, they posit that in the part of the experiment where Steele and Aronson removed the stereotype threat, the achievement gap which did remain correlated closely with the existing African American - White achievement gap on large-scale standardized testing such as the SAT. In their own words:

Thus, rather than showing that eliminating threat eliminates the large score gap on standardized tests, the research actually shows something very different. Specifically, absent stereotype threat, the African American-White difference is just what one would expect based on the African American-White difference in SAT scores, whereas in the presence of stereotype threat, the difference is larger than would be expected based on the difference in SAT scores.


In subsequent correspondence between Sackett et al. and Steele and Aronson, Sackett et al. wrote that "They [Steele and Aronson] agree that it is a misinterpretation of the Steele and Aronson (1995) results to conclude that eliminating stereotype threat eliminates the African American-White test-score gap."


In the past researchers have been able to depress scores by introducing a stereotype threat (basically, the researchers told the test subjects that they were part of a group that were dummies). Removing the threat only brought scores back up to historic averages. They have not been able, however, to actually increase scores above historic averages.

This new research, in contrast, supposedly shows that test scores can be increased by such a large amount (somewhere between 0.5 to 1.0 standard deviations) to wipe out the achievement gap that exists between blacks and whites. Educationally speaking, that's a giant effect size and a truly unprecedented result (if true). Indeed, the Obama effect must be extraordinary to achieve such a result.

It's basically the educational equivalent of cold fusion. The Times apparently forgot about that lesson in journalistic humility. And so much for acknowledging legitimate opposing viewpoints that might cast some doubts on these extraordinary, unprecedented findings. Get a load of the expert the Times dredged up, replete with some nice spin supplied by the Times.

“It’s a nice piece of work,” said G. Gage Kingsbury, a testing expert who is a director at the Northwest Evaluation Association, who read the study on Thursday.

But Dr. Kingsbury wondered whether the Obama effect would extend beyond the election, or prove transitory. “I’d want to see another study replicating their results before I get too excited about it,” he said.


Kingsbury's wants to see replication -- a prefectly reasonable response. And the Times spins the failure to achieve replication as possibly being cause by "transitory" effects. The expert provides no indication that the results might be in doubt (in fact, he praises the study as "a nice piece of work") and the Times provides no indication that any doubt exists.

The impression I get from the Times article is that the Obama effect is for real, pending peer review, but might fade due to its transitory nature. The Obama legacy is already being written.

Keep your eyes on test results this summer. If we are to believe the Times the achievement gap should be eradicated due to the Obama effect. NCLB will turn out to be a smashing success. And, the world will be a happier place. Unless those nasty transitory effects dash our hopes once again.

January 22, 2009

Today's Video

Good video demonstration showing a real world physics application.

Poster Boy for the Continued Need for Spelling Instruction



Sometimes spell check doesn't cut it.

Where's your google now

I think my physics problem (and nine step solution) demonstrated how difficult it is to think critically about Physics unless you know quite a bit physics and have had quite a lot of practice solving similar physics problems. Your 21st century skills don't seem to be much help here now do they.

RWP also has a post demonstrating the same thing with not one, not two, but three business problems.

And, Pondiscio has, I believe the best post of the week showing how much of President Obama's inaugural speech you missed out on if you lacked the needed historical and literary content knowledge.

For all you connectivists out there I see a pattern emerging.

To paraphrase Edward G. Robinson -- "Where's your google now, nyahhhh?"