What happens when you relax accountability
We continue our debate on accountability’s role in reading reforms, celebrate incremental progress, and more.
Hi, folks. I hope everyone is back from spring break and ready to wonk out. Those of you at ASU+GSV, do let us know how the techies are reacting to the screens-in-schools backlash.
Today we round up responses to Rachel Canter’s firsthand account of the Mississippi Miracle, question whether we’re over-teaching phonics, explore the “science of math,” examine a solution to grade inflation, and hear Jed Wallace celebrate the charter school sector’s “marginal revolution.”
Sign up to receive this newsletter in your inbox, usually on Tuesday and Friday mornings. SCHOOLED is free, but a few linked articles may be paywalled by other publications.
Last week, I lauded Rachel Canter’s Atlantic article on the Mississippi Marathon—as well as her longer PPI report on the same topic—for stressing that, as I put it, “reading reforms don’t work without accountability.” I also included some commentary from Karen Vaites arguing that accountability for results—in the form of A–F school grades and the like—can backfire by pushing educators to freak out and do stupid things in response.
Today we’ll keep the debate going, first with more thoughts from Karen (via her Substack), a response from Tom Larkin, and finally a take from Matt Yglesias (from his Substack) about why he thinks accountability is so critical.
For certain, Rachel is correct that accountability played a role [in Mississippi’s progress], especially inasmuch as she puts screening and retention (two reforms I emphasize, too) into the accountability bucket. Sarah [Mervosh] notes the role of accountability shifts, too.
But I have the same reaction to the “It’s Accountability” thesis anywhere I see it.
An accountability theory of change assumes that schools know what to do to raise outcomes, and if we just put the right carrots or sticks in place, they will do it. I don’t actually believe that’s the case, writ large. We’re in an education ecosystem where educators receive (and often believe) loads of misguided signals about what works to improve outcomes.
Put another way: If states implemented new, tougher accountability schema today, schools would be just as (or more) likely to embrace the newest faddish tech-enabled solutions (“just like iReady, untested for efficacy, but now AI-enabled so it’ll work this time!”) as they would to embrace better curricula, like those in Louisiana and Tennessee.
If reforms don’t ultimately influence instruction, I don’t expect gains. I hope we all agree about that. As someone who talks with a lot of teachers and school leaders, I don’t see much evidence that teachers sit in classrooms sweating the state’s new A–F school rating system, turning such efforts into a motivator. SorryNotSorry, accountability superfans.
I also struggle with the idea that Mississippi is the promised land of accountability when we don’t have a window into the curriculum used by each district. For me, transparency is an important enabling condition of accountability. Most states fall very short on curriculum transparency, and Mississippi’s one of ‘em.
Ultimately, I don’t over-emphasize state accountability approaches in my own Southern Surge writing because it didn’t emerge as one of the consistent pillars of the work across the four states with reading gains, and I am focused on replication, not single-state successes.
If state accountability systems were a big part of the reform efforts in those states and I missed it, it’s because educators on the ground weren’t talking about it. Which kinda makes my point.
All of that said: Early reading screening appears to both inform and motivate. It’s especially valuable when the screening information reaches parents, so everyone knows which kids are below-benchmark, starting in kindergarten. Third grade retention will remain controversial, but it does seem to change adult behaviors. I’ll keep focusing on these two parts of the accountability pie because they seem to move the needle.
Tom Larkin
Management 101 would say: “You get what you measure.” You don’t have to assume that someone/some organization “knows” what to do and just isn’t doing it. Accountability for results is far more basic than that. It’s great if they do know what to do, but if they don’t, accountability will force them to figure out what will improve results and then do it. That may be a long process with a whole bunch of “well, that didn’t work” but eventually even casting about in the dark will turn up something that improves things.
If I had an employee who knew what to do to get better results, but wasn’t doing it, I would conclude that their incentive structure is screwed up somewhere—maybe their definition of “better” is different than mine, maybe something is blocking them, maybe they don’t care about results—like there is no accountability.
Actual accountability would fix all three. Part of the accountability would have to be to define the “better” results, if held accountable but blocked people will always bring the blocking thing to management’s attention in clear and specific terms, if the accountability has any teeth they will care.
Part of basic accountability is setting the measurement and demanding results against it. The people doing the job should be the ones to figure out how to do it better—they know the job better than anyone. None of this is easy. Defining better is hard, and people will perform against the metric even if it’s wrongheaded if they have to. They will also work at manipulating the measurement in any way they can. I’ve been on countless sales teams who quickly realized it was way easier and more lucrative to “manage the goal” than to raise sales.
If there is no incentive to deliver better results, then knowledge of what will enable better results is not useful. Kind of neat knowledge to have, but not useful.
The reason I keep trying to drag these conversations about smartphones and ed tech and phonics and even critical race theory back to the older debate about accountability is that I think the background conditions are really important.
Schools are complicated ecosystems that involve a lot of different individual human beings inside the building, as well as a broader set of stakeholders like parents. Whole school systems are even more complicated. There are dozens of decisions being made, lots of different pressures to do this or do that, and tons of links in the implementation chain where things can fail. A state legislature can, with really good motives, decree that henceforth literacy instruction will follow the science of reading. But what actually happens at the level of school boards and superintendents and principals and second grade teachers?
If you have a system that is seriously trying to measure “Is our children learning?“ and create an incentive structure with benefits and consequences depending on the answer, it becomes that much more likely that each decision made along the chain will point in the right direction.
When you relax accountability, the reverse happens. It’s not that suddenly great teachers become mediocre or that good curricula become bad. But you’ve taken a thumb off the scale of every single decision—from whether to close schools in the snow to what technology products to buy to how to handle parental complaints—that said “We need to care about the results here.” And as a result, bad things start happening, especially for kids whose families are themselves less educated and less focused on education.
When writing SCHOOLED, I always debate how much to feature articles about instructional practice. I’m a policy wonk, and I assume that most—though not all—of my readers are too. I don’t want to get over my skis, offering takes on subjects for which I have no expertise, especially when there are folks out there like Holly Korbey and Kim Marshall who are much better positioned to tackle questions of teaching and learning.
Yet it would be nuts for those of us making and studying policy not to worry about whether our prescriptions are helping or doing harm in the real world of classrooms. So in that spirit, I encourage you to dig into Liana Loewus’s brilliant Education Next article on “over-teaching phonics.” Because for us wonky-wonks, we need to understand: Reading science is complicated, the details matter, and there are lots of ways this can go sideways.
Liana writes:
To argue that schools are over-teaching phonics is to risk being seen as a naysayer—someone who believes the tide will inevitably and rightfully turn back to more “balanced” practices that de-emphasize using systems for teaching phonics.
I’ve spent much of my career as an educator and journalist standing up for science-based instruction, and that has nothing to do with what I think will or should happen. Not a single researcher I spoke with questioned the need for explicit teaching on how the code works. Where they questioned current practices—and disagreed with each other—was on the exact skills and dosage for instruction.
“Over-teaching is happening in a few main ways,” argues Liana:
(1) spending too much time on less impactful skills,
(2) teaching extraneous skills and patterns, and
(3) teaching content that only the teacher needs to know.
Real harm can follow from this over-teaching, both in the form of bored, disengaged kids and reduced time for knowledge building, which is what students need to develop the ability to comprehend what they’re reading. We policy wonks need to make sure we’re not encouraging over-teaching phonics when we’re trying to fix the problem of under-teaching it.
Meanwhile, Danielle Hankins argues that “math needs its ‘science of reading’ moment.” Plenty of other writers have called for us to “do with math what we’ve done with reading,” but Danielle’s is probably the best article I’ve seen when it comes to getting the analogy right.
In many [math] classrooms, discovery-first instruction has become common practice. Students are encouraged to generate strategies, explore patterns, and construct algorithms before core procedures are secure. The intention is deeper understanding. But teachers frequently report student frustration, uneven mastery, and widening gaps between those who enter with strong background knowledge and those who do not…
Cognitive load theory explains why students without a strong background struggle. Research in cognitive science has long demonstrated that working memory is limited. When learners encounter too many unfamiliar elements at once, retention declines. Experts can manage complexity because they have an organized knowledge structure that was built over time. Beginners do not. Without such structures, tasks that might appear to be engaging can become cognitively overwhelming. In mathematics, asking students to derive procedures while interpreting new representations and comparing multiple solution paths increases mental demand and cognitive load significantly. What is often described as “productive struggle” in practice exceeds students’ processing capacity.
Research comparing minimally guided instruction with explicit approaches consistently shows that novice learners benefit from clear modeling and guided practice. When teachers demonstrate procedures before expecting independent generation, students build understanding more efficiently. As foundational skills become fluent, students are better able to reason flexibly. Procedural fluency and conceptual understanding should not be viewed as competing goals. Fluency reduces the mental effort required to execute basic steps, paving the way for deeper thinking.
Holly Korbey is glad that Danielle sees the connection between the science of reading, the science of learning, and the science of math. Plus:
Some states and districts have seen the connection: Alabama and Louisiana have made considerable math reforms aligned with that evidence, especially in elementary school; Kansas has created its own teacher training in evidence-based math teaching methods; Maryland is currently training every teacher about how brains learn and the basic learning principles that also apply to math.
But as with reading, the math wars are not going to go quietly into the night. Holly points to recent comments from Stanford professor Jo Boaler—whom some have rightly called “the Lucy Calkins of math”—for casting doubt on the science of math.
Boaler…encourages her customers to read a “research paper” examining the science of math group: “The Science of Math Reconsidered: A Critical Examination of Foundational Claims,” by professors Kate Raymond and Melissa Gunter. They claim that learning procedural math is a tool of authoritarianism, and accuse the science of math scientists of “promotion of skill development as the sole purpose of mathematics education.”
Ugh, here we go again.
A few months ago, we dug into the issue of grade inflation and what policymakers might do to address it. Now Alex Tabarrok takes to the Marginal Revolution blog to report on an elegant if technocratic solution—one focused on colleges but that could certainly work for high schools too: “achievement indexes based on relative comparisons.” This is better, he argues, that the sort of “grade cap” that Harvard is now contemplating, limiting the number of students who can receive A’s in any given course.
The underlying issue is informational. A grade tries to capture two things—student ability and course difficulty—with a single number. Gans and Kominers show that in general this is impossible: If some students take math and earn B’s while others take political science and earn A’s, there is no way, from grades alone, to tell whether the difference reflects ability or course difficulty.
There is, however, a solution in some cases. Clearly, if every student takes some math and political science courses, informative patterns can emerge. If math students tend to get B’s in math but A’s in political science, while political science students get A’s in their own field but C’s in math, you can begin to separate course difficulty from student ability.
Students don’t all overlap the same classes. But full overlap isn’t necessary—you just need a connected network. If Alice just takes math courses, Joe takes math and political science courses, and Bob just takes political science courses, then Alice and Bob can be compared through Joe. With enough of these links, the entire system can be stitched together. The more overlap, the more precise the estimates.
Admissions officers at highly selective colleges and universities probably do this sort of thing already, at least informally, understanding what an A signifies at various high schools and in various courses, and being more impressed with students who attend highly-competitive schools, take the toughest courses, and manage to maintain sky-high GPAs. But making this explicit instead of implicit might benefit everyone by encouraging students to seek out greater challenges and empowering teachers to give honest marks.
Speaking of Marginal Revolution, Jed Wallace argues that the charter school movement is education’s best example of the marginal revolution in action:
I wanted to frame this week’s conversation around the idea of “marginal revolution.” By that, I meant a way of thinking that Tyler Cowen, joined by other economists at George Mason University, has through decades of writing under that banner helped popularize: The idea that steady, incremental improvements can compound over time into changes large enough to feel revolutionary....
My contention is that the charter school movement’s steady progress over three and a half decades is amounting to a kind of marginal revolution in public education.
Furthermore,
That kind of persistence matters because progress compounds over time. Our task is not to lose heart over the fact that revolution has come marginally. It is to keep going, to keep figuring out how to make progress at the margins, and not to let our eye be taken off the ball by the shiny objects that so much of the education world allows itself to be distracted by.
A younger Mike used to pooh-pooh the idea of incremental progress; I suspect a younger Jed did, too. What we need is radical transformation! Now in middle age, both of us appreciate that incremental progress can add up to transformation if we stick with it long enough and avoid going backwards. This is the path, fellow reformers!
New data from the National Center for Health Statistics show that teen birth rates continue to fall, reaching a historic low in 2025. It’s great news, since children born to teen parents tend to struggle. Yet this dramatic decline—11.7 births per 1,000 female teens in 2025 compared to 61.8 per 1,000 in 1991—makes recent school achievement drops even more disappointing because the teenaged pregnancy trend should be serving as a significant tailwind to improved student outcomes. —Selena Simmons-Duffin, NPR
A new AI bot links up with learning platforms to find and complete assignments without any student involvement; its creator calls it a wake-up call reflecting the reality that educators need to rethink assignments. As AI companies race to lock in young users, both students and educators worry about the impact on learning. “Even as growing numbers of students are using the technology, a majority believe that the more they use AI for classwork, the more it will harm their critical-thinking skills,” writes Lila Shroff. —The Atlantic
Sal Khan admits that his much-hyped AI tutoring chatbot hasn’t revolutionized student learning. As he tells Matt Barnum, “AI is going to help…But I think our biggest lever is really investing in the human systems.” It’s a sentiment remarkably similar to thoughts shared last week by Google’s Head of Learning (covered here in SCHOOLED). —Chalkbeat
See you Friday!
—Mike








one thing i really like in the first half article is that you put a lot of different intellectuals in conversation with each other, in order to show the audience that even though these people all disagree on the notion of “how important is increasing accountability” there is actually a lot of common ground between them because they’re all defining “accountability” differently.
The most important accountability is student/family accountability.