60 years of classroom research keep arriving at the same uncomfortable answer. We keep building our schools as if we never heard it.
You’re sitting at the back of the room with a clipboard. Management asked you to pop into Mrs. Middleton’s class this period — a quick observation, nothing formal, just part of the term’s walkthroughs.
It’s easy to see why she gets good reports. The room hums. Students are scattered around tables, each working on their own research project on space, each one clearly engaged. Rebecca is surrounded by library books, quietly filling page after page with detailed notes and drawings on the differences between Mars and Earth. Joel has a book open in front of him too, occasionally flipping a page, occasionally getting up with permission to step out for a moment.
You jot a few notes. Positive classroom climate. Students on task. Individualized learning evident. Teacher circulating and monitoring. Mrs. Middleton walks the room, checking in, offering help where it’s asked for. Nobody’s causing trouble. Nobody looks bored. By every visible measure available to you in that forty-minute window, this is good teaching.
You have no way of knowing that Joel hasn’t written a single word in twenty minutes, and won’t. That the “interruptions” he complains about to the teacher never actually happened. That he’s not reading the text in front of him — just looking at the pictures. You have no way of knowing that Rebecca already knew everything she’s writing down before this unit began; her father taught her the names of the planets when she was five. Neither student is learning anything new this lesson. And nothing in the room — not to you, not to Mrs. Middleton, not to any checklist either of you is carrying — shows that.
Here’s the part that should stop you: this scene is from research published in 2005, describing classroom findings the researcher had been collecting since the 1960s. Walk into almost any school today, hand someone a walkthrough form, and you’ll get exactly this. Positive climate. Students on task. Teacher circulating. Sixty years, and the form hasn’t had to change.
That’s not a coincidence. It’s the whole problem.
The Myth: If It Looks Like Learning, It Must Be Learning
Graham Nuthall spent 45 years trying to answer one question: what is actually happening inside a student’s mind while we teach them? [Nuthall, G., 2005. The cultural myths and realities of classroom teaching and learning: A personal journey. Teachers college record, 107(5), pp.895-934]. This post is about what we learn from his lifetime of pedagogical research.
By the end of his career, he’d landed on an answer education still hasn’t fully absorbed, and it’s blunt enough to state in one line:
Engagement is not evidence of learning. It never was.
We built almost our entire profession on the opposite assumption. A “good lesson” is one where students look busy, look interested, ask questions, produce a worksheet, finish on time. Teacher training centers on it. Observation checklists are built from it. Even the language of professional development — time on task, active learning, student engagement — treats visible participation as a stand-in for the thing we actually care about.
Nuthall’s life’s work is the record of watching that assumption fail, over and over, under increasingly close observation — and of a profession that, sixty years later, still hasn’t fully reckoned with what he found.
The Evidence We’ve Had for Decades
It started with something almost accidental: a tape recorder hung from a light fitting in a 1959 classroom, discussing a line of poetry.
Teacher: I wonder if anyone can tell me what time in history that is likely to be? Suzy? Suzy: About the French Revolution. They’re poor people. Teacher: Very good answer. I wonder if you could give me a further reason. Peter?
Question. Answer. Comment. It turned out this exact rhythm showed up in classrooms in different countries, different languages, even in shorthand transcripts from the early 1900s. Teaching, it turns out, isn’t really a set of choices a teacher makes each day. It’s closer to a script we all absorbed over ten-plus years of sitting in classrooms ourselves — performed fluently, without ever being consciously taught. One of the largest classroom-video studies ever run, filming hundreds of lessons across Japan, Germany, and the US decades later, reached the identical conclusion from a completely different angle: teaching behaves like a cultural ritual, remarkably stable within a culture and remarkably resistant to reform.
That script is comforting, because it’s legible. An observer — any observer, in any decade — can walk in, recognize the rhythm, and confidently call it good teaching. The trouble is that the script tells you nothing about whether anyone in the room actually learned anything, and once researchers finally found a way to check, the gap turned out to be enormous.
When Nuthall’s team built miniature microphones for individual students and mounted cameras on classroom ceilings — finally able to see what a single child actually experienced across a whole day, not just what the teacher or an observer could catch by eye — the first finding should unsettle anyone who has ever run or graded a walkthrough:
Even live observers keeping continuous written records of the behaviors of individual students missed up to 40% of what was recorded on the students’ individual microphones and the video cameras.
Trained adults, actively watching, missed almost half of what was happening. And what they missed wasn’t noise. It was a second classroom, running the whole time underneath the visible one — students whispering, passing notes, negotiating friendships, occasionally trading in casual sexism or racism, all while looking, from the front of the room, perfectly on task.
This is not a historical curiosity. This is the room you or I taught in this morning. The technology to catch it existed decades ago. The habit of trusting the surface is what never went away.
The Study That Should Have Ended the Engagement Myth
If one piece of evidence deserves to sit at the center of every staffroom conversation about “good teaching,” it’s this one.
A Grade 7 class studies Antarctica. A video tells them it’s the driest continent on Earth. Nuthall’s team quietly tracks three students — Joy, Teine, and Paul — through everything that follows: the video, a class discussion, a small-group task, another discussion two days later, a written report.
Joy encounters the idea four separate times, across four different contexts, and even argues about it with a confused classmate along the way. A year later, she remembers it clearly.
Paul also encounters it four times, but never has to actively use the idea. A year later he half-remembers it — the fact survives, but not the understanding.
Teine is passing notes during the video and lands in a different small group for the discussion. She meets the idea once. A year later, it’s gone entirely.
Same lesson. Same teacher. Same room. An observer watching that one class would have logged all three students as equally attentive, equally “on task.” The visible surface was identical. The actual learning was almost nothing alike — determined not by how engaged each student looked, but by how many genuinely different encounters each one actually had with the idea.
Tested against nearly 500 concepts across eleven students in three classrooms, this pattern held roughly 80–85% of the time. That’s not an anecdote. That’s a finding solid enough to build a teaching practice around, sixty years before most classrooms have caught up to it: a student needs several different, active encounters with an idea — not one clear explanation — before it survives past the lesson it was taught in.
The Answer Was Never Ability
Here’s where the myth gets most stubborn, because it has a built-in excuse.
If a classroom is well-managed and students are clearly busy, and some of them still don’t learn — what explains the gap? For a century, the answer has been ability. Some students are just more able than others; keep everyone equally on-task, and whatever’s left over must be down to that.
Nuthall’s data says otherwise, and says it plainly:
Given the same experiences, the less able students appeared to learn from their experiences in exactly the same way as the more able students.
The difference wasn’t in how students processed an idea once they met it. It was in how many chances they actually got to meet it — and the higher-achieving students were quietly creating more of those chances for themselves: asking more questions, talking more about the actual content with peers, using unstructured time more productively.
If background interests and prior knowledge determine what students learn, then differences in apparent academic ability are just as likely to be the product of differences in classroom experiences as the other way round.
Think of it less like ability and more like diet. The underlying digestive process is the same in every student. What differs is what each one is actually getting fed. A “less able” student isn’t digesting worse — they may simply be encountering less.
Where “Ability” Actually Comes From
Look closely at Nuthall’s own numbers and the picture gets sharper — and more useful. When his team tracked exactly where each learned concept came from for three real students in the Antarctica unit — one high-achieving, one mid-range, one low-achieving — a clear pattern emerged in how each one was getting their learning:
- The highest-achieving student relied least on teacher-led activities. Most of what he learned came from his own self-chosen reading, his own follow-up questions, and — notably — spontaneous conversations with classmates that had nothing to do with an assigned task.
- The lowest-achieving student relied almost entirely on teacher-managed activities. When the whole-class lesson didn’t happen to give her a concept, she very rarely picked it up any other way.
Neither of these students processed information differently once they encountered it — the earlier finding still holds. What differed was how much of their own learning each student was quietly generating outside of what the teacher had planned. One was fluent in turning downtime, curiosity, and peer chatter into more encounters with the content. The other depended entirely on the lesson landing correctly the first time, because she had no other route in.
And background matters here too, in a way most classrooms never account for. One of the students in Nuthall’s data already knew the names of every planet before the space unit even began — her father had taught her at five years old. That’s not talent. That’s five years of head start, delivered entirely outside school, that a single term of teaching was never going to equal out. Multiply that by every unit, every subject, every year, and “ability” starts to look a lot less like a trait a child was born with and a lot more like an accumulated record of how many encounters with ideas that child has already banked — at home, with parents, in conversations that had nothing to do with a classroom at all.
This is not a small distinction. It relocates the “achievement gap” away from something fixed inside a child, and towards something adults — teachers and parents — are actively shaping every day, often without realizing it.
A Note for Parents: What Actually Grows a Child’s “Ability”
If ability isn’t a fixed trait but a record of how many real encounters a child has had with ideas, then the question for a parent changes completely. It’s no longer “is my child smart enough,” it’s “how many different ways is my child actually meeting the things they’re learning?” That’s something a parent has far more influence over than any classroom ever will.
A few things Nuthall’s own data points toward, translated into what a parent can actually do:
Talk about what they’re learning . Don’t just check that homework is done. The students who learned the most weren’t just sitting through more lessons. They were the ones re-encountering ideas in conversation, well after the lesson ended. Ask your child what the lesson was actually about, and let them explain it messily. Explaining something out loud is, itself, another encounter with the idea.
Feed curiosity before the classroom does. The student who already knew the planets at five wasn’t unusually gifted. She’d simply had years of casual, low-pressure exposure at home before it ever became “content” to be tested. Documentaries, museum visits, dinner-table tangents, a curious question answered properly instead of brushed off. None of it looks like study, and all of it quietly builds the background knowledge a child will later be able to “hang” new classroom learning onto.
Let them talk to peers and siblings about what they’re learning. Some of the richest learning Nuthall recorded happened in unplanned peer conversation — kids explaining half-understood ideas to each other, arguing, correcting one another. A child who has someone to talk ideas through with — a sibling, a friend, even a parent willing to play devil’s advocate — gets more of these unplanned “extra encounters” than a child working in isolation.
Resist explaining something once and assuming it’s landed. The single most practical finding in this research: it typically takes around three separate, genuinely different encounters with an idea before it survives in long-term memory. If your child got it wrong on a test, the answer isn’t necessarily “try harder” — it may simply be that they only ever met the idea once. Revisit it from a different angle a few days later, rather than simply repeating the same explanation louder.
Notice what “busy” is actually measuring. A child who looks studious — sitting quietly with a book, appearing to do homework — may, like Joel, be doing nothing of the sort. It’s worth occasionally asking a child to explain what they just read or wrote, not to catch them out, but because that explanation is the only real evidence of whether anything actually landed.
None of this requires turning a home into a second classroom. It requires treating “ability” as something built through repeated, varied exposure to ideas — most of which has nothing to do with formal study at all — rather than a trait a child either has or doesn’t.
Back in Mrs. Middleton’s Room
So go back to where we started — you, at the back of Mrs. Middleton’s classroom, clipboard in hand.
Nuthall’s team tested what Rebecca and Joel actually knew, before and after that lesson. When they showed Mrs. Middleton the results:
Mrs. Middleton was shocked… She assumed our evidence was somehow aberrant or mistaken. After all, she knew the students, she knew how hard they worked, she knew how successful her program was.
She wasn’t wrong to trust what she saw. She was wrong to think it was the whole story. And the reason neither of you caught it isn’t a personal failing — it’s structural. Teachers and observers both learned to run a good one-on-one conversation, reading a face, catching a nod, and unconsciously assume that same feedback loop scales up to thirty students at once. It doesn’t. Nothing in most teacher training or observation frameworks ever corrects that assumption, because both are still built on the same premise Nuthall spent his career dismantling: that visible engagement is a reliable proxy for learning.
Mrs. Middleton is a successful teacher. The students like her and try to please her… But because Mrs. Middleton believes that the busy classroom is the learning classroom and that students’ behavior is a function of their ability, she has no idea what Joel or Rebecca are actually learning.
That single sentence is the myth, in full. And an outside observer, sent in specifically to catch what a teacher might miss, walked out having confirmed the exact same illusion — because a walkthrough built to check for order and pacing will always, reliably, find exactly what it’s looking for.
Sixty Years Later, the Answer Is Already Here
This is where most articles about education research end with a shrug — learning is complicated, more research is needed. Nuthall didn’t leave it there, and neither should we.
He gave a clear answer. Engagement is not learning, and no amount of watching a room from the outside will tell you the difference. The evidence for this isn’t new, isn’t contested, and isn’t waiting on more research to confirm it. It’s been sitting in the data since at least the 1970s. What’s missing isn’t proof. It’s the will to stop building our schools — our observation forms, our appraisal systems, our own sense of a lesson “going well” — on a proxy we’ve known for decades is unreliable.
So here’s the challenge, stated plainly rather than left as a gentle question:
Stop asking whether a lesson looked engaged. Start asking how many genuinely different ways each student actually encountered the idea you wanted them to learn — and whether you have any real evidence, beyond how busy the room looked, that it landed.
That’s a harder standard to meet in a forty-minute walkthrough. It should be. The easy standard is the one that let Rebecca coast and Joel disappear for an entire year, in a classroom everyone — teacher, principal, and observer alike — was certain was working.
Nuthall spent 45 years proving the myth wrong. We've had the answer for longer than most teachers reading this have been alive. The only thing left to decide is whether we're going

Leave a comment