Assessment is the bridge between teaching and learning.
One of Wiliam's most fundamental observations sounds obvious and yet is consistently violated in practice: what teachers teach and what students learn are not the same thing. A teacher can deliver an excellent lesson on equivalent fractions. Students can sit attentively, take notes, complete the practice problems, and leave with a durable misconception that remained invisible to the teacher because no one looked for it.
This is not a failure of teaching. It is the natural result of learning: people construct understanding through the lens of what they already believe, which means students actively and unconsciously distort new information to fit existing mental models. The only way to catch this before it calcifies is to look — actively, systematically, during instruction — at what students are actually thinking, not what they appear to be doing.
The phrase "bridge between teaching and learning" is precise and important. Teaching without formative assessment is a monologue — the teacher produces content, releases it into the room, and hopes it lands. Assessment transforms teaching into dialogue: the teacher offers, the student responds (verbally or through behavior), the teacher adjusts. The loop closes. Learning becomes visible.
250 studies. One consistent finding: it works.
The 1998 meta-analysis by Paul Black and Dylan Wiliam — published as "Inside the Black Box" in Phi Delta Kappan and the accompanying full report in Assessment in Education — reviewed 250 studies conducted between 1987 and 1998 from research groups around the world. The size of this evidence base was unusual; the clarity of the conclusion was striking.
When teachers shifted their focus from assessment of learning (grading outcomes) to assessment for learning (adjusting instruction based on evidence of understanding), student achievement improved substantially and consistently. Effect sizes in the individual studies ranged from 0.4 to 0.7 — placing formative assessment among the most impactful educational interventions ever documented at scale.
The practical implication of these effect sizes is significant. Students whose teachers used embedded formative assessment consistently learned the equivalent of several additional months of schooling per year compared to students in classrooms where assessment remained primarily summative. In terms of educational intervention ROI, few strategies come close.
It works — and it's rarely done. These facts coexist.
Wiliam's 2005 paper confronts an uncomfortable tension head-on. Embedded in the same paragraph that announces the dramatic achievement gains of formative assessment is a quiet but devastating observation: despite 20 years of evidence for its effectiveness, day-to-day formative assessment remained relatively rare in actual classrooms. The research community knew it worked. Teachers largely weren't doing it. And both things were true simultaneously.
Why rare? Not because teachers lack commitment or knowledge. The research on this point is clear: formative assessment fails to become daily practice for the same reason most high-effort behaviors fail to become habits — the ratio of immediate cost to immediate benefit is unfavorable. The benefit (better student learning) accrues slowly and diffusely. The cost (additional documentation and data management during and after instruction) is immediate, concrete, and compounding.
A teacher who genuinely attempts to implement Wiliam's formative assessment principles without tools designed to make it easy will spend between 20 and 45 additional minutes per class per day. For a teacher with five classes, that's up to four hours of additional daily load — before grading, planning, communications, and the hundred other things that fill a teacher's evening. The behavior is rational to abandon, regardless of its educational value.
Feedback is not formative assessment. The distinction matters.
One of Wiliam's sharpest conceptual contributions is distinguishing between feedback and formative assessment — two terms routinely used as synonyms in educational discourse, and which are in fact meaningfully different. Feedback is information provided to a learner about their performance. Formative assessment is feedback that the learner can use to improve their performance. The difference is in whether the information is actionable.
The comedian analogy is memorable precisely because it captures the absurdity of most academic feedback. Telling a student to "work harder," "be more systematic," or "show your thinking" is structurally identical to telling a comedian to be funnier. The advice is accurate. It is not actionable. And because it is not actionable, it is not formative — it is merely feedback, with none of the learning benefits that formative assessment research documents.
For feedback to become formative, it must contain what Wiliam calls a "recipe for future action" — specific enough that a student can do something different in their next attempt. Abstract directives fail this test universally. Concrete, scaffolded, strategy-specific responses pass it.
Grades attached to feedback cancel the feedback.
One of the most counterintuitive and consequential findings in Wiliam's research comes from Ruth Butler's experimental studies on feedback and motivation. In a carefully controlled experiment, Butler gave students three types of feedback on their work: grades only, comments only, or grades alongside comments. Most educators would predict that grades + comments would produce the best results — more information, better outcomes. The data contradicts this assumption entirely.
The mechanism is psychological and predictable. When grades are present, students' attention goes immediately and almost entirely to the grade — not to the diagnostic comments. If the grade is high, students feel good and stop engaging with feedback. If the grade is low, students feel bad and interpret the feedback through a defensive lens. In both cases, the information in the comments — which is the only part with genuine learning value — is effectively invisible.
Butler's second study extended this finding to praise: students given grades and praise made no more progress than students given no feedback at all. The only measurable effect of grades and praise combined was an increase in ego-involvement without any corresponding increase in achievement. The feedback that costs the most teacher time and emotional labor — careful diagnostic marking — is structurally undermined by the grade that accompanies it.
When students assess themselves, learning doubles.
The formative assessment research points to a finding that goes beyond teacher practice into student agency: when students are given the tools and habits of self-assessment — evaluating their own work against clear criteria — their learning gains are substantially larger than when teachers do all the assessing. Wiliam cites Fontana and Fernandez's study of Portuguese primary school students as a particularly striking example.
Wiliam adds an important nuance: self-assessment doesn't need to be accurate to be beneficial. The heated debate over whether students can objectively evaluate their own performance, he argues, misses the point. What matters is whether the act of self-assessment — reflecting on one's own learning against stated criteria — improves learning. And it does, reliably, regardless of whether the student's self-assessment is precisely calibrated.
Great teaching happens in "moments of contingency."
Wiliam's theoretical framework for formative assessment centers on a concept he calls the "regulation of learning" — the idea that teaching is not the delivery of content but the creation of conditions in which students learn. Within this framework, the most important moments in a lesson are what he terms "moments of contingency": decision points at which the lesson can proceed in different directions depending on what students show they understand.
This framing has profound practical implications. It means that the quality of a lesson is not determined by how well the teacher delivers prepared material — it is determined by how effectively the teacher reads student responses and adjusts in real time. A teacher who delivers a polished lesson with no sensitivity to what students are actually understanding may be demonstrating excellent performance while producing mediocre learning outcomes.
Wiliam also highlights that the quality of teacher questioning is central to creating productive moments of contingency. He notes that extending wait time after a student answers — from under one second to three seconds — produces measurable increases in learning, because it gives students time to think rather than perform. He further argues that statements often produce richer student discourse than questions, because they require evaluation rather than simple recall.
Formative assessment must be embedded, not added on.
A critical thread running through Wiliam's practical guidance is the distinction between embedding formative assessment in daily instruction and treating it as a parallel activity. The failure mode he observes most often is the "add-on" approach: teachers who sincerely want to implement formative assessment principles and attempt to do so by adding new steps before, during, or after their existing practice. This approach adds burden without integrating benefit, and it almost always collapses under instructional load.
Wiliam is also careful to note that there is no one right way to implement formative assessment. The feedback routines in each classroom must be adapted to the teacher's style, the subject matter, the age of students, and the culture of the school. What cannot be adapted away is the core requirement: that the assessment happens during instruction, that it informs decisions about instruction, and that the burden of doing it is low enough that teachers actually sustain it.
Finally, Wiliam describes the ultimate outcome of effective formative assessment in terms that are both aspirational and concrete: students who think more often than they try to remember, who believe that working hard makes them more capable, who understand what they are working toward, and who know how they are progressing. These outcomes are not the product of any single lesson or intervention — they are the product of sustained daily practice over time.
Wiliam's findings and Trellis's responses — side by side.
Ten of Wiliam's core findings from the 1998 meta-analysis and 2005 synthesis, mapped directly to the design decisions that make Trellis the implementation infrastructure formative assessment has always needed.