Trellis is now available for the 2025–26 school year. Request your free 15-minute demo →
STUDY 02 · FORMATIVE ASSESSMENT · DYLAN WILIAM

Assessment that works. Assessment that's actually done. They are not the same problem.

Dylan Wiliam's review of 250 studies established that formative assessment — checking for understanding during instruction to adjust what happens next — is one of the most powerful tools available to teachers. It can double the speed of learning. His research also found, in the same breath, that day-to-day classroom assessment was "relatively rare." This page is about the gap between those two sentences — and how Trellis closes it.

Read the full study
Key Takeaways
The Problem
Although strong evidence shows formative assessment is known to improve learning, teachers often struggle to implement it in everyday classroom practice.
The Findings
Formative assessment was rarely embedded in everyday classroom practice, even though feedback is one of the most powerful drivers of student achievement.
The Trellis Solution
Trellis makes formative assessment effortless. 5 seconds, not 30 minutes.

The Wiliam Framework

Three questions every lesson must answer.

Wiliam identifies three "crucial processes" that effective formative assessment must address. Trellis is designed to answer all three — in real time, during instruction.

01
Where are learners in their learning?
"The coherence of these ideas can be seen more clearly by considering three crucial processes in learning." — Wiliam, 2005, p. 13
Trellis answers this
The 5-color assessment scale captures current understanding for every student in real time — not from memory at the end of the day, but during the lesson itself. The seating chart becomes a live map of where the class is right now.
02
Where are they going?
Formative assessment must connect current performance to clear learning targets, not just describe what students got wrong.
Trellis answers this
Intervention trellises make success criteria visible and explicit. When a student is marked Orange or Red, the trellis reveals the scaffolds between where they are and where they're headed — sentence frames, visual supports, peer checks — not vague remediation.
03
How do they get there?
Assessment without actionable next steps produces measurement without learning. The third question is the one most often left unanswered.
Trellis answers this
The intervention panel provides the specific "how" — documented in real time. Every documented intervention is a data point about what strategies actually worked for a specific student on a specific concept. Over time, this becomes a personalized learning record.
01

Assessment is the bridge between teaching and learning.

One of Wiliam's most fundamental observations sounds obvious and yet is consistently violated in practice: what teachers teach and what students learn are not the same thing. A teacher can deliver an excellent lesson on equivalent fractions. Students can sit attentively, take notes, complete the practice problems, and leave with a durable misconception that remained invisible to the teacher because no one looked for it.

This is not a failure of teaching. It is the natural result of learning: people construct understanding through the lens of what they already believe, which means students actively and unconsciously distort new information to fit existing mental models. The only way to catch this before it calcifies is to look — actively, systematically, during instruction — at what students are actually thinking, not what they appear to be doing.

"We must acknowledge that what students learn is not necessarily what the teacher intended, and it is essential that teachers explore students' thinking before assuming that students have 'understood' something. In this sense, assessment is the bridge between teaching and learning."
Wiliam, 2005 · Keeping Learning on Track · p. 3

The phrase "bridge between teaching and learning" is precise and important. Teaching without formative assessment is a monologue — the teacher produces content, releases it into the room, and hopes it lands. Assessment transforms teaching into dialogue: the teacher offers, the student responds (verbally or through behavior), the teacher adjusts. The loop closes. Learning becomes visible.

T How Trellis addresses this
Trellis makes the bridge literal and visible. When a teacher clicks a student's seat and assigns a color level, they are not grading — they are documenting what they observed about that student's thinking in that moment. The 5-color scale is designed to capture nuance: a student marked Yellow is approaching understanding, not failing. The seating chart becomes a real-time map of where learning is and isn't happening across every seat in the room.
02

250 studies. One consistent finding: it works.

The 1998 meta-analysis by Paul Black and Dylan Wiliam — published as "Inside the Black Box" in Phi Delta Kappan and the accompanying full report in Assessment in Education — reviewed 250 studies conducted between 1987 and 1998 from research groups around the world. The size of this evidence base was unusual; the clarity of the conclusion was striking.

When teachers shifted their focus from assessment of learning (grading outcomes) to assessment for learning (adjusting instruction based on evidence of understanding), student achievement improved substantially and consistently. Effect sizes in the individual studies ranged from 0.4 to 0.7 — placing formative assessment among the most impactful educational interventions ever documented at scale.

"In reviewing 250 studies from around the world, published between 1987 and 1998, we found that a focus by teachers on assessment for learning, as opposed to assessment of learning, produced a substantial increase in students' achievement."
Wiliam, 2005 · p. 1 · Referencing Black & Wiliam (1998) — "Inside the Black Box"

The practical implication of these effect sizes is significant. Students whose teachers used embedded formative assessment consistently learned the equivalent of several additional months of schooling per year compared to students in classrooms where assessment remained primarily summative. In terms of educational intervention ROI, few strategies come close.

📊
Key Finding · Black & Wiliam, 1998 · 250-Study Meta-Analysis
Assessment for learning — formative assessment that informs instructional decisions during teaching — produces effect sizes of 0.4–0.7, placing it among the most impactful educational interventions documented. The question isn't whether it works. The question is why it isn't happening more often.
03

It works — and it's rarely done. These facts coexist.

Wiliam's 2005 paper confronts an uncomfortable tension head-on. Embedded in the same paragraph that announces the dramatic achievement gains of formative assessment is a quiet but devastating observation: despite 20 years of evidence for its effectiveness, day-to-day formative assessment remained relatively rare in actual classrooms. The research community knew it worked. Teachers largely weren't doing it. And both things were true simultaneously.

"Since the studies also revealed that day-to-day classroom assessment was relatively rare, we felt that considerable improvements would result from supporting teachers in developing this aspect of their practice."
Wiliam, 2005 · p. 1 · The motivating observation for the entire formative assessment reform agenda

Why rare? Not because teachers lack commitment or knowledge. The research on this point is clear: formative assessment fails to become daily practice for the same reason most high-effort behaviors fail to become habits — the ratio of immediate cost to immediate benefit is unfavorable. The benefit (better student learning) accrues slowly and diffusely. The cost (additional documentation and data management during and after instruction) is immediate, concrete, and compounding.

A teacher who genuinely attempts to implement Wiliam's formative assessment principles without tools designed to make it easy will spend between 20 and 45 additional minutes per class per day. For a teacher with five classes, that's up to four hours of additional daily load — before grading, planning, communications, and the hundred other things that fill a teacher's evening. The behavior is rational to abandon, regardless of its educational value.

T How Trellis addresses this
Trellis is the tool designed to make the cost structure of formative assessment sustainable. Assessment takes 2 seconds during instruction — a color tap on a seating chart. Documentation happens in real time, not as a post-instruction task. At the end of a lesson, documentation is complete. At the end of the day, teachers leave with timestamped intervention data they never had to create separately. The behavior becomes easy enough to actually do.
04

Feedback is not formative assessment. The distinction matters.

One of Wiliam's sharpest conceptual contributions is distinguishing between feedback and formative assessment — two terms routinely used as synonyms in educational discourse, and which are in fact meaningfully different. Feedback is information provided to a learner about their performance. Formative assessment is feedback that the learner can use to improve their performance. The difference is in whether the information is actionable.

"Feedback is formative only if the information fed back to the learner is used by the learner in improving performance. If the information fed back to the learner is intended to be helpful, but cannot be used by the learner in improving her own performance it is not formative. It is rather like telling an unsuccessful comedian to 'be funnier.'"
Wiliam, 2005 · p. 8 · The comedian analogy

The comedian analogy is memorable precisely because it captures the absurdity of most academic feedback. Telling a student to "work harder," "be more systematic," or "show your thinking" is structurally identical to telling a comedian to be funnier. The advice is accurate. It is not actionable. And because it is not actionable, it is not formative — it is merely feedback, with none of the learning benefits that formative assessment research documents.

For feedback to become formative, it must contain what Wiliam calls a "recipe for future action" — specific enough that a student can do something different in their next attempt. Abstract directives fail this test universally. Concrete, scaffolded, strategy-specific responses pass it.

T How Trellis addresses this
Trellis structurally prevents vague feedback. When a teacher marks a student Orange or Red, the intervention trellis surfaces specific, named strategies — Sentence Frame, Visual Model, Partner Check, Extended Wait Time — that constitute actionable recipes. The teacher selects and documents a specific intervention. Students and teachers both know exactly what support was provided. Future lessons can build on or adjust based on what worked.
05

Grades attached to feedback cancel the feedback.

One of the most counterintuitive and consequential findings in Wiliam's research comes from Ruth Butler's experimental studies on feedback and motivation. In a carefully controlled experiment, Butler gave students three types of feedback on their work: grades only, comments only, or grades alongside comments. Most educators would predict that grades + comments would produce the best results — more information, better outcomes. The data contradicts this assumption entirely.

"Far from producing the best effects of both kinds of feedback, giving marks alongside the comments completely washed out the beneficial effects of the comments. The use of both marks and comments is probably the most widespread form of feedback used in the Anglophone world — and yet it is no more effective than marks alone. If you write careful diagnostic comments on a student's work, and then put a score or grade on it, you are wasting your time."
Wiliam, 2005 · p. 6 · Referencing Ruth Butler's experimental studies on feedback types

The mechanism is psychological and predictable. When grades are present, students' attention goes immediately and almost entirely to the grade — not to the diagnostic comments. If the grade is high, students feel good and stop engaging with feedback. If the grade is low, students feel bad and interpret the feedback through a defensive lens. In both cases, the information in the comments — which is the only part with genuine learning value — is effectively invisible.

Butler's second study extended this finding to praise: students given grades and praise made no more progress than students given no feedback at all. The only measurable effect of grades and praise combined was an increase in ego-involvement without any corresponding increase in achievement. The feedback that costs the most teacher time and emotional labor — careful diagnostic marking — is structurally undermined by the grade that accompanies it.

T How Trellis addresses this
Trellis has no grades. The 5-color assessment scale is descriptive, not evaluative — it describes where a student currently is in relation to the learning target, not how good they are as a student. Orange means "needs a scaffold right now." It is not a judgment. This structure preserves the learning-focused orientation that Wiliam's research identifies as essential: task-involvement over ego-involvement, growth over performance.
06

When students assess themselves, learning doubles.

The formative assessment research points to a finding that goes beyond teacher practice into student agency: when students are given the tools and habits of self-assessment — evaluating their own work against clear criteria — their learning gains are substantially larger than when teachers do all the assessing. Wiliam cites Fontana and Fernandez's study of Portuguese primary school students as a particularly striking example.

Students taught with self-assessment improved nearly twice as fast.
Control group improved 7.8 marks over the study period. Students whose teachers developed self-assessment practices improved 15 marks in the same period — almost exactly double. The additional variable was not more instruction time. It was structured self-reflection on their own learning.
Fontana & Fernandez study · Referenced in Wiliam, 2005, pp. 12–13

Wiliam adds an important nuance: self-assessment doesn't need to be accurate to be beneficial. The heated debate over whether students can objectively evaluate their own performance, he argues, misses the point. What matters is whether the act of self-assessment — reflecting on one's own learning against stated criteria — improves learning. And it does, reliably, regardless of whether the student's self-assessment is precisely calibrated.

"What really matters is whether self-assessment can enhance learning, and in this regard, accuracy is a secondary concern. The metacognitive engagement itself is the mechanism."
Wiliam, 2005 · p. 12 · On the self-assessment accuracy debate
T How Trellis addresses this
Trellis creates the data infrastructure for meaningful student self-assessment. When teachers share seating chart data with students, students can see their own color patterns over time: "I was Red on fractions three weeks ago, Orange last week, Yellow this week." The color progression is growth made visible — concrete evidence that effort and scaffolded support produce improvement. This is growth mindset grounded in real data, not exhortation.
07

Great teaching happens in "moments of contingency."

Wiliam's theoretical framework for formative assessment centers on a concept he calls the "regulation of learning" — the idea that teaching is not the delivery of content but the creation of conditions in which students learn. Within this framework, the most important moments in a lesson are what he terms "moments of contingency": decision points at which the lesson can proceed in different directions depending on what students show they understand.

"These 'moments of contingency' — points in the instructional sequence when the instruction can proceed in different directions according to the responses of the student — are at the heart of the regulation of learning."
Wiliam, 2005 · p. 15 · Regulation of learning theory

This framing has profound practical implications. It means that the quality of a lesson is not determined by how well the teacher delivers prepared material — it is determined by how effectively the teacher reads student responses and adjusts in real time. A teacher who delivers a polished lesson with no sensitivity to what students are actually understanding may be demonstrating excellent performance while producing mediocre learning outcomes.

Wiliam also highlights that the quality of teacher questioning is central to creating productive moments of contingency. He notes that extending wait time after a student answers — from under one second to three seconds — produces measurable increases in learning, because it gives students time to think rather than perform. He further argues that statements often produce richer student discourse than questions, because they require evaluation rather than simple recall.

Key Finding · Rowe (1986) via Wiliam, 2005 · p. 5 — Wait Time Research
Increasing the time between a student's answer and the teacher's evaluation from the average of less than one second to just three seconds produces measurable increases in learning. Extended wait time can be documented in Trellis as a specific, named intervention — making an invisible pedagogical move visible and trackable for the first time.
T How Trellis addresses this
Trellis is built for moments of contingency. Every student interaction during a lesson is a potential contingency point: the teacher notices a student struggling (sees Orange emerging on the chart), selects the appropriate scaffold from the intervention trellis, and adjusts the instructional path for that student — all in under 5 seconds, without breaking instructional flow. The lesson doesn't pause for documentation. Documentation happens because of instruction, not instead of it.
08

Formative assessment must be embedded, not added on.

A critical thread running through Wiliam's practical guidance is the distinction between embedding formative assessment in daily instruction and treating it as a parallel activity. The failure mode he observes most often is the "add-on" approach: teachers who sincerely want to implement formative assessment principles and attempt to do so by adding new steps before, during, or after their existing practice. This approach adds burden without integrating benefit, and it almost always collapses under instructional load.

"To be effective, these strategies must be embedded into the day-to-day life of the classroom, and must be integrated into whatever curriculum scheme is being used. That is why there can be no recipe that will work for everyone."
Wiliam, 2005 · p. 16 · On implementation requirements for formative assessment

Wiliam is also careful to note that there is no one right way to implement formative assessment. The feedback routines in each classroom must be adapted to the teacher's style, the subject matter, the age of students, and the culture of the school. What cannot be adapted away is the core requirement: that the assessment happens during instruction, that it informs decisions about instruction, and that the burden of doing it is low enough that teachers actually sustain it.

Finally, Wiliam describes the ultimate outcome of effective formative assessment in terms that are both aspirational and concrete: students who think more often than they try to remember, who believe that working hard makes them more capable, who understand what they are working toward, and who know how they are progressing. These outcomes are not the product of any single lesson or intervention — they are the product of sustained daily practice over time.

"Students will be thinking more often than they are trying to remember something, they will believe that by working hard, they get cleverer, they will understand what they are working towards, and will know how they are progressing."
Wiliam, 2005 · p. 16 · The outcome of effective formative assessment
T How Trellis addresses this
Trellis is not an add-on. It uses the seating chart — a visual representation of the classroom that teachers already maintain — as the primary interface for formative assessment. Teachers don't open a separate application, fill out a separate form, or navigate to a separate system. They work in the tool that reflects their actual classroom, during the lesson, in the natural flow of moving around the room and engaging students. The assessment is embedded because the tool is embedded.
09

Wiliam's findings and Trellis's responses — side by side.

Ten of Wiliam's core findings from the 1998 meta-analysis and 2005 synthesis, mapped directly to the design decisions that make Trellis the implementation infrastructure formative assessment has always needed.

Wiliam (1998 / 2005) — Research to Practice

Every finding. Every Trellis response.

Wiliam Finding What this means for MTSS Trellis Solution
Assessment is the bridge between teaching and learning Teachers must actively look for evidence of understanding during instruction — not assume it from compliance 5-color scale documents observed understanding in real time, during instruction
Formative assessment produces substantial achievement gains (0.4–0.7 effect size) Schools that implement consistent formative practice see students learn the equivalent of additional months of schooling per year Every Trellis interaction is a formative assessment event — assessment for learning, not of learning
Day-to-day classroom assessment is relatively rare The burden of documentation prevents the consistent practice that produces gains 2-second color tap during instruction — documentation complete when lesson ends
Feedback is formative only if the learner can act on it "Try harder" and "good job" are feedback — not formative assessment. Teachers need actionable response options, not vague guidance Intervention trellis provides specific, named strategies that teachers and students can act on immediately
Grades attached to comments cancel the comments entirely The most common form of feedback in English-speaking classrooms (grade + comment) has the same effect as grade alone — zero No grades in Trellis — 5-color scale is descriptive (where are you?) not evaluative (how good are you?)
Scaffolded minimal intervention beats complete solutions Students learn more from just enough help than from being given answers — but teachers need scaffolded options available in the moment Intervention trellis provides scaffolded options (sentence frames, visual models, wait time) not answer keys
Self-assessment nearly doubles achievement gains (7.8 → 15 marks) When students can see their own learning data and reflect on it, the learning effect of formative assessment doubles Teachers can share progress data with students — color patterns over time make growth visible and concrete
"Moments of contingency" are where great teaching happens The most important teaching decisions happen in real time based on student responses — not in lesson plans made the night before Trellis is designed for use during instruction — color assessment and intervention selection happen at the moment of contingency
Formative assessment must be embedded, not added Any system that requires additional parallel effort will collapse under instructional load — regardless of its educational value Seating chart interface integrates into existing classroom tools — no parallel system, no additional workflow
The classroom is a "black box" — we see outcomes but not instruction Districts have rich data on test scores and almost none on what happens during daily instruction — the thing that produces scores Trellis opens the black box — admin dashboards show what interventions are being used during Tier 1 instruction in real time

About Dylan Wiliam

Dylan Wiliam — Emeritus Professor, UCL Institute of Education, London. Founding co-editor of Assessment in Education. Author of Embedded Formative Assessment. Co-author with Paul Black of the landmark 1998 "Inside the Black Box" meta-analysis.

Black, P. J., & Wiliam, D. (1998). Inside the black box: Raising standards through classroom assessment. Phi Delta Kappan, 80(2), 139–148. · Wiliam, D. (2005). Keeping learning on track. 20th biennial meeting, Australian Association of Mathematics Teachers, Sydney.
Continue Reading

Wiliam tells us what works.
Fixsen explains why it doesn't get implemented.
Zhou shows how it becomes self-sustaining.

Study 01

Implementation Science

Dean Fixsen

Fixsen's synthesis of 250+ studies found that training alone produces 0–20% behavior change. The missing piece is organizational infrastructure — systems interventions, decision support data, and feedback mechanisms that make trained practices automatic in daily workflow.

Read study
Study 03

Collective Teacher Efficacy

YaRu Zhou

Zhou synthesizes the research showing collective teacher efficacy — the #1 influence on student achievement at d=1.57 — is built primarily through mastery experiences. Trellis creates the shared evidence infrastructure that builds this belief school-wide.

Read study