Danielson 1f: Designing Student Assessments Explained

Designing student assessments is Component 1f of the Danielson Framework, inside Domain 1: Planning and Preparation. It covers four elements: congruence with instructional outcomes, criteria and standards, design of formative assessments, and use for planning. In the Danielson Framework, 1f focuses primarily on assessment planning and design; how it is evaluated depends on the locally adopted instrument.
If your district evaluates on the Danielson Framework, 1f is the component about your assessment planning. Not your delivery, not your classroom management. The design work. That catches a lot of good teachers off guard, because most of it happens weeks before anyone walks through the door with a clipboard. Exactly how your system looks at it varies, which is part of what makes the component confusing.
I've sat on both sides of that conversation. What follows is the part of the framework that matters, then eight questions I use to pressure-test an assessment before I give it. The questions came out of my own classroom, and a few of them came out of getting it wrong first.
What Component 1f Actually Asks For
1f sits in Domain 1, which is the planning domain. That placement is the whole point. Danielson separates the design of an assessment from its use during a lesson, and the separation is deliberate.
The component breaks into four elements. Not every state or district uses Danielson, and where it is used, it's often adapted. Evaluation systems get selected or modified at the state, district, charter, or school level, so check the instrument your own system has actually adopted for the wording that applies to you:
- Congruence with instructional outcomes - Does the assessment actually measure what you said students would learn? An outcome about analysis that gets assessed with recall questions fails here, no matter how well-built the questions are.
- Criteria and standards - Are your expectations written down and clear enough that a student knows what proficient looks like before they start?
- Design of formative assessments - Did you plan the checks along the way, or are you improvising them? The framework treats planned formative assessment as design work, not instinct.
- Use for planning - What do you do with the results? Assessment data that doesn't change your next lesson isn't doing its job.
Two of those element names changed in the 2022 revision of the framework. "Design of formative assessments" became "Designing formative assessments," and "Use for planning" became "Analysis and application." If your district is still on the 2011 or 2013 instrument, you'll see the older names. Worth checking which version your evaluator is working from, because the language on your growth plan should match.
One correction I'd offer to anyone reading older material on this topic, including an earlier version of this post: the rating levels in the 2011 and 2013 instruments are Unsatisfactory, Basic, Proficient, and Distinguished. "Exemplary" isn't one of them. Some systems rename the levels in their own adaptations. Use whatever your district uses, but don't invent a level.
1f and 3d Are Not the Same Component
This is the distinction teachers get wrong most often, and it costs them evidence at conference time.
Component 1f is designing assessments. It sits in the planning domain, and it's concerned with the assessment you designed. Component 3d, Using Assessment in Instruction, sits in Domain 3 and is concerned with what you do with assessment while you're teaching: the check for understanding, the feedback, the mid-lesson course correction. What evidence gets collected for either component, and how, depends on the rubric and procedures your system has adopted.
The 2007 edition of the framework drew this line clearly for the first time, and the 2011 revision renamed 1f from "Assessing Student Learning" to "Designing Student Assessments" to make the planning emphasis obvious in the title itself.
Practically, that means a formative assessment can be evidence for both components, but for different reasons. Planning it and building the criteria is 1f. Running it, reading the room, and adjusting on the spot is 3d. When you're assembling artifacts for an evaluation, sort them that way.
Formative and Summative Assessment, Side by Side
Assessment design work generally covers both. They answer different questions, and they're built differently.
| Aspect | Formative | Summative |
|---|---|---|
| Often called | Assessment for learning | Assessment of learning |
| When it happens | During instruction, repeatedly | End of a unit, module, or course |
| What it's for | Adjusting the next lesson | Judging mastery against a standard |
| Usually graded? | Often not, and it doesn't need to be | Yes, against stated criteria |
| Broad conceptual fit | Designed under 1f, used under 3d | Designed under 1f |
Treat that last row as a conceptual mapping, not a scoring structure. How a given implementation sorts evidence between components is set by the rubric in use, not by this table.
The trap is treating formative assessment as a smaller quiz. It isn't a size difference. It's a purpose difference. An exit ticket you never read is not a formative assessment; it's a piece of paper.
How Much Testing Are We Actually Talking About?
Testing has been a four-letter word in staff rooms for two decades, and most of the frustration isn't about assessing students. It's about the ratio of testing to teaching. Between state tests, district benchmarks, unit tests, and everything a teacher writes themselves, the volume adds up fast. A study cited in a 2015 Washington Post piece on standardized testing put the average at roughly 25 hours a year per student, and that figure excludes district benchmarks and teacher-made tests.
The shape of that testing has changed a lot. Teachers who grew up in the 1980s took something like the Iowa Test of Basic Skills, mostly multiple choice, mostly recall-level questions in Bloom's terms. The No Child Left Behind era brought criterion-referenced state tests and a much bigger accountability apparatus, along with a real problem: proficiency in one state didn't mean proficiency in another. That gap is a large part of why a push for common standards happened at all. The Common Core State Standards came out of a joint effort led by the National Governors Association and the Council of Chief State School Officers, and teacher reaction was more mixed than either side of the argument usually admits (Fordham Institute, 2016).
Two things have shifted since. NCLB was replaced by the Every Student Succeeds Act in 2015, which handed a lot of accountability design back to states. And the big multi-state testing consortia have shrunk. Smarter Balanced, built around the Common Core standards, once had around 30 members and now serves roughly a dozen states plus the District of Columbia and a territory. Many states now run their own assessments under their own names.
None of that changes the classroom problem. One test on one day is a small snapshot of a learner. Even when the emphasis is growth, it's still a fraction of the story. Multiple data points, quantitative and qualitative, gathered over time, tell you far more than any single score. That's the case for taking your own assessment design seriously, and it's what the eight questions below are for.
1 - What Criteria Do You Choose When Creating an Assessment?
Strong assessments are partly in the eye of the beholder. What a teacher considers strong, an administrator may not, and a student may see it differently again. The more you can argue why an assessment is strong and back it with evidence and rationale, the stronger it tends to actually be.
This is the hardest thing an educator gets asked to do: build an assessment that tells the teacher, the student, the parent, and the administrator something useful about progress toward a learning objective. Textbook companies, states, districts, and classroom teachers are all trying to do it. Most of what they produce is research-informed, and most of it is still a snapshot at one moment.
What I look at most is formative data over time, because that's where trends and patterns show up. Growth over time is where instructional decisions actually get made. Curriculum, instruction, assessment, and engagement are the four levers a teacher controls, and the data tells you which one to pull.
Start with the end in mind. What you assess drives the lessons leading up to it. Choose learning objectives from your standards, and build the assessment around two components: knowledge of content and reasoning skills. Then bring students into the build where you can.
Some newer online assessments adapt as students answer, getting harder after a correct response and easier after an incorrect one, narrowing toward a proficiency estimate. That's useful, but it isn't a different species from a paper-and-pencil test. The interesting work is outside that box.
RELATED - Master's in Curriculum and Instruction
Here's a worked example. Take a fifth-grade English Language Arts standard:
Common Core Standard: English Language Arts: Reading: Literature: Standard 5, Component 6
CCSS.ELA-LITERACY.RL.5.6 - Describe how a narrator's or speaker's point of view influences how events are described.
A textbook assessment of that standard usually looks like this:
TASK: Read each group of sentences in the paragraph above. Decide if it is written in first person or third person point of view.
That's a legitimate item. It's also narrow. The same standard can be assessed by letting students choose the form their demonstration takes: retell a scene from a second character's perspective, chart how the account changes when the narrator changes, record a short dialogue between two narrators of the same event. Each of those requires the student to do the thing the standard describes. Each gives you more to read than a right-or-wrong bubble.
A caution on how that choice usually gets framed. Offering students different ways to demonstrate learning is well worth doing, and it isn't the same as matching instruction to a student's "learning style." Howard Gardner's theory of multiple intelligences is a claim about the structure of intelligence, and Gardner himself has objected to it being repackaged as learning styles. The evidence that teaching to a preferred style improves achievement is weak. The case for offering multiple demonstration formats is different and stronger: it gives you more evidence about the same standard, and it hands students some ownership of how they show what they know.
Formative assessments feed the summative one. Build the checks that lead up to it, and let students modify or self-assess where it makes sense.
2 - Do Your Assessments Provide for Student Choice?
A freshman history course I took at the University of Montana in 1994 had a midterm and a final, one question each. We were told to bring as many blue books as we could fill in 120 minutes and write everything we remembered. I did not fare well. I wanted choice, and there wasn't any.
Both formative and summative assessments have room for choice built into them. Assessment should drive instruction, and choice gives you better material to drive with.
Mike Anderson's piece on building student choice into formative assessment (Edutopia, March 2017) gives four tips worth keeping:
- Create good choices. They should align with your learning goals, match students' varied interests and abilities, and make sense logistically. Light prep for you, comparable time for students.
- Help students choose well. Guide without directing. Point out which option might suit a student's strengths or match where they are with a skill, then let them decide.
- Practice. The first few times, students will struggle or pick badly. That's fine. Self-assessment and decision-making are skills, and they improve with reps.
- Don't force it. If one assessment really is the best check on a piece of learning, don't offer a choice. Manufactured choice feels fake and doesn't work.
3 - Are the Assessments You Create Applicable to the Real World?
When a student asks "when will I ever use this?", you should have an answer ready. Without a rationale, students have nothing to connect the learning to, and it doesn't stick. Plenty of required assessments make that hard, which is a real constraint rather than an excuse.
What helps: give the rationale up front, offer choice in how students demonstrate learning, and do the research to find the connection when it isn't obvious. That last one takes time. It's also the difference between a task students complete and a task students care about.
4 - Do Your Assessments Adapt to Individual Students as Needs Arise?
Many districts use common assessments that give teachers summative data on students and on instructional effectiveness. Layered on top of that are individual requirements. When a student's IEP or Section 504 plan specifies assessment accommodations, those accommodations must be provided as required by the plan.
Beyond the required accommodations, there's room to individualize through customized learning plans or tiered supports such as Response to Intervention. Sometimes the adaptation is as simple as moving a student from a written assessment to a verbal one. The goal stays the same: score the student against proficiency for the learning objective, or against their own prior data for growth.
Deciding what to adapt is both art and science. You have to identify what the student is actually struggling with, then match an accommodation to it. Here's a list of testing accommodations worth having on hand:

5 - Are Students Involved in Designing Assessments and Rubrics for Their Own Work?
How much control a teacher shares over assessment design comes up often in discussions of the Danielson evaluation model. Involving students in building assessments can be evidence of high-level practice, though the exact criteria for the top of any scale depend on the framework version and the local adaptation. A teacher working that way gives up some control by teaching students how to build a strong assessment, formative or summative. You specify the learning objectives the assessment has to meet, then work with students to design something that meets them.
Handing over that ownership does two things. Students learn more from building the measure than from taking it, and they know exactly what a strong response looks like because they helped define it. Agency is the term for it, and it's not a soft outcome.
Here's an assessment built by students and teacher together in a fifth-grade science and social studies project:

6 - Do Students Complete Self-Assessments on a Rubric?
Rubrics and checklists let students score themselves against stated expectations. A simple 0-5 scale gives the reflection a number to hang on. For a 45-minute work period, that might look like:
- 5 Points = 100% on task, all required work completed, supported another student's learning
- 4 Points = 100% on task, all required work completed, 0 teacher reminders
- 3 Points = 80% on task, all required work completed, and/or teacher supported behavior once
- 2 Points = 80% on task, almost all work completed, and/or teacher supported behavior 1-2 times
- 1 Point = under 80% on task, most work incomplete, and/or teacher supported behavior 3+ times
- 0 Points = no evidence of participation, no work completed, a behavior disruption
A content-based version works the same way. Here's a 3-point rubric tied to a fifth grade social studies lesson:
- 3 Points = identifies 3 causes of the Revolutionary War and explains each clearly with detail
- 2 Points = identifies at least 1 cause and explains it clearly with detail
- 1 Point = identifies 1-3 causes but cannot explain them clearly
- 0 Points = cannot identify a cause or explain one

Asking students to reflect on a defined stretch of work is powerful, and setting those expectations before instruction starts is also strong classroom management. When a student scores themselves low, that's an opening for a one-on-one conference and a goal. The record you build is useful for the student, for you, and for the conversation with a parent about patterns rather than incidents.
7 - Do Your Formative Assessments Drive Small Group Instruction?
Small group instruction is well supported in the research. John Hattie's Visible Learning synthesis puts small group learning at an effect size of roughly 0.46. Worth knowing what that number means before you quote it: Hattie set 0.40 as the hinge point, the average across everything he studied, so small group instruction sits a little above average rather than at the top of the list. It works. It isn't magic, and how you group matters more than whether you group.
Grouping well takes data. Find the students who share a specific difficulty, and teach into it. You can group by skill, concept, readiness, or modality, and the right choice depends on what the data is telling you that week.
Set goals for the group and for each student in it. Everyone should be able to say what they're working on and how they're doing against it. Time in small group is short, so what happens there needs to count. A simple 1-4 scoring rubric during group time gives you data to set the next round of goals from.

8 - Do Your Summative Assessments Assess Reasoning as Well as Factual Knowledge?
State and district assessments should measure both. When you're writing your own, Bloom's Taxonomy is a useful check on what you're actually asking. Multiple choice and true/false questions grade fast and tell you less. An open-ended question that asks a student to evaluate or analyze takes longer to score and tells you more.
Scoring reasoning is harder than scoring recall, and the judgment gets grey. That's exactly why the criteria have to be written down before students start. Define what a proficient response looks like, then hold to it.
The Christopher Columbus rubric below asks students to build a product covering ten events from his voyage to the New World. They have to justify why those ten and not others, which means defending what they left out as well as what they included. Getting students to talk through that reasoning gives you the strongest read on whether they hit the objective. They also need the factual knowledge to make the choices at all, so both halves of the component are covered in one task.

More From the Rock My Evaluation Series
- Demonstrating Knowledge of Content and Pedagogy
- Demonstrating Knowledge of Your Students
- Setting Instructional Outcomes
- Demonstrating Knowledge of Resources
- Designing Coherent Instruction
- Designing Student Assessments (you are here)
Frequently Asked Questions
What is component 1f in the Danielson Framework?
1f is Designing Student Assessments, one of six components in Domain 1: Planning and Preparation. It covers four elements: congruence with instructional outcomes, criteria and standards, design of formative assessments, and use for planning. It focuses primarily on assessment planning and design. How that gets evaluated, and what evidence an evaluator collects, depends on the instrument your state, district, or school has adopted.
What's the difference between Danielson 1f and 3d?
1f is about designing assessments before instruction. 3d, Using Assessment in Instruction, is about what you do with assessment while teaching: checking for understanding, giving feedback, and adjusting mid-lesson. The same formative assessment can supply evidence for both, but the evidence is different. Planning and building it is 1f. Running it and responding to what you see is 3d. That's the conceptual split. How evidence is actually gathered and scored is set by your locally adopted rubric.
What evidence should I bring for a 1f conversation?
Bring the plan, not just the product. Lesson and unit plans showing where formative checks sit, the rubric or criteria you gave students before they started, an example of an assessment adapted for a specific student's needs, and something showing how a set of results changed what you taught next. That last one covers the use-for-planning element, and it's the one teachers most often leave out. Check what your own system asks for, since evidence requirements vary between implementations.
Do formative assessments have to be graded?
No. Formative assessment is designed to inform your next instructional move, and a grade isn't required for that. Plenty of the most useful checks are ungraded and take two minutes. What matters is that you planned it, you read the results, and something changed as a result.
- 1f is a planning component - It focuses on the assessment you designed rather than the lesson delivered. How that's evaluated depends on your locally adopted instrument.
- Know your four elements - Congruence with instructional outcomes, criteria and standards, design of formative assessments, and use for planning. The 2022 revision renamed the last two.
- Don't confuse 1f with 3d - Designing the assessment is 1f. Using it during instruction is 3d. That distinction holds conceptually, but confirm how your own system treats it.
- Write the criteria before students start - Students should know what proficient looks like in advance, and shared rubrics are some of the strongest 1f evidence you can produce.
- Close the loop - Assessment data that doesn't change your next lesson misses the use-for-planning element entirely.
If this is the part of the job you want to get better at, curriculum and assessment are a specialization in their own right, and it's what a lot of instructional coaches and curriculum leads studied before they moved into those roles. The sponsored programs listed on this page include online and hybrid options built for teachers who are staying in the classroom while they study.
- Learning How to Say No and Set Boundaries with Parents - November 21, 2022
- How to Implement a Check-In, Check-Out (CICO) Behavior System - September 26, 2022
- Teacher Core Values: 7 Strategies for Living Your Code - August 15, 2022





