Wait, What? Letting Students Use Their Notes Does Not Simply Make the Same Test Easier
A learner scores 68% on a closed-book assessment and 84% on an open-book assessment covering the same topic. The obvious conclusion is that the second test was easier. That may be true. It may also be incomplete.
Open-book and closed-book assessments can change preparation, search behaviour, time use, response strategy and the balance between remembering information and locating or applying it. A score produced under one condition is therefore not automatically interchangeable with a score produced under the other.
Quick Answer
Owned Bolt job: calibrate performance claims when access to external resources changes between assessments, so score differences are not automatically interpreted as learning, decline, unfair advantage or equal capability.
The useful question is not “Which format is better?” It is “What did each format ask the learner to do, and what does the resulting performance justify us believing?”
A Resource Changes the Performance Condition
Closed-book performance depends more heavily on what can be retrieved without external support. Open-book performance can shift part of the task toward locating, selecting, interpreting and applying information. Those are legitimate educational demands when they match the assessment purpose.
But “open book” is not one condition. A printed formula sheet, one page of notes, a textbook, unrestricted web access and searchable digital notes create different information environments. If the resource conditions differ, Bolt treats them as part of the measurement record rather than as invisible background.
The Same Topic Can Still Produce Different Constructs
- A closed-book factual recall item may mainly require retrieval.
- An open-book version may require search and recognition rather than unaided recall.
- A well-designed open-book application problem may demand interpretation that no note can supply directly.
- A time-limited open-book test can penalise inefficient searching.
- A poorly designed open-book test can become a race to copy rather than an assessment of understanding.
The format does not tell us the construct by itself. The task design does.
Competing Explanations When the Open-Book Score Is Higher
- The learner knew the material but retrieval difficulty suppressed the closed-book performance.
- The learner used resources effectively to compensate for incomplete internal knowledge.
- The open-book test contained easier items.
- The learner prepared differently because the format was known in advance.
- Search time helped on some items and hurt on others.
- The second assessment benefited from practice or familiarity.
- Ordinary measurement error contributed to the difference.
No single explanation should win merely because one score is larger.
School, Teacher and Student: Three Different Calibration Questions
School
If a school changes an assessment from closed book to open book, it should not silently join the old and new scores into one trend. The school should state what changed, why the format changed, and whether the intended interpretation remains comparable.
Teacher or Coach
The teacher should inspect the response process. Did the learner know which information mattered? Did resources mainly rescue forgotten facts, or did they support stronger reasoning? Did searching consume time that reduced completion? Those observations can change what the score means.
Student
The student should not treat an open-book score as proof that they “know everything” or a closed-book score as proof that they “know nothing.” Each performance answers a different question unless the assessment design establishes otherwise.
The Bolt Open-Book / Closed-Book Calibration Protocol
- Name the assessment purpose. Is the target unaided recall, application with resources, professional information use, synthesis, or something else?
- Declare the resource condition. State exactly what may be used.
- Predict the expected mode effect. Which items should change if resource access matters?
- Check item comparability. The same topic label does not guarantee equal difficulty or equal cognitive demand.
- Inspect more than the total. Compare accuracy, completion, item type, explanation quality, time use and omissions.
- Separate preparation from test-day access. Knowing an exam is open book can change how learners study beforehand.
- Use repeated evidence. One score difference should not permanently redefine the learner.
- Seek a later receipt. Use a delayed or transfer performance under declared conditions to test the interpretation.
Worked Example: The Student Who “Only Knows It With Notes”
A student performs poorly on a closed-book biology quiz but strongly on an open-book case analysis. The teacher initially concludes that the learner has weak memory but good reasoning.
That is a plausible hypothesis, not a finished diagnosis. The teacher looks closer. On the open-book task, the learner spends substantial time locating basic terms but then explains causal links accurately. On a delayed closed-book task involving the same concepts, recall improves but remains less secure than application.
The calibrated conclusion becomes more precise: application appears stronger than unaided retrieval under current conditions. That is better than the vague label “good understanding, bad memory,” and it still leaves room for future evidence to change the model.
What the Score Can and Cannot Support
An open-book score can support claims about performance with the declared resource environment. A closed-book score can support claims about performance without those resources. Neither score alone establishes how the learner will perform under the other condition, and neither is automatically the more authentic measure.
If the educational goal genuinely includes using external information well, forbidding resources may underrepresent the target. If the goal requires fluent internal access to essential knowledge, unrestricted resources may obscure that target. Validity belongs to the interpretation and use.
Teacher–Student Dialogue
Teacher: “Your open-book result was much higher. I do not want to call that improvement yet because the conditions changed.”
Student: “I could look up the facts I forgot.”
Teacher: “Good. Now we separate what the notes supplied from what you could explain and apply. Then we test whether that interpretation survives another performance.”
How Do We Know?
A systematic review in Academic Medicine, Comparing Open-Book and Closed-Book Examinations, screened more than 4,000 records and included 37 studies. It found mixed evidence across preparation, anxiety, performance, psychometrics and testing effects. Examinees often took longer on open-book examinations, and the evidence did not justify declaring one format universally superior.
A 2025 systematic review, Does the Format of an Assessment (Closed Book or Open Book) Affect Learning?, found mixed overall effects on learning, with closed-book assessments showing a lower rate of forgetting in the reviewed literature and students generally preferring open-book formats.
A 2025 meta-analysis, Does Test Format Affect Learning?, synthesised 44 comparisons involving 3,499 participants and found no statistically significant overall effect of open-book versus closed-book format on learning. That result argues against simple claims that one format inherently produces better learning.
The evidence boundary matters. Much open-book research comes from higher education, and effects vary with subject, item design, preparation and resource rules. Schools should not import a universal correction factor into younger learners or different curricula.
For Parents
If a child’s open-book score is much higher, ask what the assessment intended to measure and what the child actually did with the resource. Finding the right formula, recognising the right passage and independently explaining the idea are different pieces of evidence.
The goal is not to decide whether notes are “good” or “bad.” It is to understand which performance condition produced which evidence.
Common Misconceptions
- “Open book is automatically easier.” It can be, but search and application demands can also increase.
- “Closed book measures real knowledge.” It measures performance without external resources under the test’s other conditions; that is not the only legitimate form of knowledge use.
- “A higher open-book score proves poor memory was the only problem.” Other explanations remain possible.
- “Same content means comparable score.” Resource access changes the performance condition.
- “Open-book performance tells us what the learner can do independently later.” A later independent receipt is still needed.
Interface Handoff, MindOS Handoff and Return Receipt
After Bolt calibrates the finding—for example, “application is currently stronger than unaided retrieval, but the open-book score cannot be compared directly with the earlier closed-book score”—the Student/Studying Interface must turn that conclusion into the next executable learner situation: task, goal, criteria, first action, help rule, check and return point.
MindOS may then run the smallest appropriate learner operation selected for that situation. Bolt does not own retrieval practice, note use or any other learning mechanism.
The learner acts, produces a new performance under declared conditions, and the evidence returns to Bolt. That later performance is the receipt used to update the school, teacher and student model.
Bolt Direction Graph
Assessment purpose → resource condition → observed performance → response-process check → competing interpretations → repeated evidence → calibrated conclusion → Student Interface → MindOS if needed → later declared-condition performance → Bolt recalibration.
Useful neighbours: A Supported Answer Is Not the Same Measurement as an Independent Answer, Before You Call It Improvement, Check Whether the Scores Are Comparable, and Student/Studying Interface — Performance Handoff.
