Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — A Peer Grade Is Evidence, Not a Substitute Teacher

Wait, What? Several Students Can Agree on a Grade and Still Be Agreeing for the Wrong Reason

Peer assessment can be powerful. Students can notice strengths and weaknesses that a teacher misses, receive feedback faster, learn assessment criteria more deeply, and improve by judging other people’s work.

But a peer grade is not automatically equivalent to a calibrated teacher judgement simply because several students agree.

Peers may share the same misunderstanding, use a different standard, avoid criticising friends, overvalue presentation, or become more accurate only after training and repeated practice. The peer score is evidence. Its strength depends on how the peer-assessment system is designed.

Quick Answer

Owned Bolt job: calibrate when student peer assessment can support performance judgement, feedback and grading—and when it should remain a supplementary evidence channel rather than a final verdict.

Research supports peer assessment as a learning and formative-assessment practice. Peer scores can also show useful agreement with teacher or expert scores in some settings. But reliability and validity vary with task complexity, training, criteria, privacy, number of raters, workload and local social dynamics. The appropriate question is not “Can students mark?” It is “Under these conditions, what does this peer evidence justify?”

Peer Assessment Has Two Jobs—and They Should Be Separated

Job 1: Learning through assessment

Students inspect criteria, compare performances, generate feedback and see alternative approaches. This can improve their own learning even if the peer score is never used as an official grade.

Job 2: Measuring another student’s performance

The peer judgement becomes evidence about the quality of another learner’s work. This requires stronger calibration if the score will affect grades, selection or other consequential decisions.

These jobs can coexist, but success at Job 1 does not prove adequacy at Job 2. A peer-assessment activity can be educationally excellent even if the resulting numerical marks are not stable enough for high-stakes grading.

Current Evidence Supports Peer Assessment—but Not Blindly

A systematic review covering 449 peer-assessment studies found enormous variation in designs, purposes, objects and outcomes. The review notes that peer-assessment reliability and validity have been established in general, but also warns that each particular design still needs its own quality evidence rather than borrowing validity from the research field as a whole.

A 2024 systematic review focused on interpersonal processes found that training generally helps reduce negative effects, while privacy and format can have mixed consequences depending on the context. A 2023 study showed why “just add more peer raters” is also too simple: task complexity and reviewer workload can influence reliability and validity in different ways.

Peer assessment therefore behaves like every serious measurement system. Design matters.

Agreement Is Not Enough

Suppose four students all give a presentation 8/10. High agreement sounds reassuring. But several possibilities remain:

  • the peers applied the rubric accurately;
  • the peers shared a common but incorrect interpretation of the rubric;
  • the task was simple enough that grading was easy;
  • the group avoided giving low scores because of social discomfort;
  • the peers overvalued visible fluency and undervalued reasoning;
  • everyone used the same informal class norm rather than the declared standard.

Reliability asks whether judgements are consistent. Validity asks whether the judgement supports the intended interpretation. Agreement can help reliability while leaving validity unresolved.

Peer Judgement Can Sometimes See What Teacher Judgement Cannot

Peers may observe process evidence unavailable to the teacher. They may know who actually contributed to a group task, whether feedback was useful, how collaboration unfolded, or whether a draft changed after discussion.

That makes peer evidence potentially valuable rather than second-rate. But proximity also creates risks: friendship, conflict, reciprocity, status and fear of retaliation can affect judgement.

The right Bolt move is to preserve the unique information peers can supply while controlling the social conditions that may distort it.

School, Teacher and Student: Three Different Calibration Responsibilities

School

The school should decide whether peer assessment is formative, contributory to grading, or consequential. The higher the stakes, the stronger the need for training, explicit criteria, multiple evidence channels, review procedures and monitoring for systematic distortions.

Teacher or Coach

The teacher should teach students how to assess rather than simply handing them a rubric. Calibration can include scoring exemplar work, discussing disagreement, distinguishing evidence from preference, and practising feedback before peer marks count.

Student

The student should treat peer feedback as evidence to inspect, not a popularity vote and not an oracle. When peers disagree, the disagreement itself can be useful: which criterion produced different interpretations, and what evidence in the work resolves it?

Competing Explanations When Peer and Teacher Scores Disagree

  • The peer misunderstood the rubric.
  • The teacher overlooked evidence the peer noticed.
  • The peer is using a classroom norm rather than the formal standard.
  • The teacher and peer are weighting criteria differently.
  • The peer is influenced by friendship, rivalry or reciprocity.
  • The task is too complex for minimally trained peer raters.
  • The teacher’s own score is imperfect and needs moderation too.

Teacher disagreement does not automatically prove the peer wrong. Peer disagreement does not automatically prove the teacher biased. The disputed criterion should be brought back to observable evidence.

The Bolt Peer-Assessment Calibration Protocol

  1. Name the purpose. Learning, formative feedback, contribution evidence and summative grading require different levels of measurement quality.
  2. Use visible criteria. Peers need a shared standard rather than “give a mark.”
  3. Train with exemplars. Let students score sample work and compare their reasoning with a calibrated reference.
  4. Discuss disagreement. The aim is not instant consensus but better criterion use.
  5. Match task complexity to rater readiness. Complex, ambiguous performances require more support than simple checks.
  6. Control social pressure where needed. Anonymity or privacy may help in some settings, but design effects should be monitored rather than assumed.
  7. Use more than one peer where stakes justify it. Multiple independent judgements can reduce idiosyncratic error, but reviewer overload can degrade quality.
  8. Compare with another evidence channel. Teacher scoring, expert anchors, later individual performance or another calibrated source can test peer-score meaning.
  9. Keep feedback and grading separate where useful. Peer comments can be valuable even when peer marks are not used officially.
  10. Recalibrate the allowed use. Decide whether the evidence is strong enough for feedback only, a small grade component, or a larger summative role.

Worked Example: Three Peers Mark the Same Essay

Three students assess an essay using a four-criterion rubric. They award 14/20, 15/20 and 15/20. The teacher awards 12/20.

Instead of averaging everything to 14 and moving on, the class examines the criterion-level scores. The disagreement is concentrated in “use of evidence.” Peers rewarded the presence of several quotations; the teacher required the quotations to support the argument.

After calibration using two anchor essays, the peers rescore the original response at 12–13/20 and explain why.

The important educational result is larger than the corrected number. The students have learned what the criterion actually means. Their later peer judgements may now become more trustworthy, and their own writing may improve because the standard became operational.

Peer Assessment Can Improve the Assessor, Not Only the Assessed

A major meta-analysis found a small-to-medium positive effect of peer assessment on academic performance across primary, secondary and tertiary contexts. More recent meta-analytic work on assessment-as-learning likewise reports benefits when students actively participate in assessment activities.

This matters because peer assessment should not be judged only by whether student marks perfectly imitate teacher marks. Learning how to detect quality, apply standards and generate feedback can itself be an educational outcome.

But that benefit should not be used to smuggle weak peer grades into high-stakes decisions. Learning value and measurement validity remain separate questions.

Why More Raters Are Not Automatically Better

Adding peer raters can average out individual severity or leniency. Yet every extra review also creates workload. A 2023 study found that task complexity and reviewer load can affect reliability and validity differently; on complex tasks, adding more reviewing burden can even reduce validity.

The design problem is therefore not “get as many ratings as possible.” It is “get enough sufficiently careful, calibrated ratings for the decision being made.”

Common Misconceptions

  • “Students are not qualified to assess anything.” Research shows peer assessment can support both learning and useful judgement under appropriate designs.
  • “If peer and teacher grades correlate, peer grading is validated everywhere.” Validity is local to task, purpose, population and design.
  • “Anonymous peer assessment is always fairer.” Privacy can reduce some social pressures and worsen other processes.
  • “More peers always make the score more accurate.” Rater number interacts with task complexity and workload.
  • “Peer assessment must produce marks.” Qualitative peer feedback can be valuable without numerical grading.
  • “Teacher scores are the unquestionable gold standard.” Teacher judgement also contains measurement error and can require moderation.

How Do We Know?

The large systematic review A Systematic Review of Peer Assessment Design Elements synthesised 449 peer-assessment studies. It found wide design variation and emphasised that reliability and validity should be established for specific peer-assessment implementations rather than assumed from the general literature.

The 2024 review The impact of peer assessment design on interpersonal processes examined 27 studies and found that privacy, format and training can alter peer-assessment processes, with training generally beneficial.

The 2023 study Why increasing the number of raters only helps sometimes shows that task complexity and reviewer load can influence reliability and validity in different ways, demonstrating why rater-count rules should not be mechanical.

The meta-analysis The Impact of Peer Assessment on Academic Performance synthesised 54 studies and found an overall positive effect on academic performance. A 2024 meta-analysis of assessment-as-learning likewise reports positive learning effects from students’ active participation in assessment.

Evidence boundary: peer-assessment research spans very different ages, subjects, technologies and purposes. Agreement observed in one design should not be transported directly into another school’s high-stakes grading policy without local calibration evidence.

For Parents: Peer Feedback Can Be Useful Without Being Final

If your child says, “My classmates gave me a much higher mark than the teacher,” the useful question is not who is right by status. Ask which criterion produced the disagreement and what evidence in the work supports each judgement. If peers were using a different standard, that can be corrected. If they noticed something the teacher missed, that matters too.

Bolt Direction Graph

Student performance → peer criteria → peer judgement/feedback → inspect agreement and disagreement → compare with calibrated anchors or another evidence channel → identify social/task effects → retrain/recalibrate raters → decide appropriate use → later performance receipt.

Useful neighbours include A Group Grade Is Not Automatically an Individual Capability Score, When Two Good Teachers Give Different Marks, and Student Feedback About Teaching Is Evidence, Not a Verdict.