The Evaluator Training Gap That Brought Down CBSE OSM — And the 4-Week Protocol That Prevents It
CBSE's On-Screen Marking rollout failed in part because faculty evaluators received inadequate preparation before live deployment. Here is the structured training protocol that universities must follow before going live.

What the February 26 Warning Looked Like
On February 26, 2026, CBSE conducted a mock evaluation session intended to familiarise faculty evaluators with its On-Screen Marking (OSM) portal before live deployment. Teachers across the country logged in and were asked to evaluate sample answer sheets.
The feedback was unambiguous and immediate. Faculty reported portal failures, slow-loading answer sheets, an unfamiliar interface, difficulty navigating between question responses, and confusion about how to enter marks for different question types. Several evaluation centres reported that teachers were unable to complete even a partial mock bundle within the allotted session time.
CBSE proceeded with live deployment.
By May 2026, the consequences were visible in the national result data: the Class 12 pass percentage dropped to 85.20%, the lowest in seven years. Approximately 1,63,800 students were placed in compartment — a figure that multiple analysts traced partly to evaluator errors in the marking process, including unmarked answers and incorrect mark entry that went undetected without the quality checks a trained evaluator would catch.
The February warning was not a technical failure. It was a training signal that the institution did not act on.
Why Evaluator Training Is Different From Software Training
The common mistake institutions make when deploying on-screen marking is treating it as a software rollout — hand evaluators login credentials, show them a 45-minute tutorial video, and consider onboarding complete.
This conflates two different kinds of learning.
Software orientation covers where to click, how to log in, how to navigate. It can be completed in one session. Evaluators who complete software orientation know the portal exists and how to open it.
Evaluator calibration covers something harder: developing consistent, accurate, and defensible marks. An evaluator who has spent a career applying red pen to paper is performing a different cognitive task than an evaluator clicking through a digital marking rubric on a monitor. The transition between these modes is not automatic and is not achieved by watching a tutorial.
The evidence for this distinction comes from the OSM implementation literature globally. Cambridge Assessment, which has operated digital marking since 2012, documented that evaluators required between 15 and 25 hours of structured practice before their digital marking consistency matched their paper-based marking consistency. The gap was not skill — it was unfamiliarity with the medium.
CBSE deployed a digital medium to evaluators who had no prior experience with it and then measured their outputs as if medium familiarity were given. The result was predictable.
The 4-Week Evaluator Training Protocol
The following protocol is designed for universities deploying on-screen marking for the first time. It assumes evaluators are subject-matter experts — faculty with established marking capability — and focuses on developing medium familiarity and calibration, not subject knowledge.
Week 1: Baseline Assessment and Portal Orientation
Objective: Establish each evaluator's starting digital literacy level and familiarise the full cohort with the evaluation portal interface before any marks are entered.
Activities:
Outcome target: Every evaluator should be able to navigate to the correct response for any question in a 20-page answer sheet within 90 seconds. Time this explicitly — it is a measurable readiness indicator.
Common problem at Week 1: Evaluators accustomed to physically turning pages find scrolling or clicking between image panels disorienting. Address this explicitly rather than assuming it will resolve with practice.
Week 2: Calibration Exercises
Objective: Develop inter-rater consistency — the ability for multiple evaluators to assign marks within an acceptable variance range for the same answer.
Activities:
Inter-rater reliability target: By the end of Week 2, marks from different evaluators on the same response should be within 10% of the maximum marks for that question at least 85% of the time. This is a standard benchmark used by professional assessment bodies.
Documentation: Record each evaluator's calibration session marks and their deviation from the established benchmark. This record serves two purposes — it guides further training for evaluators who need it, and it constitutes audit evidence that calibration was conducted.
Week 3: Mock Evaluation With Feedback Loops
Objective: Complete full evaluation bundles under near-live conditions, identify systematic errors, and correct them before live deployment.
Activities:
- Questions left unmarked
- Marks entered in the wrong field
- Mark entries that appear to be outliers given the standard for that question
Screen fatigue protocol: Week 3 is where screen fatigue first becomes a real factor. Structure mock sessions to match the intended live session duration — typically 90 to 120 minutes of active evaluation, followed by a 15-minute break. Do not conduct five-hour mock sessions under the assumption that evaluators who can sustain that duration once will sustain it in live deployment. They will not.
Week 3 completion criteria: An evaluator who completes two mock bundles with a discrepancy rate (questions left unmarked, mark entry errors, moderation-triggering variance from the sample standard) below 5% is ready for supervised live evaluation.
Week 4: Supervised Live Evaluation and Escalation Protocol
Objective: Conduct live evaluation of actual answer sheets with a support system in place, and establish clear escalation paths for every foreseeable problem.
Activities:
- Answer sheet image is unreadable — what does the evaluator do?
- Portal becomes unresponsive — save state and wait, or close and reopen?
- Evaluator is uncertain how to mark a specific response type — who is the query contact?
- Marks entered incorrectly — what is the correction window?
Documentation for NAAC evidence: The Week 4 supervised session produces the first entries in what should become a continuous evaluator performance record. Document attendance, bundle completion rates, escalation queries logged, and resolution times. This documentation supports NAAC Criterion 2.5 evidence showing that evaluation was conducted under a managed, quality-assured process.
Ergonomics: The Training Component That Is Always Omitted
Faculty evaluators in India typically mark paper answer sheets at a desk with the script in front of them, making annotations in red pen for up to six or seven hours during an evaluation camp. The physical posture and visual engagement pattern for this task is different from on-screen marking.
On-screen marking requires:
Institutions that set up evaluation workstations without considering monitor height, ambient lighting, chair ergonomics, and keyboard placement will find evaluators fatiguing faster than expected — and the quality of marks entered toward the end of a session declining relative to the beginning.
Minimum ergonomic setup:
These are not expensive requirements. They are consistently ignored in the rush to deploy, and consistently cited by evaluators as contributors to error rates.
The Evaluator Performance Record
Every digital evaluation platform generates a back-end log of evaluator activity: time of each mark entry, time spent on each answer sheet, variance from co-evaluator in double valuation, moderation flags triggered. Most institutions never look at this data.
This data is not only a quality assurance tool — it is NAAC and IQAC evidence of a functioning evaluation governance process. An institution that can show its NAAC peer team:
— is presenting evidence of Criterion 2.5 compliance that very few institutions in India can currently demonstrate.
CBSE's February 26 mock session was the right idea, executed once, too late, and with no follow-up. A university has the advantage of choosing its own timeline. There is no mandate forcing a two-week deployment. The February 26 lesson is not that mock sessions are useless — it is that one session, weeks before deployment, is not a protocol.
---
Related Reading
Ready to digitize your evaluation process?
See how MAPLES OSM can transform exam evaluation at your institution.