Accuracy / Methodology

Accuracy needs a method, not a headline number

PaperGrader does not present a single universal accuracy figure for handwritten answer-sheet evaluation. Performance varies with the subject, rubric, handwriting, scan quality, question format and review process. We describe the measures an institution should examine before relying on an AI-assisted workflow.

For institutions evaluating evidence before deployment.

Examination teams, academic leaders and procurement reviewers can use this methodology to plan a representative pilot and judge whether a workflow is suitable for a particular assessment.

A single percentage can conceal the error that matters most.

OCR quality, question mapping and scoring agreement are distinct stages. A high result on one stage does not establish reliability for the complete workflow. A responsible evaluation therefore measures each stage, inspects failures and records how reviewer interventions affect the final outcome.

How the evaluation workflow operates

  1. 01

    Define the reference

    Freeze the assessment, marking rubric and independent human-reference process before comparing results.

  2. 02

    Sample representative scripts

    Include appropriate variation in handwriting, scan quality, subjects, mark ranges and answer formats.

  3. 03

    Measure the stages

    Evaluate OCR quality, question-mapping coverage, score agreement and review or override behaviour separately.

  4. 04

    Inspect failures

    Review low-confidence, unmapped and materially different results instead of reporting averages alone.

  5. 05

    Decide deployment controls

    Use the findings to set review thresholds, evaluator procedures and the assessments for which the workflow is appropriate.

Human review is part of the measured system.

The operational result is the approved score, not an unreviewed model output. Institutions should track when reviewers correct suggestions, why they do so and whether particular question types need additional checks. Metrics such as MAE, RMSE and weighted agreement can be useful only when their definitions, dataset and reference-grading method are published.

Operational benefits

  • Separates OCR, question mapping and scoring quality
  • Focuses on agreement with an independent human reference
  • Includes override and review burden in the evaluation
  • Encourages subgroup and failure analysis where lawful and meaningful
  • Avoids unsupported generalised accuracy claims

Continue the evaluation

Discuss your institution's evaluation workflow