The Development of an Automated Essay Scoring (AES) Dataset for Indonesian Language Using OCR and Inter-Rater Validation
DOI:
https://doi.org/10.31098/cset.v5i1.1133Keywords:
Automated Essay Scoring, Indonesian Natural Language Processing, Optical Character Recognition, Dataset Construction, Inter-Rater ReliabilityAbstract
The availability of Automated Essay Scoring (AES) datasets remains limited for low-resource languages such as Indonesian. This study develops and validates a reliable Indonesian AES dataset drawn from 219 authentic handwritten student exams. The methodology uses a multi-stage pipeline that includes image digitization, Optical Character Recognition (OCR) text extraction with meticulous manual verification, and a rigorous cross-validation protocol conducted by three independent expert annotators. For the per-item reliability subset (n = 59), the expert panel's consensus achieved good-to-excellent agreement (ICC = 0.819–0.914). In contrast, single-evaluator assessments agreed only moderately with the panel (QWK ≤ 0.642), showing coarse grading tendencies and wider limits of agreement. Consequently, the panel's consensus average was established as the dataset's gold-standard label. Overall, this research contributes a highly reliable, structured educational dataset that provides a robust and ecologically valid foundation for training and evaluating AES models in low-resource contexts.

