Rating accuracy is one of the fundamental standards in educational assessment to ensure the quality and integrity. Inaccuracy in academic assessment engenders negative implications towards student’s motivation and raters’ credibility. Therefore, this paper seeks to provide a discussion on rating accuracy in educational assessment based on Brunswik’s lens model. The model contends that raters’ ratings are not completed directly but through the existence of many factors including raters’ variability, rating scales and domains assessed. Raters’ idiosyncrasy is scrutinized by describing varied sources that can threaten rating accuracy. This model explains how intervening factors moderate the relationship between candidates’ capabilities and observed scores. The discussion may shed some light on the endeavors to inspire raters to be effective and uphold the values of reliable raters through the implementation of thoughtful rater training that incorporates scoring practices, exposure on rater bias and self-directed reflection. Future attempts are necessary for understanding the interaction among intervening factors that influence raters and differences of rating accuracy produced by internal and external raters.
AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing: National Council on Measurement in Education. Washington DC: American Educational Research Association.
Alla Baksh, M. A. K., Mohd Sallehhudin, A. A., & Siti Hamin, S. (2019). Examining the Factors Mediating the Intended Washback of the English Language School-Based Assessment: Pre-Service ESL teachers’ Accounts. Pertanika Journal of Social Sciences and Humanities, 27, 51-68. http://www.pertanika.upm.edu.my/Pertanika%20PAPERS/JSSH%20Vol.%2027% 20(1)%20Mar.%202019/4%20JSSH-2622-2017.pdf
Attali, Y. (2016). A Comparison of Newly-Trained and Experienced Raters on a Standardized Writing Assessment. Language Testing, 33, 99-115. https://doi.org/10.1177/0265532215582283
Bijani, H. (2018). Investigating the Validity of Oral Assessment Rater Training Program: A Mixed-Methods Study of Raters’ Perceptions and Attitudes before and after Training. Cogent Education, 33, 1-20. https://doi.org/10.1080/2331186X.2018.1460901
Brunswik, E. (1956). Perception and the Representative Design of Psychological Experiments. Berkeley, CA: University of California Press.
Coaley, K. (2009). An Introduction to Psychological Assessment and Psychometrics. Los Angeles, CA: SAGE.
Cooksey, R. W., Freebody, P., & Wyatt-Smith, C. (2007). Assessment as Judgment-in-Context: Analyzing How Teachers Evaluate Students’ Writing. Educational Research and Evaluation, 13, 401-434. https://doi.org/10.1080/13803610701728311
Davis, L. (2016). The Influence of Training and Experience on Rater Performance in Scoring Spoken Language. Language Testing, 33, 117-135. https://doi.org/10.1177/0265532215582282
Eckes, T. (2015). Introduction to Many-Facet Rasch Measurement: Analyzing and Evaluating Rater-Mediated Assessments (2nd ed.). New York: Peter Lang.
Engelhard, G., Wang, J., & Wind, S. A. (2018). A Tale of Two Models: Psychometric and Cognitive Perspectives on Rater-Mediated Assessments Using Accuracy Ratings. Psychological Test and Assessment Modeling, 60, 33-52. http://www.psychologie-aktuell.com/fileadmin/download/ptam/1-2018_20180323/3_PTAM_ Engelhard__Wang___Wind__2018-03-10__1855.pdf
Haladyna, T. M., & Rodrigues, M. C. (2013). Developing and Validating Test. New York: Routledge. https://doi.org/10.4324/9780203850381
Hennington, C. S., Bradley, L. J., Crews, C., & Hennington, E. A. (2013). The Halo Effect: Considerations for the Evaluation of Counselor Competency (pp. 1-10). Vistas Online. http://counselingoutfitters.com/vistas/VISTAS_Home.htm
Hsieh, C. N. (2011). Rater Effects in ITA Testing: ESL Teachers’ versus American Undergraduates’ Judgments of Accentedness, Comprehensibility, and Oral Proficiency. Spaan Fellow Working Papers in Second or Foreign Language Assessment, 9, 47-74. https://michiganassessment.org/wp-content/uploads/2014/12/Spaan_V9_FULL.pdf#page=55
Huang, L., Kubelec, S., Keng, N., & Hsu, L. (2018). Evaluating CEFR Rater Performance through the Analysis of Spoken Learner Corpora. Language Testing in Asia, 8, 1-17. https://doi.org/10.1186/s40468-018-0069-0
Isaacs, T., & Thomson, R. I. (2013). Rater Experience, Rating Scale Length, and Judgments of L2 Pronunciation: Revisiting Research Conventions. Language Assessment Quarterly, 10, 135-159. https://doi.org/10.1080/15434303.2013.769545
Jabeen, R. (2016). An Investigation into Native and Non Native English Speaking Instructors’ Assessment of University ESL Student’s Oral Presentation. Mankato, MN: Minnesota State University. https://cornerstone.lib.mnsu.edu/etds/647/
Kim, H. J. (2015). A Qualitative Analysis of Rater Behavior on an L2 Speaking Assessment. Language Assessment Quarterly, 12, 239-261. https://doi.org/10.1080/15434303.2015.1049353
Lee, H. (2017). The Effects of Rater’s Familiarity with Test Taker’s L1 in Assessing Accentedness and Comprehensibility of Independent Speaking Tasks. Seoul: Department of English Language and Literature, Seoul National University. http://s-space.snu.ac.kr/handle/10371/139643
Myford, C., & Wolfe, E. (2004). Detecting and Measuring Rater Effects Using Many-Facet Rasch Measurement: Part IL. Journal of Applied Measurement, 5, 189-227. https://www.researchgate.net/profile/Carol_Myford/publication/9069043_Detecting _and_Measuring_Rater_Effects_Using_Many-Facet_Rasch_Measurement_Part_I /links/54cba70e0cf298d6565848ee.pdf
Noor Lide, A. K. (2011). Judging Behaviour and Rater Errors: An Application of the Many-Facet Rasch Model. GEMA Online Journal of Language Studies, 11, 179-197. http://ejournals.ukm.my/gema/article/view/49
Oudman, S., Pol, J. Van De, Bakker, A., Moerbeek, M., & Gog, T. Van. (2018). Effects of Different Cue Types on the Accuracy of Primary School Teachers’ Judgments of Students’ Mathematical Understanding. Teaching and Teacher Education, 76, 1-13. https://doi.org/10.1016/j.tate.2018.02.007
Saadat, M., & Alavi, S. Z. (2018). The Effect of Type of Paragraph on Native and Non-Native English Speakers’ Use of Grammatical Cohesive Devices in Writing and Raters’ Evaluation. 3L: Language, Linguistics, Literature, 24, 97-111. https://doi.org/10.17576/3L-2018-2401-08
Scullen, S. E., Mount, M. K., & Goff, M. (2000). Understanding the Latent Structure of Job Performance Ratings. Journal of Applied Psychology, 85, 956-970. https://doi.org/10.1037/0021-9010.85.6.956
Song, T., Wolfe, E. W., Less-Petersen, M., Sanders, R., & Vickers, D. (2014). Relationship between Rater Background and Rater Performance. https://pdfs.semanticscholar.org/f745/93340d40576dc303eed2c7998806ef27554a.pdf
Südkamp, A., Kaiser, J., & Moller, J. (2012). Accuracy of Teachers’ Judgments of Students’ Academic Achievement: A Meta-Analysis. Journal of Educational Psychology, 104, 743-762. https://doi.org/10.1037/a0027627.supp
Sundqvist, P., Wikstrom, P., Sandlund, E., & Nyroos, L. (2018). The Teacher as Examiner of L2 Oral Tests: A Challenge to Standardization. Language Testing, 35, 217-238. https://doi.org/10.1177/0265532217690782
Vogelin, C., Jansen, T., Keller, S. D., Machts, N., & Moller, J. (2019). The Influence of Lexical Features on Teacher Judgments of ESL Argumentative Essays. Assessing Writing, 39, 50-63. https://doi.org/10.1016/j.asw.2018.12.003
Weilie, L. (2018). To What Extent Do Non-Teacher Raters Differ from Teacher Raters on Assessing Story-Retelling. Journal of Language Testing & Assessment, 1, 1-13. http://clausiuspress.com/assets/default/article/2018/08/29/article_1535590233.pdf https://doi.org/10.23977/langta.2018.11001
Wind, S. A. (2018). Examining the Impacts of Rater Effects in Performance Assessments. Applied Psychological Measurement, 43, 159-171. https://doi.org/10.1177/0146621618789391
Wind, S. A., & Schumacker, R. E. (2017). Detecting Measurement Disturbances in Rater-Mediated Assessments. Educational Measurement: Issues and Practice, 36, 44-51. https://doi.org/10.1111/emip.12164
Wind, S. A., Stager, C., & Patil, Y. J. (2017). Exploring the Relationship between Textual Characteristics and Rating Quality in Rater-Mediated Writing Assessments: An Illustration with L1 and L2 Writing Assessments. Assessing Writing, 34, 1-15. https://doi.org/10.1016/j.asw.2017.08.003
Winke, P., & Gass, S. (2013). The Influence of Second Language Experience and Accent Familiarity on Oral Proficiency Rating: A Qualitative Investigation. TESOL Quarterly, 47, 762-789. https://doi.org/10.1002/tesq.73
Wu, M. (2017). Some IRT-Based Analyses for Interpreting Rater Effects. Psychological Test and Assessment Modeling, 59, 453-470. https://www.psychologie-aktuell.com/fileadmin/download/ptam/4-2017_20171218/04_Wu.pdf
Yahya Ameen, T., Mohd Sallehhudin, A. A., Kemboja, I., & Alla Baksh, M. A. K. (2014). The Washback Effect of the General Secondary English Examination (GSEE) on Teaching and Learning. GEMA Online Journal of Language Studies, 14, 83-103. https://doi.org/10.17576/GEMA-2014-1403-06
Zhang, Y., & Elder, C. (2014). Investigating Native and Non-Native English-Speaking Teacher Raters’ Judgments of Oral Proficiency in the College English Test-Spoken English Test (CET-SET). Assessment in Education: Principles, Policy and Practice, 21, 306-325. https://doi.org/10.1080/0969594X.2013.845547
Zhao, K. (2017). Investigating the Effects of Rater’s Second Language Learning Background and Familiarity with Test-Taker’s First Language on Speaking Test Scores. Provo, UT: Brigham Young University.