CAN ARTIFICIAL INTELLIGENCE (AI) JUDGE REFLECTION? VALIDITY, RELIABILITY, AND FAIRNESS OF GENERATIVE AI-ASSISTED ASSESSMENT IN POSTGRADUATE HEALTH PROFESSIONS EDUCATION
DOI:
https://doi.org/10.37762/jgmds.13-1.835Keywords:
Artificial Intelligence, Writing, Health Professions Education, Validity, Reliability, Gibbs’ Reflective CycleAbstract
ABSTRACT
OBJECTIVES
This study aimed to evaluate the role of GenAI in assessing reflective writing among Master of Health Professions Education (MHPE) students by comparing GenAI scores with those of human raters, examining subgroup fairness, exploring stakeholder perceptions, and proposing governance recommendations.
METHODOLOGY
A sequential mixed-methods study was conducted in an MHPE programme at Khyber Medical University, Pakistan. In Phase I, 120 Gibbs-structured reflections from 40 students were scored by three trained faculty raters and a GPT-4-level GenAI model using an eight-dimensional rubric. Inter-rater reliability, AI-human agreement, and subgroup differences by gender, discipline, and career stage were examined. In Phase II, semi-structured interviews were conducted with 10 MHPE students and the three faculty raters. Data were analysed using reflexive thematic analysis and integrated with quantitative results.
RESULTS
Human scoring demonstrated strong reliability (ICC = .82). GenAI showed high alignment with human ratings for surface-level dimensions such as clarity and language mechanics (r = .81-.84), but only modest agreement for higher-order reflective constructs including feelings, analysis, and conclusion (r = .49-.59). Exploratory subgroup analyses revealed no statistically significant differences in AI-human discrepancies. However, qualitative accounts highlighted concerns about linguistic and cultural fairness. Participants valued AI for efficient, organised feedback but consistently emphasised its inability to interpret emotional nuance, contextual meaning, or developmental trajectories. Faculty stressed the irreplaceability of human judgment and the need for transparent governance and fairness monitoring.
CONCLUSION
GenAI can effectively support the assessment of structural and linguistic aspects of reflective writing but remains limited in evaluating deeper reflective constructs central to postgraduate learning. Ethical and educationally sound integration requires hybrid human-AI approaches in which AI provides formative support while human evaluators retain primary responsibility for interpretive judgment, fairness oversight, and professional mentorship. GenAI should supplement, not replace, human assessment of reflective writing.
Downloads
Metrics
References
Schön DA. The reflective practitioner: How professionals think in action. London: Temple Smith; 1983
Mann K, Gordon J, MacLeod A. Reflection and reflective practice in health professions education: a systematic review. Adv Health Sci Educ Theory Pract. 2009;14(4):595–621. https://doi.org/10.1007/s10459-007-9090-2. PMID: 18274874 DOI: https://doi.org/10.1007/s10459-007-9090-2
Sandars J. The use of reflection in medical education: AMEE Guide No. 44. Med Teach. 2009;31(8):685–95. https://doi.org/10.1080/01421590903050374. PMID: 19811128 DOI: https://doi.org/10.1080/01421590903050374
Adeani IS, Febriani RB, Syafryadin S. Using Gibbs’ reflective cycle in making reflections of literary analysis. Indonesian EFL J. 2020;6(2):139–48. https://doi.org/10.25134/ieflj.v6i2.3385 DOI: https://doi.org/10.25134/ieflj.v6i2.3382
Jasper M, Rosser M. Reflection and reflective practice. In: Jasper M, Rosser M, editors. Professional development, reflection, and decision-making in nursing and healthcare. Chichester: Wiley-Blackwell; 2013. p. 41–82
Moon J. Using reflective learning to improve the impact of short courses and workshops. J Contin Educ Health Prof. 2004;24(1):4–11. https://doi.org/10.1002/chp.1340240103.PMID: 15069907 DOI: https://doi.org/10.1002/chp.1340240103
Ryan M, Ryan M. Theorising a model for teaching and assessing reflective learning in higher education. High Educ Res Dev.2013;32(2):244–57. https://doi.org/10.1080/07294360.2012.661704 DOI: https://doi.org/10.1080/07294360.2012.661704
Wald HS, Borkan JM, Taylor JS, Anthony D, Reis SP. Fostering and evaluating reflective capacity in medical education: developing the REFLECT rubric for assessing reflective writing. Acad Med. 2012;87(1):41–50. https://doi.org/10.1097/ACM.0b013e31823b55fa.PMID:22104058 DOI: https://doi.org/10.1097/ACM.0b013e31823b55fa
Williamson S, Seewoodhary R. A review and reflection on the visual rehabilitation progress of an older person following cataract surgery two years on. Int J Ther Rehabil. 2016;23(5):242–6. https://doi.org/10.12968/ijtr.2016.23.5.242 DOI: https://doi.org/10.12968/ijtr.2016.23.5.242
Kumar P. Large language models (LLMs): survey, technical frameworks, and future challenges. Artif Intell Rev. 2024;57(10):260. https://doi.org/10.1007/s10462-024-10874-9 DOI: https://doi.org/10.1007/s10462-024-10888-y
Gallegos IO, Rossi RA, Barrow J, Tanjim MM, Kim S, Dernoncourt F, et al. Bias and fairness in large language models: a survey. Comput Linguist. 2024;50(3):1097–179. https://doi.org/10.1162/coli_a_00523 DOI: https://doi.org/10.1162/coli_a_00524
Liang J-C, Hwang G-J, Chen M-RA, Darmawansah D. Roles and research foci of artificial intelligence in language education: an integrated bibliographic analysis and systematic review approach. Interact Learn Environ. 2023;31(7):4270–96. https://doi.org/10.1080/10494820.2021.1982642 DOI: https://doi.org/10.1080/10494820.2021.1958348
Williamson B, Macgilchrist F, Potter J. Re-examining AI, automation and datafication in education. Abingdon: Routledge/Taylor & Francis; 2023. p. 1–5 DOI: https://doi.org/10.1080/17439884.2023.2167830
Lee D, Arnold M, Srivastava A, Plastow K, Strelan P, Ploeckl F, et al. The impact of generative AI on higher education learning and teaching: a study of educators’ perspectives. Comput Educ Artif Intell. 2024;6:100221. https://doi.org/10.1016/j.caeai.2024.100221 DOI: https://doi.org/10.1016/j.caeai.2024.100221
Creswell JW, Plano Clark VL. Revisiting mixed methods research designs twenty years later. Handb Mixed Methods Res Des. 2023;1(1):21–36 DOI: https://doi.org/10.4135/9781529614572.n6
Hox J, de Leeuw E, Klausch T. Mixed-mode research: issues in design and analysis. In: Biemer PP, de Leeuw ED, Eckman S, Kreuter F, Lyberg LE, Tucker C, West BT, editors. Total survey error in practice. Hoboken (NJ): Wiley; 2017. p. 511–30 DOI: https://doi.org/10.1002/9781119041702.ch23
Raudenbush SW, Bryk AS. Hierarchical linear models: applications and data analysis methods. 2nd ed. Thousand Oaks (CA): Sage Publications; 2002
Kember D, McKay J, Sinclair K, Wong FKY. A four-category scheme for coding and assessing the level of reflection in written work. Assess Eval High Educ. 2008;33(4):369–79. https://doi.org/10.1080/02602930701293355 DOI: https://doi.org/10.1080/02602930701293355
Polit DF, Beck CT, Owen SV. Is the CVI an acceptable indicator of content validity? Appraisal and recommendations. Res Nurs Health. 2007;30(4):459–67. https://doi.org/10.1002/nur.20199. PMID: 17654487 DOI: https://doi.org/10.1002/nur.20199
Jonsson A, Svingby G. The use of scoring rubrics: reliability, validity and educational consequences. Educ Res Rev. 2007;2(2):130–44. https://doi.org/10.1016/j.edurev.2007.05.002 DOI: https://doi.org/10.1016/j.edurev.2007.05.002
Chinta SV, Wang Z, Yin Z, Hoang N, Gonzalez M, Quy TL, et al. FairAIED: Navigating fairness, bias, and ethics in educational AI applications. arXiv preprint arXiv:2407.18745. 2024. Available from: https://arxiv.org/abs/2407.18745
Williamson DM, Xi X, Breyer FJ. A framework for evaluation and use of automated scoring. Educ Meas Issues Pract. 2012;31(1):2–13. https://doi.org/10.1111/j.1745-3992.2011.00223.x DOI: https://doi.org/10.1111/j.1745-3992.2011.00223.x
Braun V, Clarke V. Toward good practice in thematic analysis: avoiding common problems and becoming a knowing researcher. Int J Transgend Health. 2023;24(1):1–6. https://doi.org/10.1080/26895269.2022.2026619. PMID: 36608079 DOI: https://doi.org/10.1080/26895269.2022.2129597
Rankin J, McFadyen J. The role of gatekeepers in research: learning from reflexivity and reflection. J Nurs Health Care (JNHC). 2016;4(1):218–9
McKim C. Meaningful member-checking: a structured approach to member-checking. Am J Qual Res. 2023;7(2):41–52. https://doi.org/10.29333/ajqr/13188
Kane MT. Validation as a pragmatic, scientific activity. J Educ Meas. 2013;50(1):115–22. https://doi.org/10.1111/jedm.12007 DOI: https://doi.org/10.1111/jedm.12007
Messick S. Validity of psychological assessment: validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. Am Psychol. 1995;50(9):741–9. https://doi.org/10.1037/0003-066X.50.9.741 DOI: https://doi.org/10.1037//0003-066X.50.9.741
Gearing N. What is the ideal methodological response for the learning and teaching of critical thinking and evaluative judgement in the age of generative artificial intelligence? ATLAANZ J. 2024;7(1):1–9. Available from: https://www.atlaanzjournal.org/article/ai-critical-thinking DOI: https://doi.org/10.26473/ATLAANZ.2024/006
UNESCO. Recommendation on the ethics of artificial intelligence. Paris: United Nations Educational, Scientific and Cultural Organization (UNESCO); 2021. Available from: https://www.unesco.org/en/artificial-intelligence/recommendation-ethics
Corrêa NK, Galvão C, Santos JW, Del Pino C, Pinto EP, Barbosa C, et al. Worldwide AI ethics: a review of 200 guidelines and recommendations for AI governance. Patterns. 2023;4(10):100818. https://doi.org/10.1016/j.patter.2023.100818. PMID: 37827204 DOI: https://doi.org/10.1016/j.patter.2023.100857
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Brekhna Jamil, Nowshad Asim

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
















