Assessing Teaching and Scientific Research in Higher Education in the Era of Generative Artificial Intelligence
Abstract:
The rapid expansion of generative artificial intelligence (GenAI) is changing the conditions under which teaching, learning evidence, and scientific research are produced and assessed in higher education. This study examines these transformations through a qualitative-quantitative documentary analysis of 686 publications published between 2022 and 2026. Each document was coded using 44 binary thematic indicators grouped into six domains: assessment of teaching, assessment of learning, research supervision, assessment by academic juries, GenAI-related risks and limitations, and assessment transformations or good practices. Descriptive thematic frequencies were complemented by Multiple Correspondence Analysis (MCA), Benzécri and Greenacre corrections of inertia, and pairwise Phi coefficients. Publications from 2025 and 2026 accounted for 72.0% of the corpus, confirming the very recent and rapidly expanding character of the field. After Benzécri correction, the first three MCA dimensions explained 70.20% of corrected inertia, compared with 21.74% of raw inertia; the more conservative Greenacre correction yielded 47.86%. Phi analysis showed a strong association between authentic assessment and oral/defence-based assessment (φ = .694; χ² = 330.43; p < .001), while process traceability was positively associated with human validation (φ = .177; χ² = 21.41; p < .001). Across the corpus, integrity, transparency, explicit task-level policies, AI literacy, authentic assessment, traceability, oral justification, and human validation emerged as recurrent responses. The findings indicate a shift from assessment centred on the final product toward an evidence-based ecosystem combining product, process, traceability, explicit justification, dialogue, and accountable human judgment. On this basis, the study proposes the P-P-T-E-V model—Product, Process, Traceability, Explicitation, and Human Validation—as an integrative framework for assessing teaching-related evidence and scientific research in the GenAI era.
KeyWords:
generative artificial intelligence, higher education assessment, scientific research, research supervision, multiple correspondence analysis, academic integrity
References:
- Batista, J., Mesquita, A., & Carnaz, G. (2024). Generative AI and higher education: Trends, challenges, and future directions from a systematic literature review. Information, 15(11), 676. https://doi.org/10.3390/info15110676
- Bearman, M., Ryan, J., & Ajjawi, R. (2024). Generative AI and higher education: A review of claims from the first months of ChatGPT. Higher Education. https://doi.org/10.1007/s10734-024-01265-3
- Belkina, M., Daniel, S., Nikolic, S., Haque, R., Lyden, S., Neal, P., Grundy, S., & Hassan, G. M. (2025). Implementing generative AI (GenAI) in higher education: A systematic review of case studies. Computers and Education: Artificial Intelligence, 100407. https://doi.org/10.1016/j.caeai.2025.100407
- Dai, W., Tsai, Y.-S., Lin, J., Aldino, A., Jin, H., Li, T., Gašević, D., & Chen, G. (2024). Assessing the proficiency of large language models in automatic feedback generation: An evaluation study. Computers and Education: Artificial Intelligence, 7, 100299. https://doi.org/10.1016/j.caeai.2024.100299
- El-Adawy, S., MacDonagh, A., & Abdelhafez, M. (2024). Exploring large language models as formative feedback tools in physics. 2024 Physics Education Research Conference Proceedings, 126–131. https://doi.org/10.1119/perc.2024.pr.El-Adawy
- Escalante, J., Pack, A., & Barrett, A. (2023). AI-generated feedback on writing: Insights into efficacy and ENL student preference. International Journal of Educational Technology in Higher Education, 20, 57. https://doi.org/10.1186/s41239-023-00425-2
- Flodén, J. (2025). Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal, 51(1), 201–224. https://doi.org/10.1002/berj.4069
- Fu, Y., Weng, Z., & Wang, J. (2025). Examining AI use in educational contexts: A scoping meta-review and bibliometric analysis. International Journal of Artificial Intelligence in Education, 35, 1388–1444. https://doi.org/10.1007/s40593-024-00442-w
- Furze, L., Perkins, M., Roe, J., & MacVaugh, J. (2024). The AI Assessment Scale (AIAS) in action: A pilot implementation of GenAI-supported assessment. Australasian Journal of Educational Technology, 40(4). https://doi.org/10.14742/ajet.9434
- Lachheb, A., Leung, J., Abramenka-Lachheb, V., & Sankaranarayanan, R. (2025). AI in higher education: A bibliometric analysis, synthesis, and a critique of research. The Internet and Higher Education, 67, 101021. https://doi.org/10.1016/j.iheduc.2025.101021
- Lodge, J., Howard, S., Bearman, M., Dawson, P., & Associates. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency.
- Luo, J. (2024). A critical review of GenAI policies in higher education assessment: A call to reconsider the ‘originality’ of students’ work. Assessment & Evaluation in Higher Education, 49(5), 651–664. https://doi.org/10.1080/02602938.2024.2309963
- Lye, C. Y., & Lim, L. (2024). Generative artificial intelligence in tertiary education: Assessment redesign principles and considerations. Education Sciences, 14(6), 569. https://doi.org/10.3390/educsci14060569
- Moorhouse, B. L., Yeo, M. A., & Wan, Y. (2023). Generative AI tools and assessment: Guidelines of the world's top-ranking universities. Computers and Education Open, 5, 100151. https://doi.org/10.1016/j.caeo.2023.100151
- Nazaretsky, T., Mejia-Domenzain, P., Swamy, V., Frej, J., & Käser, T. (2024). AI or human? Evaluating student feedback perceptions in higher education. Artificial Intelligence in Education. https://doi.org/10.31219/osf.io/6zm83
- Nguyen, A., Hong, Y., Dang, B., & Huang, X. (2024). Human-AI collaboration patterns in AI-assisted academic writing. Studies in Higher Education, 49, 847–864. https://doi.org/10.1080/03075079.2024.2323593
- Pecuchova, J., Benko, Ľ., & Drlik, M. (2025). Automated grading of open-ended questions in higher education using GenAI models. International Journal of Artificial Intelligence in Education, 35(6), 3813–3846. https://doi.org/10.1007/s40593-025-00517-2
- Perkins, M., Furze, L., Roe, J., & MacVaugh, J. (2024). The Artificial Intelligence Assessment Scale (AIAS): A framework for ethical integration of generative AI in educational assessment. Journal of University Teaching & Learning Practice, 21(6). https://doi.org/10.53761/q3azde36
- Qian, Y. (2025). Pedagogical applications of generative AI in higher education: A systematic review of the field. TechTrends, 69, 1105–1120. https://doi.org/10.1007/s11528-025-01100-1
- Rubin, M. (2023). Use and misuse of corrections for multiple testing. Methods in Psychology, 8, 100120. https://doi.org/10.1016/j.metip.2023.100120
- Sahar, R., & Munawaroh, M. (2025). Artificial intelligence in higher education with bibliometric and content analysis for future research agenda. Discover Sustainability, 6, 401. https://doi.org/10.1007/s43621-025-01086-z
- Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A ‘Turing Test’ case study. PLOS ONE, 19(6), e0305354. https://doi.org/10.1371/journal.pone.0305354
- Song, N. (2024). Higher education crisis: Academic misconduct with generative AI. Journal of Contingencies and Crisis Management. https://doi.org/10.1111/1468-5973.12532
- Soulage, C. O., Van Coppenolle, F., & Guebre-Egziabher, F. (2024). The conversational AI ‘ChatGPT’ outperforms medical students on a physiology university examination. Advances in Physiology Education, 48(4). https://doi.org/10.1152/advan.00181.2023
- Susnjak, T., & McIntosh, T. R. (2024). ChatGPT: The end of online exam integrity? Education Sciences, 14(6), 656. https://doi.org/10.3390/educsci14060656
- Ma, T. (2024). Systematically visualizing ChatGPT used in higher education: Publication trend, disciplinary domains, research themes, adoption and acceptance. Computers and Education: Artificial Intelligence, 100336. https://doi.org/10.1016/j.caeai.2024.100336
- Wu, F., Dang, Y., & Li, M. (2025). A systematic review of responses, attitudes, and utilization behaviors on generative AI for teaching and learning in higher education. Behavioral Sciences, 15(4), 467. https://doi.org/10.3390/bs15040467
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, 26. https://doi.org/10.1007/s40979-023-00146-z
- Wood, D. A., Achhpilia, M. P., Adams, M. T., Aghazadeh, S., Akinyele, K., Akpan, M., et al. (2023). The ChatGPT artificial intelligence chatbot: How well does it answer accounting assessment questions? Issues in Accounting Education, 38(4), 81–108. https://doi.org/10.2308/ISSUES-2023-013
- Xia, Q., Weng, X., Ouyang, F., Lin, T. J., & Chiu, T. K. F. (2024). A scoping review on how generative artificial intelligence transforms assessment in higher education. International Journal of Educational Technology in Higher Education, 21, 40. https://doi.org/10.1186/s41239-024-00468-z