Content Validity of an Instrument Measuring Understanding of Generative Artificial Intelligence (AI) Ethics for Malay Language Assessment
Abstract
The rapid advancement of Generative Artificial Intelligence (GenAI) has significantly influenced educational assessment practices, including the teaching and assessment of the Malay Language. The integration of this technology needs a strong ethical understanding among teachers to preserve the integrity, fairness, and authenticity of assessment processes.In line with this, this study aims to evaluate the content validity of an instrument designed to measure teachers’ understanding of GenAI ethics in the context of Malay Language assessment through expert panel evaluation. This study employs a quantitative research design involving five expert panellists, consisting of professional experts and a lay expert, who assessed the instrument items in terms of content relevance, construction, and language clarity. The instrument covers six core GenAI ethics constructs: intellectual property, truthfulness, robustness, malicious use recognition, sociocultural responsibility and human-centric design. Content validity was examined using the Content Validity Ratio (CVR) method based on Lawshe’s (1975) guidelines with three likert points scale. The findings reveal that 50 items achieved the maximum CVR value and were accepted by the expert panel, while only 12 number required refinement to improve clarity and construct alignment. Overall, the results provide strong evidence that the instrument demonstrates high content validity and is suitable for assessing Malay Language teachers’ understanding of ethical GenAI use in assessment. However, this study is limited by the relatively small number of expert panellists involved in the validation process and the focus on content validity alone. Future research is recommended to conduct pilot testing with a larger sample of teachers and to examine additional psychometric properties, such as construct validity and reliability, using statistical approaches such as exploratory factor analysis or Rasch measurement analysis.
Downloads
Downloads
References
Abdullah, H. S. V., Alias, A. K., Tasir, Z., & Sharef, N. M. (2023). MQA Advisory Note No.22023 - AI Generatif (Issue 2).
Aguirre, V. C. de S. P., Dória, J. P. da S., Melo, R. A. de, Fernandes, F. E. C. V., & Mola, R. (2024). Content Validity of an Instrument for Assessing Skin Lesions in a Mother and Child Hospital. Rev Enferm UFPI, 13(1), 1–11. https://doi.org/10.26694/reufpi.v13i1.4579
Amatan, M., K Han, C. G., & Pang, V. (2021). Kesahan kandungan soal selidik faktor konteks, input dan proses terhadap penerimaan pelaksanaan elemen pendidikan STEM dalam pengajaran dan pembelajaran guru menggunakan nisbah kesahan kandungan (CVR). International Journal of Advanced Research in Future Ready Learning and Education, 23(1), 10–22.
Bujang, M., Alias, B. S., & Mansor, A. N. (2025). Evaluating the Content Validity of Instruments for Measuring Principal’s Ethical Leadership Practices Using the Content Validity Ratio (CVR) Method. Journal of Social Sciences & Humanities, 22(1), 99–108. https://doi.org/https://doi.org/10.17576/ebangi.2025.2208.08
Cardona, M. A., Rodríguez, R. J., & Ishmael, K. (2023). Artificial Intelligence and the Future of Teaching and Learning Insights and Recommendations Artificial Intelligence and the Future of Teaching and Learning (Issue May). https://tech.ed.gov
Cohen, R. J., W. Joel Schneider, & Tobin, R. M. (2022). Psychological Testing and Assessment An Introduction to Tests and Measurement (10th ed.). McGraw Hill.
Connell, J., Carlton, J., Grundy, A., Taylor Buck, E., Keetharuth, A. D., Ricketts, T., Barkham, M., Robotham, D., Rose, D., & Brazier, J. (2018). The importance of content and face validity in instrument development: lessons learnt from service users when developing the Recovering Quality of Life measure (ReQoL). Quality of Life Research?: An International Journal of Quality of Life Aspects of Treatment, Care and Rehabilitation, 27(7), 1893–1902. https://doi.org/10.1007/s11136-018-1847-y
Creswell, J. W., & Creswell, J. D. (2018). Research Design: Qualitative and Quantitative Approaches. In SAGE Publications (Fifth edit). https://doi.org/10.2307/328794
Creswell, J. W., & Guetterman, T. C. (2019). Educational Research:Planning, Conducting, and Evaluating Quantitative and Qualitative Research. In Educational Principles and Practice in Veterinary Medicine.
DeVellis, R., & Thorpe, C. (2020). Scale development: Theory and applications (fifth edit). Sage Publications.
Dwivedi, Y. K., Kshetri, N., Hughes, L., Slade, E. L., Jeyaraj, A., Kar, A. K., Baabdullah, A. M., Koohang, A., Raghavan, V., Ahuja, M., Albanna, H., Albashrawi, M. A., Al-Busaidi, A. S., Balakrishnan, J., Barlette, Y., Basu, S., Bose, I., Brooks, L., Buhalis, D., … Wright, R. (2023). “So what if ChatGPT wrote it?” Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. International Journal of Information Management, 71(March), 1–63. https://doi.org/10.1016/j.ijinfomgt.2023.102642
ElKhalil, R., Masuadi, E., Bayoumi, R., AlMekkawi, M., Ahmed, L. A., Al-Rifai, R. H., & Elbarazi, I. (2025). Content validity of the Mental Health Literacy Scale for perinatal use based on expert and patient input. Scientific Reports, 15(1), 1–11. https://doi.org/10.1038/s41598-025-17964-5
Estremera, M. L., & Mendoza -Sarmiento, M. A. (2024). Content Validity and Reliability of Questionnaires: Trends, Prospects and Innovation in the Digital Research Epoch Introduction. ASEAN Innovative and Transformative Education Journal, 1(1), 1–10. www.aitej.gov.ph
Guo, D., Chen, H., Wu, R., & Wang, Y. (2023). AIGC challenges and opportunities related to public safety: A case study of ChatGPT. Journal of Safety Science and Resilience, 4(4), 329–339. https://doi.org/10.1016/j.jnlssr.2023.08.001
Hasibuan, M. S., & Isnanto, R. (2025). Generatif AI (Januari, Issue January). Darmajaya ( DJ ) Press. https://www.researchgate.net/publication/388221462_Buku_Gen_AI_Final_20_jan_2025
Jin, Y., Martinez-Maldonado, R., Gaševi?, D., & Yan, L. (2025). GLAT: The generative AI literacy assessment test. Computers and Education: Artificial Intelligence, 9, 100436. https://doi.org/https://doi.org/10.1016/j.caeai.2025.100436
Kadaruddin, K. (2023). Empowering Education through Generative AI: Innovative Instructional Strategies for Tomorrow’s Learners. International Journal of Business, Law, and Education, 4(2), 618–625. https://doi.org/10.56442/ijble.v4i2.215
Kania, N., Kusumah, Y. S., Dahlan, J. A., Nurlaelah, E., Gürbüz, F., & Bonyah, E. (2024). Constructing and providing content validity evidence through the Aiken’s V index based on the experts’ judgments of the instrument to measure mathematical problem-solving skills. REID (Research and Evaluation in Education), 10(1), 64–79. https://doi.org/10.21831/reid.v10i1.71032
Kementerian Pendidikan Malaysia. (2017). Dasar Pendikan Kebangsaan. Edisi keempat.
Kenthapadi, K., Lakkaraju, H., & Rajani, N. (2023). Generative AI meets Responsible AI: Practical Challenges and Opportunities. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 5805–5806. https://doi.org/10.1145/3580305.3599557
Kizilcec, R. F., Huber, E., Papanastasiou, E. C., Cram, A., Makridis, C. A., Smolansky, A., Zeivots, S., & Raduescu, C. (2024). Perceived impact of generative AI on assessments: Comparing educator and student perspectives in Australia, Cyprus, and the United States. Computers and Education: Artificial Intelligence, 7(November 2023), 1–11. https://doi.org/10.1016/j.caeai.2024.100269
Laine, J., Minkkinen, M., & Mäntymäki, M. (2025). Understanding the Ethics of Generative AI: Established and New Ethical Principles. Communications of the Association for Information Systems, 56(January), 1–24. https://doi.org/10.17705/1CAIS.05601
Lawshe C. H. (1975). A Qualitative Approach to Content Validity. Personnel Psychology, 28, 563–575.
Masuwai, A., Zulkifli, H., & Hamzah, M. I. (2024). Evaluation of content validity and face validity of secondary school Islamic education teacher self-assessment instrument. Cogent Education, 11(1). https://doi.org/10.1080/2331186X.2024.2308410
Memon, M. A., Ramayah Thurasamyd, Ting, H., & Cheah, J.-H. (2025). Convenience Sampling: A Review and Guidelines for Quantitative Research. Journal of Applied Structural Equation Modeling, 9(June), 1–15. https://doi.org/10.47263/JASEM.9(2)01
Mohamat, R., Sumintono, B., & Hamid, H. S. A. (2022). Analisis Kesahan Kandungan Instrumen Kompetensi Guru untuk Melaksanakan Pentaksiran Bilik Darjah Menggunakan Model Rasch Pelbagai Faset (Content Validity Analysis of an Instrument to Measure Teacher’s Competency for Classroom Assessment Using Many Facet R. Jurnal Pendidikan Malaysia, 47(01), 1–15. https://doi.org/10.17576/jpen-2022-47.01-01
Mohammed Afandi Zainal, Mohd Effendi@Ewan Mohd Matore, Wan NorShuhadah W Musa, & Noor Hashimah Hashim. (2020). Kesahan Kandungan Instrumen Pengukuran Tingkah Laku Inovatif Guru Menggunakan Kaedah Nisbah Kesahan Kandungan (CVR). Akademika, 90(Isu Khas 3), 43–54. http://ejournals.ukm.my/akademika/article/view/42052
Mohd Matore, M. E. E., Zainal, M. A., Mohd Noh, M. F., Khairani, A. Z., & Abd Razak, N. (2021). The Development and Psychometric Assessment of Malaysian Youth Adversity Quotient Instrument (MY-AQi) by Combining Rasch Model and Confirmatory Factor Analysis. IEEE Access, 9(Mi), 13314–13329. https://doi.org/10.1109/ACCESS.2021.3050311
Mohd Matore, M. E., Idris, H., Normawati, A. R., & Khairani, A. Z. (2017). Kesahan Kandungan Pakar Instrumen IKBAR Bagi Pengukuran AQ Menggunakan Nisbah Kesahan Kandungan. Proseeding of International Conference On Global Education V (ICGE V), Padang Indonesia [10-11 April 2017], April, 979–997.
Mökander, J., Schuett, J., Kirk, H. R., & Floridi, L. (2023). Auditing Large Language Models: A Three-Layered Approach. In SSRN Electronic Journal (Vol. 4, Issue 4). Springer International Publishing. https://doi.org/10.2139/ssrn.4361607
Mokkink, L. B., Herbelet, S., Tuinman, P. R., & Terwee, C. B. (2025). Content validity: judging the relevance, comprehensiveness, and comprehensibility of an outcome measurement instrument – a COSMIN perspective. Journal of Clinical Epidemiology, 185, 1–5. https://doi.org/10.1016/j.jclinepi.2025.111879
Polit, D. F., & Beck, C. T. (2006). The content validity index: Are you sure you know what’s being reported? critique and recommendations. Research in Nursing & Health, 29(5), 489–497. https://doi.org/https://doi.org/10.1002/nur.20147
Raub, S. A., Mansor, M., Ishak, R., Pengurusan, F., Pendidikan, U., & Idris, S. (2023). Kesahan Kandungan Instrumen Pengukuran Pengurusan Pengetahuan Organisasi Pembelajaran Menggunakan Kaedah Nisbah Kesahan Kandungan (CVR). Jurnal Pendidikan Bitara UPSI, 16(2), 59–70.
Roebianto, A., Savitri, S. I., Aulia, I., Suciyana, A., & Mubarokah, L. (2023). Content Validity: Definition and Procedure of Content Validation in Psychological Research. TPM - Testing, Psychometrics, Methodology in Applied Psychology, 30(1), 5–18. https://doi.org/10.4473/TPM30.1.1
Rubio, D. M. G., Berg-Weger, M., Tebb, S. S., Lee, E. S., & Rauch, S. (2003). Objectifyng content validity: Conducting a content validity study in social work research. Social Work Research, 27(2), 94–104. https://doi.org/10.1093/swr/27.2.94
Samsudin, S. binti, Saleh, H. bin M., & Ahmad, A. S. bin. (2024). Persepsi Bakal Guru Terhadap Kesan Aplikasi kecerdasan Buatan (AI) dalam Pengajaran dan Pembelajaran. International Journal of Educational Research on Andragogy and Pedagogy, 2(1), 112–124. https://ijerap.com/index.php/ijerap/article/view/27
Sovey, S., Osman, K., & Matore, M. E. E. M. (2022). Rasch Analysis for Disposition Levels of Computational Thinking Instrument Among Secondary School Students. Eurasia Journal of Mathematics, Science and Technology Education, 18(3), 2–15. https://doi.org/10.29333/ejmste/11794
Tahir, M. H. M., Saputra, S., Othman, S., Shah, D. S. M., Sulaiman, S. H., Azhar, M. A., & Mohandas, E. S. (2025). Online Assessment in Higher Education: A Systematic Review. Multidisciplinary Reviews, 9(1), 1–10. https://doi.org/10.24059/olj.v27i1.3398
Wang, F., & Sahid, S. (2024). Content validation and content validity index calculation for entrepreneurial behavior instruments among vocational college students in China. Multidisciplinary Reviews, 7(9). https://doi.org/10.31893/multirev.2024187
Williams, R. T. (2023). The ethical implications of using generative chatbots in higher education. Frontiers in Education, 8(January). https://doi.org/10.3389/feduc.2023.1331607
Zare, Z., Khosravi, M., Marzaleh, M. A., Jalali, F. S., Izadi, R., & Naseh, H. (2025). Psychometric evaluation of an instrument measuring artificial intelligence utilization in decision-making domains of healthcare organizations. Scientific Reports, 15(1), 1–12. https://doi.org/10.1038/s41598-025-20753-9

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

