| 8,390 | 74 | 665 |
| 下载次数 | 被引频次 | 阅读次数 |
人机协同评价结合人的主观判断和机器的数据处理能力,可创造出更高效、准确和个性化的教育评价体系,从而推动决策优化、增强教育效果和提升服务质量。生成式人工智能具备智能交互、上下文语义理解能力,是人机“协同”理念落地的重要手段。基于此,文章构建了一个生成式人工智能支持的人机协同评价实践模式,包括确立了以协同论为支撑的人机协同评价模型,设计人机协同评价系统,形成确定评价指标体系、评价文本预处理、设定评价提示语、输出评价反馈和综合分析的人机协同评价实践路径,并进一步将实践路径转化为可供教育评价者操作的人机协同评价支架。在此基础上,文章以上海市H大学开展的基于问题解决的主观作业评价活动为例,解释了如何应用生成式人工智能支持人机协同评价。研究发现,人机协同评价与教师评价结果一致性高,但教育评价者仍需持续发展人机协同能力与批判性思维,以促进人机协同评价的发展。
Abstract:[1]朱丽,李卫霞.新时代教育评价改革的伦理审思[J].上海教育评估研究,2021(1):7-11.
[2]何文涛,张梦丽,逯行,等.人工智能视域下人机协同教学模式构建[J].现代远距离教育,2023(2):78-87.
[3]宋乃庆,郑智勇,周圆林翰.新时代基础教育评价改革的大数据赋能与路向[J].中国电化教育,2021(2):1-7.
[4][美]丹尼尔·戈尔曼,[法]奥利弗·布兰查德.共生:4.0时代的人机关系[M].杨薇,译.北京:中国科学技术出版社,2023.
[5]詹泽慧,季瑜,牛世婧,等.ChatGPT嵌入教育生态的内在机理、表征形态及风险化解[J].现代远距离教育,2023(4):3-13.
[6]江进林,陈丹丹.主观题自动评分研究——回顾、反思与展望[J].中国外语,2021(6):58-64.
[7]Das B,Majumder M,Sekh A A,et al.Automatic question generation and answer assessment for subjective examination[J].Cognitive systems research,2022(72):14-22.
[8]Zhu Y,Li Z,Li Y,et al.Automatic Generation of Graduation Thesis Comments Based on Multilevel Analysis[C]//International Conference of Pioneering Computer Scientists,Engineers and Educators.Singapore:Springer Nature Singapore,2022:67-79.
[9]Bernius J P,Krusche S,Bruegge B.Machine learning based feedback on textual student answers in large courses[J].Computers and Education:Artificial Intelligence,2022(3):100081.
[10]Kozierok R,Aberdeen J,Clark C,et al.Assessing open-ended human-computer collaboration systems:applying a hallmarks approach[J].Frontiers in artificial intelligence,2021(4):670009.
[11]del Gobbo E,Guarino A,Cafarelli B,et al.Automatic evaluation of open-ended questions for online learning.A systematic mapping[J].Studies in Educational Evaluation,2023(77):101258.
[12]李艳,刘淑君,李小丽,等.人机协同作文评价能促进写作教学吗?——来自 Z 校拓展课的证据[J].现代远程教育研究,2022(1):63-74.
[13]杨晓哲,王若昕.困局与破局:教育数字化转型的下一步[J].华东师范大学学报 (教育科学版),2023(3):82-90.
[14]董艳,李心怡,郑娅峰,等.智能教育应用的人机双向反馈:机理,模型与实施原则[J].开放教育研究,2021(2):26-33.
[15]肖君,白庆春,陈沫,等.生成式人工智能赋能在线学习场景与实施路径[J].电化教育研究,2023(9):57-63+99.
[16]李海峰,王炜.生成式人工智能时代的学生作业设计与评价[J].开放教育研究,2023(3):31-39.
[17]Guo K,Wang D.To resist it or to embrace it?Examining ChatGPT’s potential to support teacher feedback in EFL writing[J].Education and Information Technologies,2023(8):1-29.
[18]李汉卿.协同治理理论探析[J].理论月刊,2014(1):138-142.
[19]Corning P A.“The synergism hypothesis”:On the concept of synergy and its role in the evolution of complex systems[J].Journal of social and evolutionary systems,1998(2):133-172.
[20]Wang Q,Wan Y,Feng F.Human-machine collaborative scoring of subjective assignments based on sequential three-way decisions[J].Expert Systems with Applications,2023(216):119466.
[21]Li Q,Cui L,Kong L,et al.Collaborative Evaluation:Exploring the Synergy of Large Language Models and Humans for Open-ended Generation Evaluation[J].arXiv preprint arXiv 2023,2310.19740.
[22]Geng B,Varshney P K.Human-machine collaboration for smart decision making:current trends and future opportunities[C]//2022 IEEE 8th International Conference on Collaboration and Internet Computing (CIC).IEEE,2022:61-67.
[23]Zhang Y,Ren P,de Rijke M.A human-machine collaborative framework for evaluating malevolence in dialogues[C]//Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1:Long Papers).2021(8):5612-5623.
[24]Ren M,Chen N,Qiu H.Human-machine collaborative decision-making:An evolutionary roadmap based on cognitive intelligence[J].International Journal of Social Robotics,2023(7):1101-1114.
[25]连燕华,马晓光.评价要素系统结构分析及模型的建立[J].研究与发展管理,2000(4):17-20+44.
[26]肖新发.评价要素论[J].武汉大学学报(人文科学版),2004(5):523-528.
[27]李亦菲.教育评价的要素和结构[J].基础教育课程,2005(1):43-45.
[28]张远增.论教育评价系统的结构与机制[J].上海教育评估研究,2013(4):1-5.
[29]童雨溪.新时代我国教育评价制度创新研究[D].长沙:湖南大学,2021.
[30]王焕霞.学业评价与课程标准一致性指标体系的构建研究[J].现代基础教育研究,2020(3):79-86.
[31]周光礼,袁晓萍.聚焦“四个评价”深化教育评价机制改革[J].中国考试,2020(8):1-5.
[32]顾明远.对深化新时代教育评价改革的几点认识[J].教育测量与评价,2020(8):3-5+18.
[33]Zhao W X,Zhou K,Li J,et al.A survey of large language models[J].arXiv preprint arXiv 2023,2303.18223.
[34]Van Horne K,Penuel W R,Bell P.Integrating science practices into assessment tasks[J].STEM Teaching Tools,2016(2):1-16.
[35]Hackl V,Müller A E,Granitzer M,et al.Is GPT-4 a reliable rater?Evaluating Consistency in GPT-4 Text Ratings[J].arXiv preprint arXiv2023,2308.02575.
[36]Bridgeman B,Trapani C,Attali Y.Comparison of human and machine scoring of essays:Differences by gender,ethnicity,and country[J].Applied Measurement in Education,2012(1):27-40.
[37]Maestrales S,Zhai X,Touitou I,et al.Using machine learning to score multi-dimensional assessments of chemistry and physics[J].Journal of Science Education and Technology,2021(30):239-254.
[38]Freitag M,Foster G,Grangier D,et al.Experts,errors,and context:A large-scale study of human evaluation for machine translation[J].Transactions of the Association for Computational Linguistics,2021(9):1460-1474.
[39]Joulin A,Grave E,Bojanowski P,et al.Bag of tricks for efficient text classification[J].arXiv preprint arXiv 2016,1607.01759.
[40]王一岩,刘淇,郑永和.人机协同学习:实践逻辑与典型模式[J].开放教育研究,2024(1):65-72.
[41]Hendrycks D,Burns C,Basart S,et al.Aligning ai with shared human values[J].arXiv preprint arXiv 2020,2008.02275.
[42]Burston J.Twenty years of MALL project implementation:A meta-analysis of learning outcomes[J].ReCALL,2015(1):4-20.
基本信息:
DOI:10.13927/j.cnki.yuan.20240422.001
中图分类号:G434
引用信息:
[1]宛平,顾小清.生成式人工智能支持的人机协同评价:实践模式与解释案例[J].现代远距离教育,2024,No.212(02):33-41.DOI:10.13927/j.cnki.yuan.20240422.001.
基金信息:
2019年度国家社科基金重大项目“人工智能促进未来教育发展研究”(编号:19ZDA364)
2024-04-23
2024-04-23
2024-04-23