Categories
Effective Teaching Approach Higher Education

Supporting self-regulated learning through generative AI feedback in online higher education

A recent mixed-methods study by Yilmaz and colleagues examined whether feedback generated by artificial intelligence can help university students develop self-regulated learning skills (SRLs) in an online course, and how students’ perceptions of the feedback source shape their attitudes. The study was conducted within a nine-week distance-learning “Basic Statistics” module and involved 46 higher education students who were randomly and blindly assigned to either a GenAI feedback group (n = 23) or a tutor feedback group (n = 23), with participants unaware of which type of feedback they were receiving during the intervention itself.

The researchers drew on Nazaretsky and colleagues’ (2024) four-dimension feedback perception framework, covering objectivity, usefulness, genuineness, and provider credibility, together with Barnard and colleagues’ (2009) six-dimension SRL model, comprising goal setting, task strategies, environment structuring, time management, help-seeking, and self-evaluation. SRL was measured using trace-based indicators adapted from Ye and Pennisi’s (2022) framework, derived from roughly 48,000 MOODLE log records that were mined, standardised, and mapped onto SRL proxy scores before and after two feedback interventions. GenAI feedback was produced using GPT-4 through structured prompts fed with each student’s individual proxy z-scores, while tutor feedback was generated using a purpose-built support tool named RefleXED, based on the same underlying data.

The results showed that students rated GenAI-generated feedback more favourably than tutor-generated feedback across every perception dimension, with a statistically significant advantage found specifically for Genuineness (p = .036, r = .36). In terms of SRL development, the GenAI feedback group demonstrated a significant improvement in the Task Strategies dimension (p = .044, r = .35), and a near-significant trend emerged for Time Management (p = .057, r = .33), while no significant between-group differences were found for the remaining SRL dimensions. Qualitative analysis of open-ended responses from 17 treatment-group students revealed considerable variation in awareness of the feedback source: students who recognised the feedback as AI-generated generally reported no change in attitude, prioritising content quality over provenance, whereas a smaller number of unaware students indicated their views might have shifted had they known the source in advance.

The findings suggest that carefully designed, learning-analytics-informed GenAI feedback holds real potential to scale personalised support for self-regulation in online higher education, but the researchers caution against viewing GenAI as a replacement for tutors. Instead, they argue for a complementary model in which GenAI’s scalability and adaptability work alongside tutors’ pedagogical judgement, while institutions remain attentive to how students’ awareness and perceptions of the feedback source can influence its ultimate impact. The authors note that the modest sample size (n = 46) drawn from a single course limits generalisability, and they call for future research involving larger, more diverse cohorts, no-feedback control conditions, and longer-term tracking of SRL outcomes.

Source (Open Access): Yilmaz, M., Temur, H. B., Emmungil, L., Çelik, E., Gauthier, A., & Cukurova, M. (2026). Supporting self-regulated learning through generative AI feedback in online higher education: the importance of student perceptions of the source of feedback. International Journal of Educational Technology in Higher Education, 23(1), 16.

https://doi.org/10.1186/s41239-026-00592-y… Read the rest

Categories
Higher Education Language Development

Task Offloading to GenAI in the Writing Feedback Process and Its Effects on Writing Development

Lai and colleagues used a quasi-experimental design to investigate whether offloading different tasks to GenAI in the “draft–assess–revise” workflow affects English-as-a-foreign-language (EFL) students’ writing development. Conducted over seven weeks, the study involved 101 Chinese university students and compared three conditions: (1) GenAI drafting, student assessment, student revision; (2) student drafting, GenAI assessment, student revision; and (3) student drafting, student self-assessment, student revision.

Results showed that the GenAI drafting group performed best in writing ability (M = 18.21, SD = 2.21), significantly higher than both the GenAI assessment group (M = 16.30, SD = 1.70) and the no GenAI group (M = 15.15, SD = 1.79). This indicates that including GenAI in the feedback workflow generally benefits writing, but offloading the drafting task rather than the assessment task produces the strongest effects.

Regarding cognitive processes, the GenAI drafting group showed significantly higher cognitive engagement, utilizing more total prompts (median 20.50 vs. 9.00, U = 195, p < .001) and more learning-oriented prompts (median 10.50 vs. 8.00, U = 398.50, p < .05) than the assessment group. They also outperformed the other two groups in metacognition in problem identification (M = 4.41), argumentation (M = 2.82), and providing constructive feedback (M = 4.97) during peer review.

In writing self-efficacy, the GenAI drafting group demonstrated the largest improvement (M = 4.74, SD = .55), significantly higher than the other groups, whereas the GenAI assessment group (M = 4.42, SD = .53) and no GenAI group (M = 4.45, SD = .78) did not differ significantly. Interviews revealed these students experienced a stronger sense of control and competence because they could critically evaluate and revise AI-generated drafts. In contrast, some students in the GenAI assessment group worried about overreliance on AI feedback, leading to “false confidence.”

Overall, the study highlights that the educational value of GenAI in writing feedback depends on which tasks are offloaded to AI. Offloading the drafting task to GenAI while having students handle assessment and revision promotes deeper cognitive engagement, stronger metacognitive understanding, and higher self-efficacy. The authors recommend that writing instruction should prioritize AI-generated drafts paired with student evaluation and revision, serving as a catalyst for critique and reflection rather than replacing student thinking.

Source (Open Access): Lai, C., Pan, M., Guo, K., & Cui, Y. (2026). What task to offload to GenAI in the writing feedback process? – Effects of task-offloading approaches on EFL learners’ writing skill development. Computers & Education, 252, 105675.

https://doi.org/10.1016/j.compedu.2026.105675… Read the rest

Categories
Effective Teaching Approach K-12 Education

The effect of AI-driven intelligent tutoring systems on K-12 students’ learning and performance: A systematic review

A recent systematic review published in npj Science of Learning examines the effects of intelligent tutoring systems (ITSs) on students’ learning and performance in K-12 education. As artificial intelligence in education (AIEd) has expanded rapidly, ITSs have emerged as a key application with the potential to personalize learning and improve educational outcomes. However, despite their growing adoption, their actual educational value remains uncertain. While some studies suggest that ITSs can enhance learning outcomes and even outperform traditional instruction, others report limited or inconsistent effects. In addition, existing research often conflates different educational contexts or focuses on broader AI applications, leaving a lack of systematic understanding of ITS effectiveness specifically in K-12 settings. This study therefore aims to assess the effects of ITSs on K-12 students’ learning and performance and to examine the experimental designs used to evaluate these systems.

The authors conducted a systematic review of 28 empirical studies involving a total of 4,597 students. Most studies adopted quasi-experimental designs, typically comparing an ITS-based intervention group with control conditions such as traditional teacher-led instruction, non-intelligent tutoring systems, modified ITSs, or no control group. The studies covered a range of countries, subjects, and school levels, with a strong concentration in middle and high school STEM education. Intervention durations varied considerably, from a single class session to several weeks or months. The review categorized studies based on educational context, experimental design, and intervention characteristics to enable a structured comparison of findings.

The review finds that ITSs generally have a positive effect on students’ learning and performance in K-12 education, particularly when compared to traditional teacher-led instruction, where most studies report medium to large effects. However, when compared with non-intelligent tutoring systems, the results are more mixed, with several studies finding no significant differences. Substantial heterogeneity is observed across studies due to differences in design, duration, and context. Importantly, the effectiveness of ITSs depends on key features such as personalization, adaptivity, and real-time feedback, as well as on implementation conditions. ITSs that are integrated with teacher support, encourage self-regulated learning, and are used over longer periods tend to produce better outcomes. In contrast, short interventions may be influenced by novelty effects, and learner characteristics such as prior knowledge and educational level also shape outcomes.

Taken together, the findings suggest that ITSs can enhance learning and performance in K-12 education, but their effectiveness is contingent upon pedagogical design and implementation conditions rather than technology alone. ITSs are most effective when aligned with sound instructional principles and used in combination with teacher guidance. The study also highlights limitations in the existing literature, including short intervention durations, limited sample diversity, and a lack of attention to ethical considerations. It calls for future research with more robust experimental designs, longer interventions, and greater attention to ethical issues, particularly as AI technologies continue to evolve and play an increasing role in education.

Source (Open Access): Létourneau, A., Deslandes Martineau, M., Charland, P., Karran, J. A., Boasen, J., & Léger, P. M. (2025). A systematic review of AI-driven intelligent tutoring systems (ITS) in K-12 education. npj Science of Learning, 10(1), 29.

https://doi.org/10.1038/s41539-025-00320-7… Read the rest

Categories
Higher Education

The Role of Undergraduates’ Critical Thinking in Generative AI Reliance Behaviors

Hou and colleagues conducted a large-scale survey study using structural equation modeling (SEM) to examine how undergraduates’ critical thinking influences their different types of reliance on generative AI during problem-solving tasks. The study analyzed 808 valid responses, measuring students’ critical thinking skills and dispositions, AI literacy, trust in AI, and four types of AI-use behaviors—reflective, cautious, collaborative, and thoughtless use. The authors conceptualized reliance behavior as the way learners evaluate and make use of the differences between AI and human abilities, and proposed that critical thinking may play a key moderating role in this process.

The results showed that AI literacy strongly predicted both critical thinking skills (β = .66, p < .001) and dispositions (β = .41, p < .001), whereas trust in AI was negatively related to both (skills: β = –.16, p < .05; disposition: β = –.11, p < .001). Regarding reliance behaviors, critical thinking skills were positively associated with collaborative use (β = .25), reflective use (β = .21), and cautious use (β = .24), with similar effects found for critical thinking disposition. These findings highlight the importance of critical thinking in supporting desirable forms of AI use. In contrast, trust strongly predicted thoughtless use (β = .47, p < .001) and also slightly increased collaborative use (β = .15, p < .05) and reflective use (β = .19, p < .001), indicating a dual role of trust in both strengthening and weakening ideal reliance behaviors. More importantly, AI literacy promoted collaborative (β = .25), reflective (β = .20), and cautious use (β = .22) through the mediation of critical thinking, whereas trust produced negative indirect effects on these desirable behaviors because it reduced critical thinking (β = –.05 to –.06, p < .001). This means that critical thinking both enhances the positive influence of AI literacy and suppresses the potential blind reliance brought by high trust, guiding learners toward more reflective, careful, and collaborative ways of using AI.

Overall, the study provides strong evidence that critical thinking does not simply reduce AI reliance; instead, it shapes how students rely on AI, encouraging forms of use that are more reflective, collaborative, and prudent. The authors argue that the development of AI literacy must be accompanied by the cultivation of critical thinking to reduce thoughtless dependence and to promote healthier human–AI collaboration. They also emphasize that educational interventions should clearly define “ideal reliance behaviors” and help students develop responsible and thoughtful habits of AI use in an era where generative AI is becoming increasingly widespread.

Source (Open Access): Hou, C., Zhu, G., & Sudarshan, V. (2025). The role of critical thinking on undergraduates’ reliance behaviours on generative AI in problem‐solving. British Journal of Educational Technology, 56(5), 1919-1941.

https://doi.org/10.1111/bjet.13613… Read the rest

Categories
Higher Education Social and Motivational Outcomes

Learners’ Preferences for Feedback from AI and Human Instructors

Le and his team examined whether learners’ preferences for feedback from human instructors versus generative artificial intelligence (AI) would change after receiving feedback from different sources and interface types in an academic English writing task. The study recruited 114 university students who were non-native English speakers and randomly assigned them to four groups: no feedback (control), human instructor feedback, ChatGPT 4.0 in a free-conversation interface, and a structured writing analysis tool powered by ChatGPT. Learners’ preferences were measured both before and after the task using rating scales and binary-choice questions, and the four groups were compared in terms of post-task preference and preference change.

The results showed that learners already had a clear preference for human instructors before the task (87.2% chose human), and this preference remained stable after the task (86.0% chose human), reflecting a phenomenon of algorithm aversion in educational settings. However, post-test preference scores differed significantly among the four groups: the human instructor group rated significantly higher than both the free-conversation AI group and the control group. On the binary human/AI choice measure, significant differences were also found — the human instructor and structured AI tool groups both scored higher than the free-conversation AI group. Regarding preference change, the overall mean shift was close to zero, but the differences among groups were significant: the free-conversation AI group showed a slight increase in preference for AI, whereas the human instructor and structured AI tool groups remained more favorable toward humans. In other words, although all three feedback types were effective, the free-conversation interface was the only one that reduced algorithm aversion and increased learners’ acceptance of AI, while the structured, one-time feedback tool further reinforced their preference for human instructors.

Based on these findings, the authors argue that enhancing the interactivity and dialogic nature of AI-based learning tools may influence learners’ preferences more effectively than purely improving their technical performance. Interactive dialogue allows for clarification and correction, which reduces learners’ unrealistic expectations that algorithms must be perfect and mitigates distrust. Overall, the study situates human preference within the context of interface design, providing both empirical insights and cautions for the adoption, product design, and pedagogical integration of AI in education.

 

Source (Open Access): Le, H., Shen, Y., Li, Z., Xia, M., Tang, L., Li, X., … & Fan, Y. (2025). Breaking human dominance: Investigating learners’ preferences for learning feedback from generative AI and human tutors. British Journal of Educational Technology.

https://doi.org/10.1111/bjet.13614… Read the rest

Categories
Educational Administration and Leadership Effective Teaching Approach Secondary School Education

Leveraging AI to predict young learners’ online learning engagement

With many schools rushing to adopt Generative AI, it is important to consider the real learning gains (or lack thereof) that these tools offer. A 2023 study by Pardos & Bhandari examined the use of AI-generated hints as a scaffolding mechanism with Algebra students.

Seventy-seven participants (high school graduates selected via Amazon’s MTURK system) were assigned to a control group (which provided human-generated hints) or an experimental group (which provided AI-generated hints). The researchers wanted to learn the rate of “low quality” AI-generated hints, as well as if the hints produced learning gains compared to the control group. The questions from the lesson were fed, verbatim, to ChatGPT in order to generate the hints. Quality checks were performed manually to ensure that all AI-generated hints were correct and showed the proper steps. This was then contrasted with the control group, whose hints were generated by undergraduate tutors. Pre and post tests were administered to check for learning gains between the two groups.

The results showed that 70% of the hints generated by ChatGPT were considered to be good quality, and that there was a statistically significant learning gain in the control group. A major limitation of the study is that the researchers did not prompt the AI to use any scaffolding strategies. Therefore, the quality of the hints between groups not only differed by human or AI creator, but also by pedagogical theory. Human tutors were probably more likely to employ Vygotsky-esque scaffolds, while ChatGPT was more likely to provide an immediate answer. Future work could improve upon the prompts used in this study and create a multi-tiered approach with less consequential hints being revealed at first.

 

Source (Open Access): Pardos, Z. A., & Bhandari, S. (2023). Learning gain differences between ChatGPT and human tutor generated algebra hints (No. arXiv:2302.06871). https://doi.org/10.48550/arXiv.2302.06871… Read the rest