Categories
Achievement Higher Education Primary School Education Secondary School Education

The impact of test preparation on performance in large-scale educational tests: A meta-analysis of experimental studies

A recent meta-analysis examines whether test preparation improves performance in large-scale educational tests. Although schools, commercial institutions, and students invest heavily in preparation courses and materials, existing studies have reported inconsistent effects. Earlier reviews were also concentrated on US college admission and cognitive ability tests and often relied on evidence published before 2000. This study therefore estimates the overall effect of test preparation and examines how it varies across intervention, test, student, and research-design characteristics.

The authors synthesized 28 experimental and quasi-experimental studies, including 92 effect sizes and 16,741 participants across primary, secondary, and tertiary education. The studies covered college admission tests, English proficiency tests, and other large-scale assessments. A hierarchical random-effects model with robust variance estimation was used, while moderator analyses examined intervention breadth, strategy and material use, test characteristics, previous test experience, and study design.

Test preparation produced a significant positive overall effect on test performance (g=.26, 95% CI [.10,.42], p<.001), but heterogeneity was high (I^2=82.44%). Test-specific and narrowly focused interventions were more effective (g=.33) and (g=.32) than broad programs aimed at general knowledge or transferable skills (g=-.04). Interventions teaching test-taking strategies also showed stronger effects (g=.35) than those without such instruction (g=.05). Students with previous testing experience benefited substantially more (g=.69) than those without it (g=.06). However, preparation effects did not transfer significantly to other tests in the same domain (g=.06, p=.502).

Taken together, the findings suggest that test preparation can improve large-scale test scores, but its effectiveness mainly reflects alignment with the target test rather than broader learning gains. The limited transfer effect raises concerns about whether score improvement represents genuine development in the abilities being assessed. Future research should use stronger experimental designs, report intervention processes more clearly, and distinguish test-score gains from transferable learning.

Source (Open Access): Hao, Z., Baird, J. A., Masri, Y. E., & Double, K. (2025). The impact of test preparation on performance of large-scale educational tests: A meta-analysis of experimental studies. Review of Educational Research, 00346543251360775.

https://doi.org/10.3102/00346543251360775Read the rest

Categories
Educational Administration and Leadership Higher Education K-12 Education

The impact of generative artificial intelligence on students’ higher order thinking: Evidence from a three-level meta-analysis

Given the current trends in AIEd, it is crucial to synthesize the overall impact of Gen-AI on HOT, as Gen-AI has been creatively applied by researchers to assist in teaching and learning. Concerns exist in current discourse that Gen-AI may harm HOT. However, empirical research on the specific effects of Gen-AI on HOT remains scattered, and no consensus has been reached. Nie et al.’s (2025) study synthesized 19 experimental and quasi-experimental studies conducted in different contexts (k=68), involving 2,347 participants. A three-level meta-analysis was employed to account for within- and between-study variability, assessing the impact of Gen-AI on HOT and exploring the effects of moderators.

The results revealed that Gen-AI had a significant positive effect on enhancing students’ HOT (Hedges’s g=0.851, p<0.001). Moderator analyses were conducted based on impact target, Gen-AI elements, study contexts, and methodological characteristics. When the sample size was less than 80, the promotion effect was more significant. Both short-term interventions (less than 4 weeks) and long-term interventions (more than 8 weeks) have the potential to produce significant positive effects.

These findings offer vital policy implications for integrating Gen-AI into education. First, educational authorities should mandate smaller class sizes or optimized student-to-teacher ratios in AI-driven curricula, as large groups dilute Gen-AI’s efficacy in fostering HOT. Second, curriculum frameworks must transition from short-term pilots to extended, long-term intervention timelines, given that 4-8 weeks is insufficient for deep cognitive development. Lastly, policymakers should fund targeted, granular research into AI’s specific impacts on different dimensions of HOT to guide the scaling of intelligent, personalized, and adaptive educational ecosystems.

Source (Open Access): Nie, X., Tian, Y., Liu, M., Wu, D., & Guo, Y. (2025). The impact of generative artificial intelligence on students’ higher order thinking: Evidence from a three-level meta-analysis. Education and Information Technologies30(17), 25359-25390.https://doi.org/10.1007/s10639-025-13735-xRead the rest

Categories
Higher Education Language Development

Task Offloading to GenAI in the Writing Feedback Process and Its Effects on Writing Development

Lai and colleagues used a quasi-experimental design to investigate whether offloading different tasks to GenAI in the “draft–assess–revise” workflow affects English-as-a-foreign-language (EFL) students’ writing development. Conducted over seven weeks, the study involved 101 Chinese university students and compared three conditions: (1) GenAI drafting, student assessment, student revision; (2) student drafting, GenAI assessment, student revision; and (3) student drafting, student self-assessment, student revision.

Results showed that the GenAI drafting group performed best in writing ability (M = 18.21, SD = 2.21), significantly higher than both the GenAI assessment group (M = 16.30, SD = 1.70) and the no GenAI group (M = 15.15, SD = 1.79). This indicates that including GenAI in the feedback workflow generally benefits writing, but offloading the drafting task rather than the assessment task produces the strongest effects.

Regarding cognitive processes, the GenAI drafting group showed significantly higher cognitive engagement, utilizing more total prompts (median 20.50 vs. 9.00, U = 195, p < .001) and more learning-oriented prompts (median 10.50 vs. 8.00, U = 398.50, p < .05) than the assessment group. They also outperformed the other two groups in metacognition in problem identification (M = 4.41), argumentation (M = 2.82), and providing constructive feedback (M = 4.97) during peer review.

In writing self-efficacy, the GenAI drafting group demonstrated the largest improvement (M = 4.74, SD = .55), significantly higher than the other groups, whereas the GenAI assessment group (M = 4.42, SD = .53) and no GenAI group (M = 4.45, SD = .78) did not differ significantly. Interviews revealed these students experienced a stronger sense of control and competence because they could critically evaluate and revise AI-generated drafts. In contrast, some students in the GenAI assessment group worried about overreliance on AI feedback, leading to “false confidence.”

Overall, the study highlights that the educational value of GenAI in writing feedback depends on which tasks are offloaded to AI. Offloading the drafting task to GenAI while having students handle assessment and revision promotes deeper cognitive engagement, stronger metacognitive understanding, and higher self-efficacy. The authors recommend that writing instruction should prioritize AI-generated drafts paired with student evaluation and revision, serving as a catalyst for critique and reflection rather than replacing student thinking.

Source (Open Access): Lai, C., Pan, M., Guo, K., & Cui, Y. (2026). What task to offload to GenAI in the writing feedback process? – Effects of task-offloading approaches on EFL learners’ writing skill development. Computers & Education252, 105675.

https://doi.org/10.1016/j.compedu.2026.105675Read the rest

Categories
Achievement Higher Education

Student agency in Colombian higher education: a dual-pathway model of mediation and moderation between student engagement and academic achievement

A recent quantitative study by Torres Castro and Pineda-Baéz examines the relationship between student engagement and academic achievement, with particular attention to how student agency functions as both a mediating and a moderating mechanism in this relationship.

The study surveyed 1,713 final-year students from six accredited private universities across five regions in Colombia. Grounded in social cognitive theory and the ecological model of student engagement, the study conceptualizes student engagement as a multidimensional construct. It encompasses ten indicators across four domains: academic challenge, learning with peers, experiences with faculty, and campus environment. Student engagement is assessed through the National Survey of Student Engagement, with scores standardized on a 0 to 60 scale. Student agency is operationalized as agentic engagement through the five-item Agentic Engagement Scale. This scale captures proactive behaviors such as expressing preferences, asking questions, and suggesting adjustments to instruction. Academic achievement was measured by self-reported cumulative grade point average on Colombia’s standardized 0.0 to 5.0 scale. The study employed hierarchical regression with robust standard errors, structural equation modeling, and bootstrap-based mediation analysis using R version 4.3.1.

The findings reveal that among the ten engagement indicators, only collaborative learning and student-faculty interaction are positively and independently associated with academic achievement. Among the five agentic engagement behaviors, only expressive voice emerged as a significant mediator, channeling approximately six percent of the total effect of collaborative learning on grade point average. Expressive voice also moderated the relationship between student-faculty interaction and grade point average in a compensatory pattern. The association between faculty interaction and achievement was stronger among students with lower levels of expressive voice, whereas it was attenuated yet remained positive among those with higher levels. Although effect sizes were modest, the findings demonstrate that student agency operates as a context-dependent dual mechanism.

The study further suggests that instructional practices should combine structured guidance with opportunities for student-initiated contribution, and that support systems should prioritize students with lower levels of agentic engagement.

Source (Open Access): Torres Castro, U. E., & Pineda-Baéz, C. (2026). Student agency in Colombian higher education: a dual-pathway model of mediation and moderation between student engagement and academic achievement. Higher Education, 1-20.

https://doi.org/10.1007/s10734-026-01669-3Read the rest

Categories
Higher Education Language Development

Comparing the effects of ChatGPT and automated writing evaluation on students’ writing and ideal L2 writing self

Using a randomized controlled experimental design, Shi et al. (2025) compared the effects of ChatGPT-based feedback and traditional automated writing evaluation (AWE) systems on English-as-a-foreign-language (EFL) students’ writing performance and their ideal L2 writing self. One hundred and fifty second-year university students from three writing classes in a Chinese public university were recruited and randomly divided into a ChatGPT group, an AWE group, and a control group.

After an eleven-week intervention, results showed that ChatGPT helped students perform better in their writing compared to the control group and the AWE group, but compared to the AWE group, ChatGPT significantly lowered students’ ideal L2 writing self. Qualitative results shed light on possible causes: while participants were fully aware of the affordances of ChatGPT feedback, they were also concerned with their (over) reliance on the tool and the accompanying loss of creativity and agency and expressed their reserved attitude toward future intention to use ChatGPT.

Educators should refine learning objectives based on students’ ZPD and design prompts accordingly, so that ChatGPT supports learning rather than completing tasks, while also teaching prompt-engineering skills. For lower-intermediate to intermediate learners, AWE’s systematic and rule-based feedback can provide stronger scaffolding and better preserve authorship. However, ChatGPT’s richer affordances may lead to over-reliance, weakening learner agency and diminishing the ideal L2 writing self. Therefore, language-education goals should be redefined to incorporate AI literacy and critical thinking, safeguarding teacher and learner agency and promoting responsible use.

 

Source (Open Access): Shi, H., Chai, C. S., Zhou, S., & Aubrey, S. (2025). Comparing the effects of ChatGPT and automated writing evaluation on students’ writing and ideal L2 writing self. Computer Assisted Language Learning, 1-28.

https://doi.org/10.1080/09588221.2025.2454541Read the rest

Categories
Higher Education

The Role of Undergraduates’ Critical Thinking in Generative AI Reliance Behaviors

Hou and colleagues conducted a large-scale survey study using structural equation modeling (SEM) to examine how undergraduates’ critical thinking influences their different types of reliance on generative AI during problem-solving tasks. The study analyzed 808 valid responses, measuring students’ critical thinking skills and dispositions, AI literacy, trust in AI, and four types of AI-use behaviors—reflective, cautious, collaborative, and thoughtless use. The authors conceptualized reliance behavior as the way learners evaluate and make use of the differences between AI and human abilities, and proposed that critical thinking may play a key moderating role in this process.

The results showed that AI literacy strongly predicted both critical thinking skills (β = .66, p < .001) and dispositions (β = .41, p < .001), whereas trust in AI was negatively related to both (skills: β = –.16, p < .05; disposition: β = –.11, p < .001). Regarding reliance behaviors, critical thinking skills were positively associated with collaborative use (β = .25), reflective use (β = .21), and cautious use (β = .24), with similar effects found for critical thinking disposition. These findings highlight the importance of critical thinking in supporting desirable forms of AI use. In contrast, trust strongly predicted thoughtless use (β = .47, p < .001) and also slightly increased collaborative use (β = .15, p < .05) and reflective use (β = .19, p < .001), indicating a dual role of trust in both strengthening and weakening ideal reliance behaviors. More importantly, AI literacy promoted collaborative (β = .25), reflective (β = .20), and cautious use (β = .22) through the mediation of critical thinking, whereas trust produced negative indirect effects on these desirable behaviors because it reduced critical thinking (β = –.05 to –.06, p < .001). This means that critical thinking both enhances the positive influence of AI literacy and suppresses the potential blind reliance brought by high trust, guiding learners toward more reflective, careful, and collaborative ways of using AI.

Overall, the study provides strong evidence that critical thinking does not simply reduce AI reliance; instead, it shapes how students rely on AI, encouraging forms of use that are more reflective, collaborative, and prudent. The authors argue that the development of AI literacy must be accompanied by the cultivation of critical thinking to reduce thoughtless dependence and to promote healthier human–AI collaboration. They also emphasize that educational interventions should clearly define “ideal reliance behaviors” and help students develop responsible and thoughtful habits of AI use in an era where generative AI is becoming increasingly widespread.

Source (Open Access): Hou, C., Zhu, G., & Sudarshan, V. (2025). The role of critical thinking on undergraduates’ reliance behaviours on generative AI in problem‐solving. British Journal of Educational Technology56(5), 1919-1941.

https://doi.org/10.1111/bjet.13613Read the rest