Technology-Enhanced Assessment in Bahasa Melayu Education: Designing Adaptive Digital Tools to Improve Language Proficiency Outcomes ()
1. Introduction
Bahasa Melayu occupies a central position in the Malaysian education system as the national language, the main medium of instruction in national schools, and a compulsory subject that all students must pass in high-stakes national examinations. The language also carries considerable cultural and civic weight, functioning as a vehicle of national identity and social cohesion in a multilingual society (Ministry of Education Malaysia, 2013). Ensuring that all learners achieve strong proficiency in Bahasa Melayu is therefore not only an educational objective but also a matter of national policy. Yet the assessment practices meant to support that objective have evolved far more slowly than the curriculum they serve. Classroom assessment of Bahasa Melayu continues to rely heavily on uniform written tests, standardized worksheets, and summative examinations that treat all learners as though they occupied the same point on the proficiency continuum (Arumugham, 2020; Said et al., 2021). Such instruments generate scores, but they rarely generate the fine-grained, timely, and actionable information that teachers and learners need in order to close specific proficiency gaps.
The limitations of static assessment are well documented. Traditional fixed-form tests deliver the same items to every candidate regardless of ability, which means that weaker learners face many items too difficult to be informative, while stronger learners answer many items too easy to be challenging (Wainer, 2000). Feedback, when it arrives, is often delayed, grade-focused, and detached from the specific linguistic behaviors that produced the errors, which sharply reduces its formative value (Hattie & Timperley, 2007; Shute, 2008). At the same time, decades of research have established that assessment, when used formatively, is among the most powerful levers available for improving learning outcomes (Black & Wiliam, 1998; Wiliam, 2011). The challenge is therefore not whether to assess but how to redesign assessment so that it functions as an engine of learning rather than merely a mechanism of certification (Stobart, 2008).
Digital technology has transformed what is possible in this regard. Technology-enhanced assessment can present multimodal stimuli, capture rich response data, score constructed responses automatically, adapt task difficulty in real time, and return immediate, individualized feedback (Redecker & Johannessen, 2013; Timmis et al., 2016; Madland et al., 2024). Computerized adaptive testing tailors item selection to each learner’s estimated ability, producing more precise measurement with fewer items (Weiss & Kingsbury, 1984; van der Linden & Glas, 2010). More recently, advances in artificial intelligence, learning analytics, and natural language processing have enabled systems that model individual learners over time, diagnose specific misconceptions, and personalize both content and feedback (Hwang et al., 2020; Ouyang & Jiao, 2021; Chiu et al., 2023; Yang et al., 2022). In language education specifically, automated writing evaluation, speech recognition, and intelligent tutoring have matured into practical classroom tools for major world languages (Chapelle & Voss, 2016; Li et al., 2015; Ranalli, 2018).
The benefits of these developments, however, have accrued unevenly across languages. The overwhelming majority of research on adaptive and intelligent language assessment concerns English and a small number of other high-resource languages (Golonka et al., 2014; Godwin-Jones, 2021). Bahasa Melayu, despite being spoken across the Malay world and holding official status in Malaysia, Brunei, and Singapore, remains strikingly underserved by assessment technology. Published research on digital assessment tools designed specifically for Bahasa Melayu proficiency is scarce, and the tools that do circulate in Malaysian classrooms tend to be generic quiz platforms that digitize traditional item formats without exploiting adaptivity, diagnosis, or intelligent feedback (Yunus et al., 2013; Maxwell-Smith et al., 2022). There is, in short, a substantial gap between what assessment technology can now do and what Bahasa Melayu education currently does with it.
This conceptual paper addresses that gap. Its purpose is to synthesize the relevant theoretical and empirical literature and, on that basis, to propose a coherent framework for designing adaptive digital assessment tools for Bahasa Melayu education. Three questions guide the discussion. First, what does existing scholarship on formative assessment, adaptive testing, and language assessment technology imply for the design of technology-enhanced assessment in Bahasa Melayu? Second, what theoretical foundations should underpin such tools so that they improve, rather than merely measure, language proficiency? Third, what design principles and system architecture can translate these foundations into a practical blueprint for developers, teachers, and policymakers? As a conceptual contribution, the paper does not report new empirical data; it follows instead the tradition of integrative scholarship that builds frameworks by synthesizing and extending existing bodies of knowledge.
The framework proposed here is deliberately scoped, because the standards it must anchor to, the task types it can use, and the degree of learner agency that is developmentally appropriate all vary across school populations. ATEA-BM is designed in the first instance for Bahasa Melayu in Malaysian national and national-type schools at Year 4 to Year 6 of primary schooling and Form 1 to Form 3 of lower secondary schooling, that is, for learners aged approximately ten to fifteen. This range is chosen for three reasons. Learners at these levels have generally consolidated basic decoding and handwriting, so that adaptive tasks can address the full span of listening, speaking, reading, writing, grammar, and vocabulary rather than early literacy alone; the DSKP content and learning standards for these years are differentiated finely enough to support subskill diagnosis; and classroom-based assessment obligations at these levels are continuous rather than dominated by terminal examination preparation, so that a formative system has room to operate (Kementerian Pendidikan Malaysia, 2018; Ministry of Education Malaysia, 2013). Three learner profiles are assumed within this range and are treated throughout as design parameters rather than as edge cases: pupils for whom Bahasa Melayu is a first language; pupils for whom it is a second or additional language, including those whose dominant home language is Mandarin, Tamil, or a Bornean or other regional language; and pupils with persistent low attainment who require sustained work at standards below their nominal year level. These profiles differ in the lexical and listening support they need, in the pacing of tasks, and in the language and elaboration of feedback, and Section 4 specifies how the framework accommodates those differences. Early primary (Year 1 to Year 3), upper secondary examination preparation, and adult or foreign-language learners of Malay lie outside the present scope; the design principles may well transfer, but the task inventory, the standards to which tasks must be mapped, and the breadth of learner choice would each require re-specification, and that work is not attempted here.
The paper proceeds as follows. Section 2 reviews the literature on Bahasa Melayu assessment, technology-enhanced assessment, adaptive testing, and artificial intelligence in language assessment, beginning with an account of how that literature was identified, selected, and synthesized. Section 3 sets out the theoretical foundations of the proposed framework. Section 4 presents the framework itself: its design principles and system architecture, the boundary it draws between formative adaptive practice and adaptive measurement, the evidence required before a subskill may be diagnosed, the safeguards governing automated analysis of writing and speech, and the way learner choice is bounded by curriculum entitlement. Section 5 discusses implications for practice, teacher education, and policy. Section 6 considers challenges and limitations, Section 7 outlines a research agenda, and Section 8 concludes.
2. Literature Review
Because this paper builds a framework rather than estimates an effect, the literature reviewed below was assembled through an integrative review, a form of synthesis suited to bodies of work that are theoretically and methodologically diverse and whose purpose is to generate a conceptual model (Torraco, 2005; Snyder, 2019). The procedure was as follows. Searches were run in Scopus, Web of Science, ERIC, and Google Scholar, supplemented by targeted searching of the Malaysian databases MyJurnal and MyCite and of Ministry of Education curriculum and policy documents, using combinations of the terms formative assessment, classroom-based assessment, pentaksiran bilik darjah, technology-enhanced assessment, computerized adaptive testing, adaptive learning, automated writing evaluation, automated speech scoring, artificial intelligence in education, language assessment, Bahasa Melayu, Malay language, and pentaksiran bahasa, in both English and Malay. Sources were eligible if they were peer-reviewed empirical studies, systematic or narrative reviews, foundational theoretical works, or official curriculum and policy documents, and if they addressed at least one of four strands: the theory and effects of formative assessment and feedback; the design and psychometrics of adaptive testing; artificial intelligence and the automated analysis of learner language; and assessment policy or practice in Malaysian and Bahasa Melayu settings. Technology-specific empirical work was concentrated on the period from 2010 onwards, since claims about what systems can now do date quickly, while foundational theoretical and psychometric works were retained irrespective of date, because the anchors of the field remain earlier contributions such as Vygotsky (1978), Sadler (1989), Black and Wiliam (1998), and Wainer (2000). Reference lists of the reviews located were searched backwards to capture work the databases missed. Synthesis was thematic rather than statistical: sources were coded against the four strands, points of convergence and tension across strands were identified, and the design principles set out in Section 4 were derived from those points and then checked back against the sources that generated them. Because selection and coding were carried out by the authors without independent double screening, the review makes no claim to the exhaustiveness of a systematic review, and it is reported here so that the basis of the framework can be inspected and contested.
The claim that Bahasa Melayu is underserved by assessment technology warrants particular scrutiny, since an absence is harder to evidence than a presence, and it rests here on three converging sources rather than on an empty result from a single search. First, the scoping study by Maxwell-Smith et al. (2022) of 657 publications in Indonesian and Malay language technology found the field dominated by corpus construction, machine reading of web-scraped text, and sentiment analysis, with education applications almost entirely absent. Second, the searches described above returned no peer-reviewed study reporting an operational adaptive or automatically scored assessment system for Bahasa Melayu proficiency in school settings; the Malaysian work located addressed instead the design of digital environments for teaching and engagement (Jie & Kamrozzaman, 2025), teacher readiness for assessing writing (Said et al., 2021), and the implementation of classroom-based assessment (Arumugham, 2020; Hock et al., 2022). Third, reviews of language learning technology and of learner data resources report the same asymmetry structurally, with research effort concentrated on English and a small group of high-resource languages (Golonka et al., 2014; Godwin-Jones, 2021). The claim should therefore be read as a claim about the accessible published record, not as an assertion that no such tool exists anywhere in commercial or unpublished form; it is nonetheless sufficient to establish that design for Bahasa Melayu cannot proceed by importing a validated model from the English-language literature.
2.1. Assessment in Bahasa Melayu Education: The Current Landscape
The Bahasa Melayu curriculum in Malaysian schools is organized around the integrated development of four language skills, namely listening (kemahiran mendengar), speaking (kemahiran bertutur), reading (kemahiran membaca), and writing (kemahiran menulis), together with systematic mastery of grammar (tatabahasa) and vocabulary as codified in authoritative references such as Tatabahasa Dewan (Nik Safiah Karim et al., 2015). The Standard Curriculum and Assessment Documents (Dokumen Standard Kurikulum dan Pentaksiran, DSKP) specify, for each skill and year level, content standards, learning standards, and six-level performance standards (tahap penguasaan) describing progression from basic recognition to sophisticated, autonomous language use (Kementerian Pendidikan Malaysia, 2018). In principle, this architecture provides exactly the kind of fine-grained proficiency map on which modern diagnostic assessment can be built.
In practice, implementation has struggled to match the ambition of the curriculum. The Malaysia Education Blueprint 2013-2025 called for a decisive shift away from examination-oriented schooling towards school-based and classroom-based assessment supporting higher-order thinking (Ministry of Education Malaysia, 2013). The subsequent institutionalization of classroom-based assessment (Pentaksiran Bilik Darjah, PBD) formally requires teachers to assess continuously, interpret pupil performance against the tahap penguasaan, and use the results to adjust teaching. Yet studies of implementation report persistent difficulties: teachers experience heavy recording workloads, hold uneven understandings of the performance standards, and frequently default to familiar test-and-grade routines rather than genuinely formative practice (Arumugham, 2020; Hock et al., 2022). Evidence specific to Bahasa Melayu points in the same direction. In a national survey of 182 Malay language teachers across seven zones, Said et al. (2021) found that although teachers rated their own mastery of writing assessment highly, their reported practice clustered around procedural implementation rather than diagnostic interpretation, a pattern consistent with the wider finding that the rhetoric of assessment for learning coexists with classroom realities that remain substantially summative in character (Bennett, 2011).
For Bahasa Melayu specifically, this situation has several consequences. Diagnosis is coarse: a pupil may be recorded at a given tahap penguasaan for writing without any systematic record of whether the underlying difficulty lies in sentence construction (pembinaan ayat), affixation (imbuhan), spelling (ejaan), discourse organization, or idea development. Feedback is slow and generic, typically arriving days after performance in the form of marks and brief written comments. Practice is undifferentiated, since a single worksheet or essay title is normally assigned to an entire class regardless of the wide proficiency range within it, including the needs of pupils for whom Bahasa Melayu is a second or additional language. Each of these problems is precisely the kind that technology-enhanced and adaptive assessment is designed to address.
2.2. Technology-Enhanced Assessment: Concepts and Evolution
Technology-enhanced assessment refers to the use of digital technologies to design, deliver, score, and report assessment and, more ambitiously, to transform the pedagogical function of assessment itself (Timmis et al., 2016). Scholars commonly distinguish a migratory strategy, in which existing paper instruments are simply transferred to screen, from a transformative strategy, in which technology enables assessment forms that were previously impossible, such as simulation-based tasks, continuous embedded assessment, and instant adaptive feedback (Redecker & Johannessen, 2013; Pellegrino & Quellmalz, 2010). The literature consistently finds that the educational value of digitization lies almost entirely in the transformative strategy: putting a static multiple-choice test on a tablet changes little, whereas redesigning assessment around adaptivity, interactivity, and analytics can change the learning process itself (Shute & Rahimi, 2017).
Recent scholarship has sharpened this distinction. Bearman et al. (2023) argue that the digital can no longer be treated as an optional add-on to an otherwise stable assessment design, because assessment now takes place within a digitally mediated society; their organizing framework distinguishes designing the digital in as a tool for improvement, as a means of developing digital literacies, and as a means of credentialing uniquely human capabilities. Madland et al. (2024) similarly observe that frameworks for technology-integrated assessment have tended to outpace the empirical literature, which remains concentrated on tool adoption rather than on assessment design. Both analyses support a position taken throughout this paper: the question for Bahasa Melayu is not which platform to purchase, but what assessment argument the technology is being asked to serve.
A substantial body of evidence links well-designed technology-mediated formative assessment to improved engagement and achievement. Gikandi et al. (2011) concluded from a systematic review that online formative assessment can foster meaningful interaction, timely feedback, and learner autonomy when it is embedded in a coherent pedagogical design rather than bolted on as an afterthought. Reviews of computer-based assessment for learning in school settings similarly report positive effects on achievement and motivation, while cautioning that effects depend on feedback quality and implementation fidelity (Shute & Rahimi, 2017). Meta-analytic work on mobile technology integration finds moderate positive effects on learning performance across subjects, which matters in the Malaysian context, where smartphones and tablets are frequently the most accessible devices for pupils (Sung et al., 2016; Kukulska-Hulme & Shield, 2008).
Within language education, technology-enhanced assessment has a long pedigree. Computer-assisted language testing has developed from early item banks to sophisticated systems capable of assessing all four skills, and language testing researchers have articulated validation frameworks specific to computer-delivered language assessment (Chapelle, 2001; Chapelle & Douglas, 2006). Two decades of work in this field show a clear trajectory from efficiency-driven applications towards assessment designs that exploit the computer’s capacity for interactivity, multimedia, and automated analysis of learner language (Chapelle & Voss, 2016; Kessler, 2018). Malaysian research has begun to follow this trajectory: Jie and Kamrozzaman (2025) report on the design of an immersive environment for the National Language course in which listening, speaking, reading, and writing are practiced through simulated authentic situations, illustrating both the pedagogical appeal of technology-mediated Bahasa Melayu tasks and the current absence of an assessment layer capable of turning such activity into diagnostic evidence.
2.3. Adaptive Assessment Technologies
The core technology of adaptive assessment is computerized adaptive testing (CAT). In a CAT, the system estimates a candidate’s ability after each response and selects the next item to be maximally informative at that estimated ability level, so that every candidate receives a test tailored to their own proficiency (Weiss & Kingsbury, 1984; Wainer, 2000). The psychometric machinery underlying CAT is item response theory, which models the probability of a correct response as a function of latent ability and item parameters such as difficulty and discrimination (Embretson & Reise, 2000). The practical advantages are considerable: adaptive tests achieve target measurement precision with substantially fewer items than fixed forms, reduce testing time, minimize frustration for low-proficiency learners and boredom for high-proficiency learners, and yield ability estimates on a common scale that supports the tracking of growth over time (van der Linden & Glas, 2010).
Adaptivity, however, need not be confined to item selection for measurement efficiency. Evidence-centered design reconceptualizes assessment as an evidentiary argument connecting observable performances to claims about learner competencies through explicit student, evidence, and task models (Mislevy et al., 2003). Within this broader view, adaptive systems can branch not only on estimated ability but also on diagnosed error patterns, selected scaffolds, and learner choices, thereby merging assessment with instruction. Research on personalized e-learning demonstrates that systems combining ability estimation with adaptive learning-path guidance can improve both learning efficiency and outcomes (Chen, 2008). A particularly instructive recent example is the adaptive formative assessment system reported by Yang et al. (2022), which couples computerized adaptive testing with a learning memory cycle so that item selection serves review scheduling as well as measurement; the system is a working demonstration that adaptivity can be designed to teach rather than only to test. Contemporary artificial intelligence in education extends this logic further through learner models that accumulate evidence across time and tasks, enabling prediction, diagnosis, and recommendation at a granularity that classroom teachers cannot feasibly achieve manually (Hwang et al., 2020; Luckin & Cukurova, 2019; Zawacki-Richter et al., 2019). Ley et al. (2023) argue that the most productive configuration of such systems is an explicit partnership in which the machine’s learner model and the teacher’s professional model of the class inform one another, rather than one silently substituting for the other.
The paradigm within which such systems are designed matters. Ouyang and Jiao (2021) distinguish three paradigms of artificial intelligence in education: AI-directed systems that treat the learner as a recipient, AI-supported systems that treat the learner as a collaborator, and AI-empowered systems that treat the learner as a leader of their own learning. Reviews of the field caution that much current work remains technology-centered and weakly theorized pedagogically (Zawacki-Richter et al., 2019; Chiu et al., 2023; Crompton & Burke, 2023). Hopfenbeck et al. (2023) sharpen the warning for the specific case of classroom formative assessment, noting that artificial intelligence may either strengthen teachers’ feedback and peer-assessment practices or quietly displace them, depending on how the tools are positioned relative to teacher judgment. For school language education, the design implication is clear: adaptive assessment should be conceived within the AI-supported and AI-empowered paradigms, in which the system augments teacher judgment and learner agency rather than replacing them.
2.4. Artificial Intelligence and Automated Analysis in Language Assessment
Automated scoring of learner language is the technology that most directly enables adaptive assessment of productive skills. Automated essay scoring systems extract linguistic features from written responses, or learn them from data, and produce scores that in many settings agree with human raters at levels comparable to inter-rater agreement among humans (Attali & Burstein, 2006; Shermis, 2014). Beyond scoring, automated writing evaluation systems provide immediate diagnostic feedback on grammar, mechanics, vocabulary, and organization, and classroom research shows that such feedback can support revision behavior and writing development when it is integrated into instruction and mediated by teachers (Warschauer & Grimes, 2008; Li et al., 2015). The arrival of generative language models has widened the range of feedback that can be produced automatically, but Huang et al. (2025) find in their review of this literature that although such models can approximate human scoring and generate richer rubric-aligned commentary, questions of fairness, validity, and ethical governance remain largely unresolved. Studies of learner uptake reinforce the point that generation is not the same as use: students do not automatically understand or act on automated feedback, which underscores the need for feedback literacy and careful message design (Ranalli, 2018). Language testing scholars have likewise warned that automated scoring changes the validity argument of an assessment and must be evaluated against the constructs it claims to measure, not merely against human-machine score agreement (Xi, 2010).
Speech technology plays the corresponding role for oral proficiency. Automatic speech recognition and pronunciation scoring now underpin speaking practice and assessment in widely used language learning platforms, and case-study evidence indicates that learners can make measurable proficiency gains through sustained use of such mobile applications (Loewen et al., 2019). Sun (2023) reports measurable gains in both pronunciation and speaking from sustained use of speech recognition feedback, although Rogerson-Revell (2021) cautions that computer-assisted pronunciation training still struggles with the accented and multilingual speech typical of classrooms outside the inner circle, which is precisely the condition of Malaysian schools. Reviews of technologies for language learning conclude that effectiveness evidence is strongest where the technology delivers individualized practice with immediate feedback, which is exactly the profile of adaptive assessment (Golonka et al., 2014).
For Bahasa Melayu, the application of these technologies is constrained by the resource gap between English and most other languages. Malay is comparatively under-resourced in annotated learner corpora, validated automated scoring models, and speech systems trained on the range of accents and sociolects found in Malaysian classrooms (Godwin-Jones, 2021). Maxwell-Smith et al. (2022) quantify the imbalance directly: their scoping study of 657 publications on Indonesian and Malay language technology found the field dominated by exploratory corpus work, machine reading of web-scraped text, and sentiment analysis, with education applications almost entirely absent. The morphology of Malay, in particular its rich and pedagogically salient system of affixation, offers both a challenge and an opportunity: a challenge because scoring models developed for English do not transfer directly, and an opportunity because affixation errors are systematic, curriculum-mapped, and therefore highly amenable to rule-informed automated diagnosis. The scarcity of published work on adaptive assessment for Bahasa Melayu is itself a central justification for the framework developed here: design must be theorized before large-scale development and validation can proceed in a principled way.
3. Theoretical Foundations
A framework for adaptive Bahasa Melayu assessment must rest on theory that explains how assessment produces learning, how learners develop with support, how they come to direct their own learning, and how teachers integrate technology with pedagogy and content. Four bodies of theory jointly provide this foundation.
3.1. Formative Assessment and Feedback Theory
The first foundation is the theory of formative assessment. Black and Wiliam (1998) demonstrated in their landmark review that assessment practices generating and using evidence of learning during instruction can yield some of the largest achievement gains documented in educational research, particularly for lower-attaining learners. Their later theorization specifies the mechanism: formative assessment works by clarifying learning intentions and success criteria, engineering effective evidence-eliciting tasks, providing feedback that moves learning forward, and activating learners as instructional resources for one another and as owners of their own learning (Black & Wiliam, 2009; Wiliam, 2011). Sadler (1989) supplies the classic conceptual core, arguing that formative assessment requires the learner to hold a concept of the target standard, to compare current performance with that standard, and to take action to close the gap.
Feedback research refines these ideas into design guidance. Hattie and Timperley (2007) show that feedback is most powerful when it answers three questions, namely where the learner is going, how the learner is doing, and where the learner should go next, and when it operates at the level of the task, the process, and self-regulation rather than at the level of the person. Shute (2008) adds that formative feedback should be specific, non-evaluative, and appropriately timed, with elaborated feedback generally outperforming simple verification. Nicol and Macfarlane-Dick (2006) reposition feedback within self-regulated learning, arguing that good feedback practice develops the learner’s own capacity to monitor and evaluate performance. That this mechanism can be carried by automated systems is no longer merely plausible: in a longitudinal randomized field experiment, Bellhäuser et al. (2023) found that automatically generated daily feedback measurably improved students’ use of self-regulated learning strategies. For an adaptive Bahasa Melayu system, these findings translate directly. Every automated response should carry diagnostic information mapped to curriculum standards, should model the process for improvement, and should progressively transfer evaluative responsibility to the learner, for example through self-assessment features whose benefits for self-regulation and self-efficacy are meta-analytically supported (Panadero et al., 2017).
3.2. Sociocultural Theory and the Zone of Proximal
Development
The second foundation is Vygotskian sociocultural theory. Vygotsky (1978) defined the zone of proximal development as the distance between what a learner can do independently and what the learner can achieve with guidance, and argued that instruction is most productive when it targets this zone. Adaptive assessment is, in a precise sense, a technological operationalization of this idea: by continuously estimating what the learner can currently do and selecting tasks slightly beyond that level, an adaptive system keeps the learner working within the zone of proximal development rather than below it or hopelessly above it. Dynamic scaffolding follows the same logic, with hints, models, and worked examples offered contingently and withdrawn as competence grows. Sociocultural theory also guards against an impoverished, machine-only conception of assessment. Language is learned in social interaction, and the mediating role of the teacher and of peers remains essential; the adaptive system should therefore be designed as a mediating artifact within classroom activity, informing and enriching human interaction around Bahasa Melayu rather than substituting for it.
3.3. Heutagogy and Learner Agency
The third foundation is heutagogy, the study of self-determined learning. Heutagogy extends andragogy by placing the learner at the center of decisions about what, how, and when to learn, and by emphasizing capability, namely the capacity to use competencies in novel situations, alongside competency itself (Hase & Kenyon, 2007). Blaschke (2012) argues that digital technologies are natural enablers of heutagogical practice because they give learners direct control over pathways, pacing, and resources while generating records that support reflection. A systematic review of heutagogical approaches in language learning by Kamrozzaman and Jie (2025) qualifies this promise in ways directly relevant to the present framework: across 33 studies, self-determined approaches were consistently associated with greater autonomy, metacognitive awareness, and proficiency, but their success depended on learner readiness, culturally responsive adaptation, and supportive technological ecosystems rather than following automatically from the provision of choice. Applied to adaptive assessment, heutagogy implies that adaptivity should not be a black box silently routing learners through items. The system should instead expose the learner’s proficiency profile in an age-appropriate form, allow the learner to set goals, choose practice domains, and request challenges, and prompt reflection on progress. This orientation aligns with self-determination theory, which holds that motivation flourishes when learners experience autonomy, competence, and relatedness (Deci & Ryan, 2000; Ryan & Deci, 2000). An adaptive system that keeps success rates in a productive range supports perceived competence; one that offers meaningful choice supports autonomy; and one that feeds into classroom discussion and peer activity supports relatedness.
3.4. TPACK and Multimedia Learning
The fourth foundation concerns the teacher and the interface. The Technological Pedagogical Content Knowledge (TPACK) framework holds that effective technology integration arises from the interaction of technological knowledge, pedagogical knowledge, and content knowledge rather than from any of these alone (Mishra & Koehler, 2006). An adaptive Bahasa Melayu assessment tool is therefore not a neutral instrument that any teacher can simply be handed; its productive use requires teachers who understand the Bahasa Melayu constructs being assessed, the formative pedagogy the tool is meant to serve, and the affordances and limits of the technology itself. The framework developed below accordingly treats teacher dashboards, professional development, and teacher override of system decisions as first-class design components. Finally, because assessment tasks are themselves instructional messages, their design should respect the principles of multimedia learning and cognitive load theory: multimodal Bahasa Melayu tasks should combine audio, text, and image in ways that direct attention to the target construct, avoid extraneous processing, and manage intrinsic load through sequencing and worked support (Mayer, 2017; Sweller et al., 2019).
4. A Conceptual Framework for Adaptive Bahasa Melayu Assessment
Building on the reviewed literature and theoretical foundations, this section proposes the Adaptive Technology-Enhanced Assessment for Bahasa Melayu (ATEA-BM) framework. The framework has two components: a set of six design principles stating what any adaptive Bahasa Melayu assessment tool should achieve, and a four-layer system architecture stating how a tool can be organized to achieve it. The framework is deliberately technology-agnostic at the level of specific products, so that it can guide diverse implementations from lightweight mobile applications to full learning-management-system modules. Table 1 contrasts the resulting model of assessment with prevailing practice, in order to make clear what the framework is intended to change.
Table 1. Conventional and adaptive assessment of Bahasa Melayu compared.
Dimension |
Conventional Bahasa Melayu
assessment |
ATEA-BM adaptive assessment |
Task allocation |
One fixed form for the whole class |
Task difficulty and scaffolding matched to each learner |
Diagnostic grain |
A single tahap penguasaan per skill |
Subskill profile: imbuhan, ayat, ejaan, discourse, ideas |
Feedback latency |
Days to weeks after performance |
Seconds, within the task itself |
Feedback content |
Marks and brief general comments |
Feature-specific guidance linked to a revision action |
Learner role |
Recipient of a grade |
Sets goals, chooses pathways, self-assesses against a visible profile |
Teacher role |
Marker and recorder |
Interpreter of diagnostic evidence and orchestrator of pedagogy |
Evidence for PBD |
Compiled manually from
separate records |
Generated continuously and exported in PBD-compatible form |
4.1. Design Principles
Principle 1: Curriculum-anchored construct definition. Every item, task, and feedback message must be explicitly mapped to the content standards, learning standards, and tahap penguasaan of the Bahasa Melayu DSKP (Kementerian Pendidikan Malaysia, 2018) and to authoritative descriptions of the language system (Nik Safiah Karim et al., 2015). This anchoring follows evidence-centered design, which requires a clear chain of reasoning from observed performance to claims about competence (Mislevy et al., 2003). Curriculum anchoring is what converts an adaptive score into information a Malaysian teacher can act on within PBD.
Principle 2: Adaptivity in the zone of proximal development. The system should select tasks so that each learner works at the edge of current competence, with success neither guaranteed nor improbable, operationalizing Vygotsky (1978) through item response theory calibration (Embretson & Reise, 2000) and adaptive selection algorithms (van der Linden & Glas, 2010). Adaptivity should extend beyond difficulty to the type of scaffold offered, so that the system adapts support as well as challenge (Yang et al., 2022).
Principle 3: Diagnostic, forward-looking feedback. Feedback must be immediate, specific to the linguistic feature involved, expressed in encouraging and non-evaluative language, and structured around where the learner is going, how the learner is doing, and what to do next (Hattie & Timperley, 2007; Shute, 2008). For Bahasa Melayu, this means feedback at the level of, for example, a particular imbuhan choice, a sentence-structure pattern, or a discourse feature of karangan writing, never merely a percentage score.
Principle 4: Learner agency and transparency. Consistent with heutagogy and self-determination theory, learners should see an intelligible model of their own proficiency, set goals, exercise choice over practice pathways, and be prompted to reflect and self-assess (Blaschke, 2012; Ryan & Deci, 2000; Panadero et al., 2017; Kamrozzaman & Jie, 2025). Adaptive routing decisions should be explainable to learners and teachers in plain language. This requirement is not cosmetic: Bearman and Ajjawi (2023) argue that learning to work with algorithmic systems whose reasoning is partly opaque is itself an educational task, and that pedagogy should equip students to interrogate such systems rather than simply comply with them. Khosravi et al. (2022) set out what explainability requires in educational settings specifically, distinguishing explanations of what the system did from explanations a learner can act on, and it is the latter that an adaptive Bahasa Melayu tool must provide.
Principle 5: Teacher-in-the-loop orchestration. The system augments rather than replaces teacher judgment (Luckin & Cukurova, 2019; Ouyang & Jiao, 2021; Hopfenbeck et al., 2023). Teachers receive dashboards aggregating diagnostic information at individual and class level, can override system decisions, can assign targeted tasks, and can export evidence in forms compatible with PBD recording requirements, thereby reducing rather than adding to workload (Arumugham, 2020; Hock et al., 2022). There is direct evidence that this design decision matters: in K-12 classrooms, Knoop-van Campen et al. (2023) found that teacher dashboards had an equalizing effect on feedback, extending teacher attention to pupils who would otherwise have received less of it.
Principle 6: Equity, access, and cultural validity. Task content should reflect the sociocultural worlds of Malaysian learners, accommodate second-language learners of Bahasa Melayu, function on low-cost mobile devices and under intermittent connectivity (Sung et al., 2016; Kukulska-Hulme & Shield, 2008), and be monitored for fairness across gender, location, language background, and socioeconomic groups (Timmis et al., 2016; UNESCO, 2023).
Table 2 summarizes the six principles, their principal theoretical grounding, and their main design implications. The principles are interdependent: adaptivity without curriculum anchoring produces efficient measurement of the wrong things; diagnostic feedback without learner agency produces dependence; and any of these without teacher orchestration produces a tool that classrooms will not sustain.
Table 2. Design principles of the ATEA-BM framework.
Design principle |
Principal theoretical grounding |
Main design implication |
1. Curriculum-anchored construct definition |
Evidence-centered design (Mislevy et al., 2003); DSKP standards (Kementerian Pendidikan Malaysia, 2018) |
Every task and feedback message is mapped to content standards, learning standards, and tahap penguasaan. |
2. Adaptivity in the zone of proximal development |
Sociocultural theory (Vygotsky, 1978); item response theory (Embretson & Reise, 2000) |
Task difficulty and scaffolding are continuously matched to each learner’s estimated proficiency. |
3. Diagnostic,
forward-looking feedback |
Feedback theory (Hattie & Timperley, 2007; Shute, 2008) |
Immediate, specific, non-evaluative feedback identifies the linguistic feature involved and the next step. |
4. Learner agency and
transparency |
Heutagogy (Blaschke, 2012; Kamrozzaman & Jie, 2025); self-determination theory (Ryan & Deci, 2000) |
Learners view their proficiency profile, set goals,
choose pathways, and self-assess; adaptive decisions
are explainable. |
5. Teacher-in-the-loop
orchestration |
TPACK (Mishra & Koehler, 2006);
AI-supported paradigm
(Ouyang & Jiao, 2021) |
Dashboards, override controls, and PBD-compatible
evidence exports keep teachers in control and reduce workload. |
6. Equity, access,
and cultural validity |
Critical and equity perspectives (Selwyn, 2016; UNESCO, 2023) |
Mobile-first, low-bandwidth delivery; culturally relevant content; fairness monitoring across learner subgroups. |
4.2. System Architecture
The ATEA-BM architecture comprises four interacting layers, shown in Figure 1. The first is the domain model layer, a structured representation of the Bahasa Melayu curriculum: a map of skills, subskills, grammatical systems, vocabulary strata, and text genres, each linked to DSKP standards and tahap penguasaan descriptors, and each associated with a calibrated bank of tasks spanning selected-response, short constructed-response, extended writing, listening, and speaking formats. The domain model also encodes typical error taxonomies for Malay, such as affixation errors, sentence-structure errors, spelling and orthographic errors, and register or discourse errors, which the diagnostic components consume.
The second is the learner model layer, which maintains for each pupil a multidimensional, probabilistic profile of proficiency across the domain model, updated after every interaction in the manner of item response theory ability estimation (Embretson & Reise, 2000) and enriched with error-pattern histories, response-time data, scaffold usage, goal selections, and reflection records. The learner model is the system’s memory of the pupil and the substrate for all adaptation, and it must be visible in appropriate forms to pupil, teacher, and parent (Luckin & Cukurova, 2019).
The third is the adaptive engine layer, which makes three kinds of decisions. Selection decisions choose the next task to keep the learner within the zone of proximal development. Scaffolding decisions choose whether and how to offer hints, models, sentence starters, glossaries, or worked examples, and when to fade them. Feedback decisions compose diagnostic messages by combining automated analysis of the response with the error taxonomy and feedback templates built according to the principles of Hattie and Timperley (2007) and Shute (2008). For constructed responses, the engine draws on automated scoring components, namely automated writing evaluation for karangan and short answers (Attali & Burstein, 2006; Shermis, 2014; Warschauer & Grimes, 2008) and speech recognition with pronunciation and fluency scoring for oral tasks (Loewen et al., 2019), with routing of low-confidence machine judgments to the teacher in accordance with validity guidance for automated scoring (Xi, 2010; Huang et al., 2025).
![]()
Figure 1. The four-layer ATEA-BM system architecture.
The fourth is the feedback and reporting layer, which delivers learner-facing feedback, dashboards, and reflection prompts; teacher-facing class analytics, misconception alerts, and PBD-compatible evidence exports; and system-facing analytics for item calibration, fairness monitoring, and continuous improvement. Across all four layers, the architecture assumes the AI-supported and AI-empowered paradigms, in which the technology collaborates with, and progressively empowers, human actors (Ouyang & Jiao, 2021).
Figure 1 also makes explicit a feature that distinguishes the architecture from a conventional testing engine. Evidence does not simply flow outward to a report and stop. The feedback and reporting layer returns diagnostic information to the pupil, and the pupil’s subsequent responses re-enter the learner model as fresh evidence, so that the system’s picture of the learner is continuously revised. The teacher sits on this loop rather than outside it, receiving dashboard evidence and retaining the authority to override adaptive decisions and assign targeted tasks.
The same logic can be expressed temporally rather than structurally. Figure 2 presents the adaptive formative assessment cycle that the architecture is designed to support, in which elicitation, capture, interpretation, feedback, and action follow one another continuously rather than being separated by the days or weeks that characterize conventional marking cycles.
Figure 2. The adaptive formative assessment cycle in Bahasa Melayu.
4.3. Formative Adaptive Practice and Adaptive Measurement Distinguished
Adaptivity serves two distinct purposes within ATEA-BM, and conflating them would compromise both. The first is adaptive formative practice: continuous, low-stakes cycles in which task selection, scaffolding, and feedback are organized to move learning forward, in the manner of the adaptive formative assessment system reported by Yang et al. (2022). In this mode the proficiency estimate is a working hypothesis used to route the next task; newly authored items may be seeded before they are fully calibrated; hints, glossaries, and worked models are available on request; and feedback is delivered inside the task. Precisely the features that make this mode pedagogically valuable make its observations unsuitable as measurement data, because responses are neither independent of the support received nor collected under standardized conditions. The second purpose is adaptive measurement: periodic, deliberately bounded check-in events delivered under standardized conditions from a calibrated bank, without scaffolds or in-task feedback, with item-exposure control and a stopping rule based on the standard error of the ability estimate, in the manner of conventional computerized adaptive testing (Weiss & Kingsbury, 1984; Wainer, 2000; van der Linden & Glas, 2010). The two modes share a domain model and a learner model, but they are separate events, are visibly distinct to learners and teachers, and yield outputs carrying different warrants.
It follows that not every output of the learner model carries the same evidential weight, and the framework therefore assigns outputs to three tiers of use, summarized in Table 3. Tier 1 comprises decisions internal to the learning process: next-task selection, scaffold choice and fading, practice recommendations, the learner-facing proficiency profile, and class-level misconception alerts. These require content-alignment review and evidence that learners and teachers understand them, but no psychometric warrant, because a poor routing decision is corrected by the next observation rather than recorded about a pupil. Tier 2 comprises diagnostic statements offered to the teacher as candidate evidence for PBD: subskill profiles, provisional indications of tahap penguasaan, and exported work samples with their automated analyses. These may inform PBD, but only through the teacher and only when the evidence conditions set out in Section 4.4 are met; the system presents them as claims to be confirmed, amended, or rejected, and what is recorded is the teacher’s decision rather than the machine’s estimate. The teacher, not the engine, assigns the tahap. Tier 3 comprises any use in which a system-generated estimate is itself reported as attainment: formal reporting without teacher confirmation, growth reporting, comparison across classes or schools, placement into intervention or streaming, and research or accountability uses. No ATEA-BM output may be used at Tier 3 until it is supported by a validity argument of the kind language testing scholarship requires (Chapelle & Douglas, 2006; Xi, 2010): item calibration on samples representative of the intended population, evidence of dimensionality and model fit, conditional standard errors of measurement within stated bounds, differential item functioning analyses across the subgroups named in Principle 6, and, for automatically scored writing and speech, human-machine agreement together with evidence that the machine is scoring the construct rather than a correlate of it (Xi, 2010; Huang et al., 2025).
Table 3. Tiers of use for ATEA-BM learner-model outputs.
Tier of use |
Example outputs |
Evidence required before use |
Tier 1: Pedagogical
routing, internal to the learning process |
Next-task selection; scaffold selection and
fading; practice recommendations;
learner-facing proficiency profile;
class misconception alerts |
Alignment of tasks and feedback with DSKP standards; expert content review; evidence that learners and
teachers understand the displays; no psychometric
warrant required |
Tier 2: Teacher-mediated evidence for PBD |
Subskill diagnostic profiles; provisional tahap penguasaan indications; work samples
exported with their automated analyses |
Evidence-sufficiency conditions in Table 4; documented error taxonomy; reported reliability and confidence thresholds for automated components; teacher
confirmation before anything is recorded |
Tier 3: Reported
measurement, including attainment, growth,
comparison, placement, and accountability uses |
System-generated proficiency estimates
reported as attainment or growth;
comparisons across classes or schools;
placement into intervention or streaming;
research outcome measures |
IRT calibration on representative samples;
dimensionality and model-fit evidence; conditional standard errors within stated bounds; DIF analyses across subgroups; human-machine agreement and
construct-validity evidence for automated scoring |
Two practical consequences follow. First, the interface must mark the mode, so that learners know whether they are practicing or being measured; the validity of a formative system depends on learners engaging with it as an opportunity to learn rather than an occasion to perform (Stobart, 2008; Bennett, 2011). Second, the learner model must carry provenance with every observation, recording the mode in which it was obtained, the scaffolds in effect, and the confidence of any automated score, so that measurement-grade estimates can be computed from measurement-grade evidence alone. Without such provenance the convenience of a single accumulating profile would quietly import practice data into reported results, which is the most likely route by which a formative system becomes an unvalidated high-stakes one.
4.4. Evidence Sufficiency for Subskill Diagnosis
The diagnostic ambition of the framework carries a corresponding risk. A learner who selects the wrong affix in a single sentence has made an error; a learner who is told that imbuhan is a weakness has been given something closer to an identity. Because a subskill diagnosis in ATEA-BM propagates into the learner’s visible profile, the teacher’s dashboard, and potentially a PBD record, the system must not convert isolated errors into stable proficiency deficits. The framework therefore treats a diagnosis as a claim licensed by stated evidence conditions rather than as a byproduct of scoring, following the evidence-centered logic that observed performances warrant claims about competence only through an explicit evidence model (Mislevy et al., 2003).
Six conditions, summarized in Table 4, must be satisfied before the system may report a subskill such as imbuhan, pembinaan ayat, or discourse organization as an area of weakness. The first is volume: a minimum number of scorable opportunities addressing that subskill, with an illustrative default of eight for discrete features such as affixation or spelling and three extended performances for text-level constructs such as organization or idea development, since a judgment about discourse cannot rest on a single composition. The second is distribution: evidence drawn from at least three distinct tasks on at least two occasions separated in time, so that no diagnosis can be generated within a single sitting, where fatigue, a misread prompt, or an unfamiliar topic may account for the pattern. The third is format variety: at least two response formats, including at least one constructed response for productive subskills, so that a finding is not an artifact of one task type; a learner who identifies the correct affix in a selected-response item but not in free writing has a production difficulty rather than a knowledge gap, and the two warrant different feedback. The fourth is confidence: observations produced by automated writing or speech analysis count towards a threshold only when the component’s confidence exceeds its documented operating threshold, with low-confidence cases retained for teacher review rather than counted. The fifth is specificity: the pattern must be distinguishable from plausible alternative explanations, so that the system withholds a subskill diagnosis when the same responses are equally consistent with a vocabulary gap, a failure to comprehend the prompt, or off-task responding, and reports the ambiguity to the teacher instead of resolving it silently. The sixth is recency: evidence contributes with a weight that declines across a defined window, by default one school term, so that a diagnosis lapses unless it is re-confirmed and a remediated weakness does not persist in the profile as a permanent label.
Table 4. Minimum evidence conditions before a subskill may be diagnosed.
Condition |
Operational requirement (illustrative defaults) |
What it guards against |
Volume |
At least eight scorable opportunities for discrete features such as imbuhan or ejaan; at least three extended performances for text-level constructs such as organization or idea development |
A diagnosis resting on one or
two errors |
Distribution |
Evidence from at least three distinct tasks on at least two occasions
separated in time |
Single-session artifacts such as
fatigue, a misread prompt, or an
unfamiliar topic |
Format variety |
At least two response formats, including at least one constructed response for productive subskills |
Confusing a production difficulty with a knowledge gap, or the reverse |
Confidence |
Automated observations counted only above the component’s documented confidence threshold; low-confidence cases routed to the teacher |
Machine error accumulating into a learner-facing conclusion |
Specificity |
Alternative attributions such as vocabulary, prompt comprehension,
or off-task responding ruled out, or the ambiguity reported to the teacher |
Attributing to one subskill an error produced by another |
Recency |
Evidence weighted down across a defined window, by default one school term, with the diagnosis lapsing unless re-confirmed |
Stale labels persisting after a
difficulty has been remediated |
Diagnoses are also graded in strength rather than switched on. Below threshold, an emerging pattern may be shown to the teacher only, described in behavioral terms and marked as unconfirmed; at threshold, it becomes a Tier 2 diagnostic claim, visible to the learner and eligible for teacher-confirmed use in PBD. Learner-facing wording follows the same logic, describing what happened rather than what the pupil is: a learner should read that the meN-prefix was required in four of the last six sentences in which it appeared, not that they have failed to master affixation. The numerical values proposed here are illustrative rather than settled. They are stated because the design requirement is that such thresholds be explicit, documented, and auditable rather than buried in code; their calibration is an empirical task belonging to the research agenda in Section 7, and it should be undertaken separately for each subskill, since the evidence needed to identify a stable affixation weakness is unlikely to equal that needed for discourse organization.
4.5. Safeguards for Automated Writing and Speech Analysis
Automated analysis of karangan and of oral performance is the component of the framework with the greatest capacity to do harm, both because Malay is under-resourced relative to English (Godwin-Jones, 2021; Maxwell-Smith et al., 2022) and because pronunciation scoring is least reliable for exactly the accented and multilingual speech that Malaysian classrooms contain (Rogerson-Revell, 2021). Three safeguards are therefore treated as parts of the design rather than as implementation details. The first is a set of human review thresholds. Any automated score whose confidence falls below the component’s documented operating threshold is routed to the teacher rather than shown to the learner as a judgment. Any score that would move a learner across a reporting boundary is confirmed by a teacher before it is recorded. A random audit sample of automatically scored work, with an illustrative default of ten per cent per class per reporting cycle, is independently rated by teachers, so that drift is detected in normal operation rather than after complaint. And no automated score is used alone to place a learner into intervention, grouping, or streaming. Where audited agreement with teacher ratings falls below a component’s stated operating standard for a particular task type or learner group, that component reverts to advisory use for those cases until it has been retrained and re-audited.
The second safeguard is a learner correction and appeal mechanism. Every automated judgment carries a visible control by which the learner can dispute it, and the framework treats the use of that control as an act of feedback literacy rather than as a complaint (Ranalli, 2018). A disputed observation is quarantined immediately: it is suspended from the learner model so that it cannot contribute to a subskill diagnosis while under review, and it is queued to the teacher together with the original response, the automated analysis, and the learner’s stated reason. The teacher’s resolution is recorded, is visible to the learner, and updates the learner model, whether by reinstating, amending, or discarding the observation. Disputes are then aggregated and analyzed, since a cluster of appeals against a particular feedback template, item, or speaking task is diagnostic of a fault in the system rather than in the learners, and this analysis feeds the item and model revision cycle. Making automated reasoning contestable in this way is itself educative, since learning to interrogate rather than comply with algorithmic judgment is a legitimate curricular aim (Bearman & Ajjawi, 2023), and it requires explanations a learner can act on rather than descriptions of what the system did (Khosravi et al., 2022).
The third safeguard is routine auditing of error rates across learner subgroups, because aggregate accuracy can conceal systematic disadvantage and the groups at risk in Malaysian classrooms are identifiable in advance. Automated components should be evaluated, and their results reported, disaggregated by gender, home language background, urban or rural location, state or region, and, where an acceptable indicator exists, socioeconomic status. For writing this means agreement with human raters together with the direction and size of any systematic score difference; for speech it means recognition error rates and pronunciation-score bias by first-language and accent group. Differential item functioning analysis provides the corresponding check at item level (Embretson & Reise, 2000). Audits are repeated after every model or item-bank update, since a retrained component is in effect a new instrument, and each component is documented with a description of its training data, its intended range of use, and its known limitations. Where a disparity exceeds a pre-declared tolerance, the response is graduated: the component is restricted to advisory use for the affected group, teacher review becomes mandatory for those cases, and the component is withdrawn if the disparity persists after retraining. Audit summaries should be reported to participating schools rather than held privately by developers, learners and parents should be told plainly when work is analyzed automatically, and learner writing and speech should not be used to train models without informed consent and de-identification. These provisions give operational content to Principle 6 and to international guidance requiring transparency and fairness monitoring in educational technology (UNESCO, 2023; Timmis et al., 2016).
4.6. Balancing Learner Control with Curriculum Progression
Principle 4 gives learners goals, choices, and a visible profile; Principle 1 anchors the system in a curriculum that is a national entitlement rather than a menu. The tension between them is real, since learners tend to avoid the domains in which they experience repeated failure, which are precisely the domains in which the DSKP requires progress. The framework resolves this through bounded autonomy, an arrangement in which choice is genuine but operates above a curriculum floor. Self-determination theory supports rather than qualifies this construction, since autonomy support in that literature means meaningful choice within structure and not the absence of structure (Ryan & Deci, 2000), and the review of heutagogical language learning by Kamrozzaman and Jie (2025) found that self-determined approaches succeed where learner readiness and supportive structures are present rather than wherever choice is offered.
Four mechanisms implement bounded autonomy. First, a curriculum floor: the learning standards prescribed for the year remain a required set, and learner choice operates over sequence, task format, topic and text content, pacing within a window, and the destination of optional additional practice, but never over whether a required domain is addressed. Second, choice within the necessary: when a required domain falls due, the system offers a constrained menu in which every option targets that domain while differing in format, context, or length, so that the learner exercises real control over how the difficult work is done rather than whether it is done. Third, coverage monitoring with avoidance detection: the engine tracks evidence accumulated in each domain against the year’s requirements and compares the learner’s choices with their estimated proficiency, so that a pattern of selecting away from the domains in which proficiency is lowest is identified as avoidance rather than preference, restored to the queue, and flagged to the teacher, who can address the aversion pedagogically. Fourth, support before compulsion: where avoidance is detected the system’s first response is to lower entry difficulty and raise scaffolding in that domain so that the learner re-enters it at a success rate within the productive band, since sustained failure is the mechanism that generated the avoidance and compulsion without that adjustment would merely reproduce it.
Two further provisions keep the arrangement legible and developmentally appropriate. The trade-off is made explicit to the learner rather than enforced silently: the coverage map is visible, the system states which required domain is now due and why, and the learner chooses within it, consistent with the requirement that adaptive decisions be explainable in terms a learner can act on (Khosravi et al., 2022). And the breadth of choice is graded by level and readiness rather than fixed. At Year 4 it is narrow, with goal setting scaffolded and conducted in conference with the teacher; through lower secondary it widens, so that learners set term goals, request challenge, and schedule their own review. Throughout, the teacher retains the authority to widen, narrow, or suspend choice for an individual, so that bounded autonomy remains a professional judgment about a particular learner rather than a fixed system setting (Principle 5).
4.7. Applying the Framework across the Four Language Skills
The framework’s generality is tested by whether it yields concrete design guidance for each skill in the Bahasa Melayu curriculum. For reading (membaca), the system adaptively selects texts by lexical difficulty, syntactic complexity, and genre, assesses literal, inferential, and evaluative comprehension, and diagnoses whether difficulties stem from vocabulary, sentence processing, or discourse integration. Feedback directs learners to the textual evidence relevant to each question, modeling comprehension strategy rather than merely marking answers, in keeping with process-level feedback (Hattie & Timperley, 2007).
For writing (menulis), automated writing evaluation provides immediate feedback on mechanics, imbuhan and sentence structure, vocabulary range, and organization, staged so that early drafts receive meaning-focused guidance and later drafts receive accuracy-focused guidance, a sequencing consistent with classroom research on automated feedback integration (Li et al., 2015; Warschauer & Grimes, 2008). Because learners vary in their ability to interpret automated feedback, messages must be short, exemplified, and linked to revision actions, and teachers should explicitly teach feedback literacy (Ranalli, 2018).
For listening (mendengar), adaptive selection varies speech rate, accent, text length, and task type, using authentic Malaysian audio. Diagnosis distinguishes perception difficulties from comprehension difficulties by combining item types, and multimedia design principles govern the pairing of audio with visual support (Mayer, 2017). For speaking (bertutur), speech-enabled tasks range from read-aloud and repetition through picture description to simulated interaction, with automated scoring of pronunciation, fluency, and, progressively, content, and with periodic teacher-rated oral tasks retained both for validity and for the social dimension of oral language emphasized by sociocultural theory (Vygotsky, 1978). Retaining teacher-rated oral tasks also hedges a technical limitation, since pronunciation scoring remains least reliable for precisely the accent range Malaysian classrooms contain (Rogerson-Revell, 2021). Grammar and vocabulary run through all four skills as cross-cutting strands, assessed both discretely for diagnosis and integratively within skill tasks, mirroring the integrated design of the DSKP (Kementerian Pendidikan Malaysia, 2018). Table 5 summarizes these applications.
Table 5. Application of the ATEA-BM framework across the Bahasa Melayu language skills.
Skill |
What the engine adapts |
What the system diagnoses |
Feedback emphasis |
Reading (membaca) |
Lexical difficulty, syntactic
complexity, genre, question type |
Vocabulary, sentence processing, or discourse integration as the source of difficulty |
Directs the learner to the textual evidence and models
comprehension strategy |
Writing (menulis) |
Prompt complexity, planning support, sentence starters,
glossary access |
Imbuhan and sentence structure, spelling, vocabulary range,
organization, idea development |
Meaning-focused on early drafts; accuracy-focused on later drafts |
Listening (mendengar) |
Speech rate, accent, text length, visual support, task type |
Perception difficulty distinguished from comprehension difficulty |
Replays targeted segments and names the listening strategy
required |
Speaking (bertutur) |
Task openness, from read-aloud through picture description to simulated interaction |
Pronunciation, fluency, and
progressively content; low-confidence cases routed to the teacher |
Models the target production;
periodic teacher-rated tasks
retained for validity |
Grammar and
vocabulary
(tatabahasa dan kosa kata) |
Item type and spacing across the review cycle |
Systematic error patterns mapped to the Malay error taxonomy |
Explains the rule through a
contrasting example,
then re-tests it later |
5. Discussion and Implications
5.1. Implications for Classroom Practice
The most immediate implication of the ATEA-BM framework is a reconfiguration of the assessment cycle in Bahasa Melayu classrooms. Where the conventional cycle runs from teaching to testing to grading over days or weeks, an adaptive tool compresses the elicitation, interpretation, and feedback phases into seconds and distributes them across every lesson. This does not diminish the teacher’s role; it changes its center of gravity. Freed from a large share of routine marking, the teacher can concentrate on the activities that research identifies as most consequential: designing rich tasks, orchestrating discussion of common errors surfaced by the dashboard, conferencing with individual pupils about their proficiency profiles, and activating peer and self-assessment (Black & Wiliam, 2009; Wiliam, 2011). The tool supplies the evidence; the teacher supplies the pedagogy. Practically, schools can embed adaptive assessment in three recurring routines: short diagnostic warm-ups at the start of instructional units, differentiated independent practice during lessons, and periodic proficiency check-ins whose results feed directly into PBD records, addressing the workload concerns documented in implementation research (Arumugham, 2020; Hock et al., 2022).
A second practice implication concerns differentiation. Because the adaptive engine individualizes task difficulty and scaffolding, mixed-ability classes no longer force a choice between boring the proficient and overwhelming the struggling. This matters particularly for pupils learning Bahasa Melayu as a second language, who can receive additional lexical support and slower-paced listening input without public differentiation, preserving dignity while personalizing support. The motivational consequences predicted by self-determination theory, namely sustained perceived competence and autonomy, are among the framework’s most important intended outcomes (Ryan & Deci, 2000).
5.2. Implications for Teacher Professional Development
The TPACK framework implies that the introduction of adaptive assessment tools must be accompanied by professional development integrating technological, pedagogical, and content knowledge rather than treating tool training as a technical matter (Mishra & Koehler, 2006). Teachers need to understand what the proficiency estimates and diagnostic categories mean linguistically, how to convert dashboard information into instructional moves, and how to cultivate pupils’ feedback literacy so that automated feedback is understood and used (Ranalli, 2018; Nicol & Macfarlane-Dick, 2006). The evidence that teachers’ assessment practice in Bahasa Melayu is already stretched by procedural demands (Said et al., 2021; Hock et al., 2022) suggests that professional development should lead with workload relief and diagnostic interpretation rather than with system features. Pre-service Bahasa Melayu teacher education should likewise incorporate technology-enhanced assessment as a core component of assessment literacy (Kessler, 2018). The systematic review by Børte et al. (2023) reaches a compatible conclusion from the international literature: where teachers lack assessment-specific technological knowledge, they assimilate digital tools into existing test-and-grade routines instead of using them to change practice, which is exactly the failure mode an adaptive tool introduced into a PBD context would be prone to. Professional development must also cover the safeguards themselves, since human review thresholds, learner appeals, and the authority to override a system judgment protect learners only if teachers know when and how to exercise them, and since the boundary between evidence that may inform PBD and estimates that require psychometric validation before formal reporting is a matter of assessment literacy rather than of software training (Sections 4.3 to 4.5). Attention is also owed to teachers’ and students’ broader dispositions towards artificial intelligence: in a survey of 150 Malaysian university students, Jie and Kamrozzaman (2024) found that over-reliance on AI and limited understanding of how it works were both significantly associated with poorer learning experiences, a reminder that introducing intelligent assessment without accompanying AI literacy risks producing dependence rather than capability. Without such investment, the international experience is unambiguous: technology-centered deployments that ignore teacher knowledge and school culture underperform or fail (Selwyn, 2016; Zawacki-Richter et al., 2019).
5.3. Implications for Policy
At system level, the framework aligns with and operationalizes the assessment aspirations of the Malaysia Education Blueprint 2013-2025, which explicitly seeks to leverage information and communications technology to raise the quality of learning and to strengthen school-based assessment (Ministry of Education Malaysia, 2013), and it is consistent with the national commitment to artificial intelligence capability set out in the Malaysia National Artificial Intelligence Roadmap (Ministry of Science, Technology and Innovation, 2021). Three policy moves would accelerate progress. First, the development of a national, openly documented item and task bank for Bahasa Melayu, calibrated with item response theory and mapped to the DSKP, would provide the public infrastructure on which multiple tools could be built (van der Linden & Glas, 2010; Kementerian Pendidikan Malaysia, 2018). Second, investment in Malay language technology resources, including annotated learner corpora and speech datasets representative of Malaysian classrooms, is a precondition for high-quality automated scoring in Bahasa Melayu (Godwin-Jones, 2021; Maxwell-Smith et al., 2022). Third, procurement and data-governance standards should require transparency of adaptive algorithms, fairness monitoring, and strong protection of pupil data, in line with international guidance on technology in education (UNESCO, 2023; Timmis et al., 2016). Policymakers should also resist the temptation to conscript adaptive formative tools into high-stakes accountability, since the validity of formative systems depends on learners engaging with them as opportunities to learn rather than occasions to perform (Stobart, 2008; Bennett, 2011). Malaysia’s comparative performance on international assessments provides the motivation for reform (OECD, 2023), but it should not become the metric against which a formative system is judged.
6. Challenges and Limitations
Several challenges temper the promise of the framework. The first is the language-resource challenge already noted: automated scoring and speech recognition for Bahasa Melayu require investment in data and models that cannot simply be imported from English-language systems, and until such components reach acceptable accuracy, implementations should weight machine judgment conservatively and route uncertain cases to teachers (Xi, 2010; Shermis, 2014; Huang et al., 2025). The second is the equity challenge. Although mobile-first design mitigates access barriers, disparities in devices, connectivity, and home support remain real in parts of Malaysia, and an adaptive system deployed without attention to these disparities could widen rather than narrow proficiency gaps (Sung et al., 2016; UNESCO, 2023; Selwyn, 2016). The third is the pedagogical-culture challenge: in an assessment culture historically oriented to examinations, there is a risk that adaptive tools are used as endless test-preparation machines, which would betray their formative purpose (Stobart, 2008; Arumugham, 2020). The fourth is construct coverage. Some valued outcomes of Bahasa Melayu education, including extended interactive speaking, literary appreciation, and the cultural and civic dimensions of the language, resist full automation and must remain substantially teacher-assessed; the framework treats this as a design feature rather than a defect, consistent with its teacher-in-the-loop principle (Chapelle & Douglas, 2006; Luckin & Cukurova, 2019).
As a piece of scholarship, this paper also has limitations. It is conceptual and therefore proposes rather than proves; the framework’s components are grounded in international evidence, but their combined effect in Bahasa Melayu classrooms is an empirical question. The literature base is dominated by studies of English and other high-resource languages, and transfer to Malay involves assumptions about morphology, orthography, and classroom context that require testing. These limitations define the research agenda that follows.
7. Future Research Directions
A program of research is needed to move the ATEA-BM framework from blueprint to validated practice. Design-based research should come first, iteratively developing and refining prototype task banks, feedback templates, and dashboards with Bahasa Melayu teachers and pupils, following the learning-sciences-driven design approach advocated for educational artificial intelligence (Luckin & Cukurova, 2019). Psychometric research should calibrate Bahasa Melayu item banks with item response theory, examine dimensionality across the four skills, and evaluate the precision and efficiency of adaptive delivery against fixed forms (Embretson & Reise, 2000; Weiss & Kingsbury, 1984). The same programme should calibrate the evidence-sufficiency thresholds proposed in Section 4.4 separately for each subskill, estimating how many observations, across how many tasks and occasions, are needed before a diagnosis is stable, and should establish the operating and audit standards for automated components required by Section 4.5, including the subgroup tolerances at which a component reverts to advisory use. Language-technology research should build and validate automated writing evaluation and speech scoring for Malay, reporting accuracy, error analysis, and fairness across learner subgroups, and should articulate validity arguments in the tradition of language assessment research (Attali & Burstein, 2006; Xi, 2010; Chapelle & Voss, 2016; Maxwell-Smith et al., 2022). Intervention research, including quasi-experimental and longitudinal designs, should then estimate effects on proficiency outcomes, self-regulation, and motivation, testing the mediating mechanisms predicted by formative assessment and self-determination theory (Black & Wiliam, 2009; Ryan & Deci, 2000; Panadero et al., 2017; Bellhäuser et al., 2023). Finally, implementation research should examine teacher adoption, professional development models, and system-level scaling conditions, drawing on TPACK and on the critical literature on educational technology to anticipate points of failure (Mishra & Koehler, 2006; Selwyn, 2016; Crompton & Burke, 2023).
8. Conclusion
This paper set out to conceptualize how technology-enhanced assessment, and adaptive digital tools in particular, can be designed to improve language proficiency outcomes in Bahasa Melayu education. Synthesizing research on formative assessment, computerized adaptive testing, artificial intelligence in education, and language assessment technology, it proposed the ATEA-BM framework: six design principles, anchored in formative assessment theory, sociocultural theory, heutagogy, and TPACK, together with a four-layer architecture linking a curriculum-anchored domain model, a rich learner model, an adaptive engine, and a feedback and reporting layer across the four language skills. The central argument is that adaptivity should be understood pedagogically and not merely psychometrically: its value lies in keeping every learner within the zone of proximal development, in converting every response into diagnostic, forward-looking feedback, in cultivating learner agency, and in equipping teachers with evidence they can act on. For Bahasa Melayu, a language of national importance that has been underserved by assessment technology, the framework offers researchers a testable design theory, developers a principled blueprint, and policymakers a route to realizing the assessment ambitions of national education policy. The task ahead is empirical: to build, calibrate, and evaluate adaptive Bahasa Melayu assessment tools in real classrooms, so that the promise theorized here becomes measurable improvement in the language proficiency of Malaysian learners.
Acknowledgements
The authors thank UNITAR International University for supporting this research.
Author Contributions
Conceptualization, N.A.K. and S.E.; methodology, N.A.K.; formal analysis, N.A.K.; investigation (literature identification and screening), N.A.K. and S.E.; writing—original draft preparation, N.A.K.; writing—review and editing, N.A.K. and S.E.; visualization (framework figures and tables), N.A.K.; supervision, N.A.K.; project administration, N.A.K. All authors have read and agreed to the published version of the manuscript.