Abstract
Immersive virtual reality (VR) offers transformative potential for second language acquisition (SLA) by enabling contextualized, task-based collaborative learning. However, the degree to which avatar visual realism and tracking-driven embodiment influence learners' social presence, communicative anxiety, and interactional dynamics remains insufficiently understood. This study investigates the interplay between avatar visual realism (stylized vs. photorealistic) and embodiment fidelity (low-fidelity controller-based inverse kinematics vs. high-fidelity sensor tracking with real-time facial and eye expression mapping) within a collaborative VR language learning environment. A 2×2 between-subjects experimental design was conducted with 192 intermediate English-as-a-Foreign-Language (EFL) university students paired into 96 collaborative dyads. Participants completed an interactive information-gap negotiation task in a bespoke multi-user VR environment. Objective measures of communicative fluency, turn-taking latency, and vocabulary retention were combined with standardized psychometric scales assessing social presence, foreign language anxiety, and perceived cognitive load. The findings reveal that high embodiment fidelity significantly enhanced copresence, mutual understanding, and learners' willingness to communicate. Conversely, photorealistic avatars induced an uncanny valley effect and heightened social anxiety when paired with low tracking fidelity, whereas stylized avatars paired with high embodiment yielded the highest interactional fluency and lowest language anxiety. These findings suggest that behavioral fidelity takes precedence over graphical realism in collaborative educational VR, offering actionable design frameworks for immersive language learning platforms.