Abstract
Rapid urbanization and escalating vehicular volumes present critical challenges to modern metropolitan mobility, exacerbating street-level congestion and driving up transportation-related greenhouse gas emissions. Traditional adaptive traffic signal control heuristics frequently fail to capture high-order, non-linear spatio-temporal correlations across interconnected arterial networks. In this study, we propose a novel Spatio-Temporal Transformer-based Multi-Agent Reinforcement Learning (ST-TARL) framework tailored for decentralized adaptive traffic signal control in smart city ecosystems. By integrating self-attention mechanisms across both temporal sequence horizons and dynamic inter-intersection spatial graphs, each signal agent learns cooperative control policies that explicitly optimize vehicular throughput while mitigating stop-and-go driving patterns. Furthermore, we incorporate an emissions-aware multi-objective reward formulation derived from instantaneous vehicular energy consumption dynamics. Extensive microscopic traffic simulations conducted in SUMO (Simulation of Urban MObility) across both synthetic grid topologies and real-world urban road networks demonstrate that ST-TARL significantly outperforms state-of-the-art baselines. Specifically, our model achieves a 26.4% reduction in average vehicular delay, a 31.8% decrease in cumulative intersection queue lengths, and an 18.2% drop in tailpipe carbon dioxide (CO2) emissions compared to leading deep multi-agent reinforcement learning approaches. These findings underscore the viability of transformer-driven policy optimization as an enabling methodology for sustainable, responsive, and eco-friendly urban intelligent transportation infrastructures.