Abstract
Automated news generation powered by Large Language Models (LLMs) offers unprecedented scalability in journalistic workflows but risks propagating systemic social biases, political polarization, and factual distortions. Conventional alignment techniques, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), predominantly collapse normative value judgments into a single aggregated reward model, inadvertently marginalizing minority perspectives and failing to capture the complex, multi-faceted ethical standards demanded by professional journalism. In this paper, we propose a Multi-Stakeholder Preference Learning (MSPL) framework designed to achieve nuanced ethical alignment in automated news generation. Our approach disaggregates preference signals across three vital cohorts: professional journalists adhering to canonical editorial ethics, domain-specific fact-checkers evaluating epistemic fidelity, and demographically diverse reader panels representing varied socio-political identities. By formulating ethical alignment as a multi-objective constrained optimization problem along a Pareto frontier, MSPL balances journalistic integrity, epistemic accuracy, and representative neutrality. Empirical evaluations conducted on a novel benchmark of contentious political, socio-economic, and cultural news events reveal that MSPL reduces ideological and demographic bias by 38.4% relative to standard DPO and monolithic RLHF baselines, while simultaneously preserving factual precision and narrative fluency. These findings demonstrate the necessity of multi-stakeholder pluralism in aligning generative AI systems deployed within sensitive socio-technical ecosystems.