Abstract
The proliferation of artificial intelligence in creative domains necessitates robust methods for evaluating its outputs, moving beyond mere indistinguishability to assess intrinsic qualities like 'creativity.' This study addresses the challenge of evaluating AI-generated poetry by employing a dual methodology: expert human peer review and computational stylometry. A corpus of 50 AI-generated poems, produced by a transformer-based neural network model, was presented alongside 50 human-authored poems of similar style in a double-blind review to a panel of 15 literary experts. Concurrently, both corpora were subjected to extensive computational stylometric analysis, examining features such as lexical diversity, syntactic complexity, metaphorical density, and sentiment. Results indicate that while AI-generated poems often achieved high scores in technical proficiency and adherence to form, human evaluators frequently distinguished a subtle lack of profound emotional depth or unique human experience. Stylometric analysis revealed significant similarities in surface-level linguistic features but also quantifiable differences in deeper semantic and structural patterns. This research suggests that while AI can convincingly emulate poetic forms, the nuanced perception of 'creativity' by human experts involves criteria still challenging for computational assessment, thereby refining our understanding of AI's creative potential and its evaluation.