Benchmarking Human and AI-Generated Writings in Higher Education: A Multilayered Study on Creativity, Coherence, and Student Perception across Linguistic, Semantic, and Evaluative Features
IEEE Access, cilt.14, ss.105209-105235, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 14
- Basım Tarihi: 2026
- Doi Numarası: 10.1109/access.2026.3711912
- Dergi Adı: IEEE Access
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals
- Sayfa Sayıları: ss.105209-105235
- Anahtar Kelimeler: artificial intelligence, creative writing evaluation, creativity, Educational settings, human evaluation, large language models (LLMs)
- Karadeniz Teknik Üniversitesi Adresli: Evet
Özet
In contemporary educational settings, creativity stands out as a critical competency among higher-order thinking skills. However, the evaluation of this skill is hindered by the predominance of subjective judgments and the lack of standardized criteria, posing a significant challenge. Additionally, the role and comparative proficiency of Large Language Models (LLMs) in the production and evaluation of creative writing, relative to human output, remain an underexplored problem area. This study systematically examines the linguistic, structural, and semantic differences between human-written texts and those generated by LLMs (ChatGPT, Claude, Gemini, DeepSeek, Mistral, Grok) in the context of creative writing. The study involved 91 preservice teachers, yielding 546 human-written texts across six scenarios. For each scenario response, six corresponding LLM-generated texts were produced, resulting in a total corpus of 3,822 texts. Differences in lexical diversity, syntactic complexity, topic distribution, and semantic coherence were analyzed multidimensionally. Employing a mixed-methods approach that included Natural Language Processing (NLP), topic modeling, and machine learning techniques, a total of 65 linguistic features were analyzed. The findings indicate that human texts exhibit higher lexical diversity (TTR=0.71, MTLD=94.3) and simpler, more readable structures (Atesman=68.5), while LLM-generated texts feature longer, more complex sentence structures (average parse tree depth=6.3) and lower semantic coherence (Sentence-BERT similarity=0.65). Human texts demonstrate superiority in descriptive elements (noun and adjective ratio) and thematic focus (entropy=1.84), whereas LLMs tend to produce formulaic narratives. Student evaluations reveal that LLMs perform strongly in logical coherence (QRL) and level of detail (LD), but human texts are preferred for originality (OR) and overall satisfaction (SAT). However, models such as ChatGPT and Grok demonstrate near-human performance in fluency and logical reasoning. A RoBERTa-based classifier achieved high performance in distinguishing human and LLM texts, with 93% accuracy and 97% ROC-AUC. Overall, the findings elucidate the structural and perceptual differences between human and LLM-generated texts and offer significant pedagogical implications for the integration of artificial intelligence into creative writing education.