"Your Language is Not My Language": Understanding Human Mechanisms and Machine Perceptions of Written Production Across Language Variations
Restricted (Penn State Only)
- Author:
- Tang, Zixin
- Graduate Program:
- Information Sciences and Technology
- Degree:
- Doctor of Philosophy
- Document Type:
- Dissertation
- Date of Defense:
- June 18, 2026
- Committee Members:
- Carleen Maitland, Program Head/Chair
Dongwon Lee, Major Field Member
Mahir Akgun, Major Field Member
Wenpeng Yin, Outside Unit & Field Member
Janet van Hell, Co-Chair & Dissertation Advisor
Kenneth Huang, Co-Chair & Dissertation Advisor - Keywords:
- Natural Language Processing (NLP)
Language Diversity
Large language models (LLMs)
Language variation
psycholinguistic - Abstract:
- With easier and more convenient access to the Internet and modern technology, more and more people have started utilizing artificial intelligence (AI), particularly large language models (LLMs), in their daily activities and duties, making machines experience various types of language input as humans do. However, the growing prevalence of language variation presents new challenges for LLMs. Unlike earlier NLP systems, modern LLMs are primarily trained and fine-tuned on web-based data, which is dominated by language varieties with abundant online resources, such as Standard American English for English models and Mainland Mandarin Chinese for Mandarin Chinese models. This imbalance in training data can limit and impact model performance when processing non-standardized language variations (\ie, African American English, Taiwan Mandarin Chinese, languages produced by low-proficiency speakers, etc.), leading to failures to understand users' language input, potential biases and discrimination, and less comprehensible language generation towards specific groups of users. At the same time, LLMs are increasingly being used to study human language and cognition. Despite their growing capabilities, their potential applications in language science remain underexplored. Understanding both the limitations and opportunities of LLMs is therefore important for developing more inclusive language technologies and advancing research on human language processing. Against this backdrop, this dissertation examines how LLMs perceive, process, and respond to language diversity in written communication. It seeks to provide a deeper understanding of potential biases and harms that may arise across language varieties and to identify factors contributing to disparities in model performance. In addition, the dissertation explores ways to leverage LLMs as computational tools for language science research. By quantifying and analyzing model outputs, researchers can use LLMs to conduct simulations and computational analyses that offer new insights into human language processing mechanisms. This dissertation first presents a systematic review of how NLP communities collect data, design experiments, and compare performance results for different language varieties. The review reveals substantial imbalances in the availability of data across languages and varieties, as well as a limited diversity of evaluation tasks. It highlights the need for more nuanced task design and greater investment of resources for studying language diversity within individual languages. Building on the survey, the following studies provide a data collection method and examine whether and how LLMs perform differently in sentiment tasks across language varieties. These studies indicate that performance disparities do occur, although the impacted groups may not be those traditionally assumed to be at greatest risk. Finally, a quantitative analysis of essay writing illustrates how LLM-derived measures can be used to investigate human language production, demonstrating the value of LLMs as tools for language science. Overall, this dissertation contributes methods for collecting large-scale resources on language varieties, provides evidence of performance disparities across user groups, and demonstrates how LLMs can support computational studies of human language. Together, these contributions advance the development of more inclusive and user-centered language technologies for speakers of underrepresented language varieties.
Accessible Version in Progress
We're generating an accessible version of this file to meet ADA Title II requirements. This process may take up to one hour. Please return later to access the accessible copy once it's ready.
You can still download the current version by clicking "OK".
What's happening:
An accessible PDF is being generated using Adobe with AI used to generate alternative text (alt text) for images in the PDF.