Text Production and Comprehension by Human and Artificial Intelligence
Workshop will take place in
Days
Hours
Minutes
Seconds
Talk Summaries
Morten Christiansen: Large Language Models, Creativity, and Textual Culture
Language, particularly as mediated through writing, has been the primary driver for the evolution of human culture for millennia (Christiansen & Chater, 2022). Throughout this long period of cultural evolution, humans have been in the driver’s seat as the sole creator of what might be termed textual culture. However, the explosive growth and sophistication of Large Language Models (LLMs) would appear to challenge humanity’s hegemony over textual culture. In several interrelated lines of research, I am aiming to determine the possible implications of LLMs for textual culture, considering both pitfalls and promises.

In a first line of work, I have been involved in developing an interdisciplinary institutional framework for studying the impact of LLMs on textual culture in a just-submitted Center of Excellence application for a Center for Contemporary Cultures of Text to be hosted at Aarhus University in Denmark. The center will bring together computer scientists, cognitive scientists, linguists, humanists, and other relevant researchers to evaluate how textual culture might be changing, develop more culturally sensitive LLMs for long-form text generation (including for under-resourced languages), assess the opportunities for human-LLM co-creation, and determine in what ways LLMs might help us understand human linguistic abilities (cf. Contreras Kallens et al., 2023). As an initial step toward creating a supporting environment for LLM-related research at Aarhus University, I helped co-found the Center for Language Generation and AI (https://cc.au.dk/en/clai) in 2023. Moreover, I was also part of a university-wide committee at Cornell, tasked with developing guidelines for using Generative AI in education (Bala et al., 2023).

To understand the potential linguistic creativity that LLMs may display, I have been looking into their abilities to generate poetry in a second line of research with my colleague, Laurent Dubreuil (Romance Languages and Comparative Literature) as well as graduate students Pablo Contreras Kallens and Jacob Matthews (and funded by a New Frontiers Grant from Cornell). Poetry involves one of humanity’s most creative use of language, requiring the writer to have excellent command of language and knowledge of how to employ it to express complex ideas and emotional nuances. Using a curated database of poetry representing different time periods and genres, from William Shakespeare, Lord Byron, and Emily Dickinson to Claude McKay, Walt Whitman, and Gertrude Stein, we prompted GPT-3 to generate completions of poem fragments. These were presented in a Poetry Turing Test, along with completions by Cornell undergraduates and the original authors, to a pool of participants with differing knowledge of poetry. The results showed that GPT-3 could produce passable (though perhaps not great) poetry, better than the average undergraduate and often not distinguishable from completions written by the original authors except by experts (Contreras Kallens et al., 2024). Future work will explore the possibilities of human-LLM co-creation of text for various purposes (within the purview of the Center for Contemporary Cultures of Text).

A third line of research that has just begun to take shape will investigate the possible impact of LLMs on the fundamental human skill of literacy. Learning how to read and write has important primary effects on perception and cognition as well as secondary effects on knowledge acquisition. Yet, literacy currently appears to be in decline in many parts of the world, and it seems very possible that a widespread deployment of LLMs might exacerbate this trend. However, there is also a chance that LLMs can help turn things around—but more research is needed to reveal potential pitfalls and help steer future literacy development in the right direction (see Huettig & Christiansen, 2024, for some initial suggestions).
These three lines of research are complimented by additional research conducted with Pablo Contreras Kallens (Saarland University) and Ross Kristensen-McLachlan (Aarhus University) to understand the nature of meaning in LLMs and the role of human feedback in achieving human-like language performance (see separate summary by Pablo Contreras Kallens). Altogether this work on LLMs aims to gain new insights into the positive and negative implications of LLMs on textual culture.

Bill Cope, Mary Kalantzis, & Patrick Bolger: The Language of Generative AI: Theory and Practical Application
The recent success of large language models (LLMs) in generating useful, well-formed responses to a wide variety of human prompts has both captivated and frightened educators. The potential for this technology to contribute to our scientific understanding of how humans produce text is intriguing, though (we would argue) currently confounded by the idea that machine and human intelligence, and their representation of that intelligence in language, are on a similar or even identical continuum. In this presentation, we will speak both theoretically, from the point of view of philosophy of language, and practically, about an application we have developed and applied in the higher education context.

Theoretically, both LLMs and humans design coherent textual meaning from a vocabulary whose elemental units of construction are tokens and morphemes, respectively. Because the target activity of both machine and person is human meaning, the boundaries of both token and morpheme are defined semantically. Tokens are collocations of characters represented in Unicode, and beneath that, binary notation. Their boundaries have been hand-crafted by humans to align with morphemes. Trivially, the machine has no ability to mean. It is unable to operate semantically in the human sense of interest-oriented, world-directed human meaning. It cannot self-create meaningful tokens. Nor can it in any sense “understand” them. More profoundly, the manner of the construction of humanly understandable meaning is quite unlike that of the human mind.

Our theoretical proposition is this: On the ground of tokens, Generative AI works statistically and empirically; on the ground of morphemes, human language works grammatically and theoretically. From a corpus of billions of words (if we may, for the sake of comparison, lapse from tokens and morphemes to a linguistic vernacular), by weighting the relations of a word to surrounding words, the machine builds a vocabulary that is massive in scale. Effectively, it attributes many meanings to the same word. Take these three sentences: “She walked to work.” “She walked the dog.” “She walked the prisoners to their cells.” These are three different kinds of walked, the first direction-oriented, the second for the sake of walking (or, straining at the leash, is the dog walking her?), the third a forced walking. By virtue of the word collocation, these three different kinds of walk happen to be different—though the machine cannot “know” this of course. The machine might generate a new sentence which, by a process of reiteration, uses the right kind of walk in the right kind of way. But this is a purely empirical, word-to-word achievement. Only a machine could achieve this through the tedium of statistical work grounded in the last analysis in binary notation. The results are extracted by empirical brute force (Kalantzis and Cope 2024).

Humans can’t know anywhere near as many such “words,” let alone when all are reduced to Unicode, then reduced again to binary notation. They can, however, produce coherent, novel sentences through the theoretical mechanism of grammar. Parts of speech (to use the old school notion) can substitute for each other as elementary kinds of meaning that are patterned into states, events, actions, and such like. Grammarians can distinguish the different kinds of walk with notions of transitivity, case, voice, and mood, for example. In this sense, grammar serves as a theory of everyday life. Even when we don’t have the technical words for it, our minds operate on the nuances of grammar, making theoretical connections across morphemes. Through the combinatorial power of grammar, humans can create entirely new meanings in never-before-said sentences. Lest this sound Chomskyian, this is a semantic process, not purely syntactic. Rather, it is an interest-oriented activity where language is inseparable from its pragmatic context. This is quite unlike LLMs which will forever be limited, having been cut adrift, siloed into written text because that is all they can process. And to the extent that Generative AI is multimodal, it is derivatively so, where the non-textual artifacts that it produces are captive to textual labelling and animated only by textual prompts (Cope and Kalantzis 2023). Going beyond language-centered accounts of meaning we've been working on a semantic multimodal, “transpositional” grammar that extends the systemic-functional paradigm in linguistics into other forms of meaning-in-context (Cope and Kalantzis 2020, Kalantzis and Cope 2020). Now we are hoping with this to layer a metaontology over the empiricism of LLMs.

Practically, Mary and Bill have been building Generative AI applications for education and applying them in our Learning Design and Leadership masters and doctoral program at the University of Illinois. Patrick has been working with us to process and interpret a mountain of empirical data. Our base platform has been CGScholar (Common Ground Scholar). With this tool, we capture multimodal student projects—typically about 5,000 words plus video, infographics, and other media embeds. After a first draft, the student requests the AI both to review and to offer feedback on their work. We do this with multi-pass prompt engineering—ten highly elaborated, highly abstract prompts about empirical representations, theoretical frameworks, critical analyses, practical applications and such like (Cope and Kalantzis 2015). The LLM is supplemented, via retrieval-augmented generation (RAG), with a vector database containing both all our graduate students’ work and all instructors' academic publications (35m tokens) for the past five years. Then we connect to ChatGPT via API for reviews of student work. But our ambition is to break free from this as soon as open-source models make this feasible. We also ask the students to moderate the AI with peer review. Since the beginning of 2023, the application has been through eight versions, across 25 classes, reviewing 524 major student projects (Saini et al. 2024, Tzirides et al. 2023, Tzirides et al. 2024, Zapata et al. 2024).

Frankly, we have been shocked at the quality of feedback, and so have our graduate students. AI feedback has offered more extensive and finely-targeted formative suggestions for revision than would ever be feasible to us as instructors. The run time is about 10 minutes per student work, and costs about one US dollar per run. Moreover, if in these trials we have been working at the highest end of the education continuum, the downstream implications are enormous. With a new grant from the Institute of Education Sciences, we will begin trials in middle and high schools this fall.

We are now working on another major release of the CGScholar software. This time the AI will not only be used to review and assess the finished work, it will also help the students produce it. One of the challenges here is to track when the AI is producing the work rather than the student. Naturally, the AI should only serve to help the student, not act in their stead. Because we capture keystrokes in CGScholar, we are going to build prompting into CGScholar to create two parallel texts: (1) the work which is the subject of instruction, and (2) a “Cyber-social Learning Log” which tracks AI interaction. Students will naturally use AI to help them with their work—these days they would be silly not to. But teachers can strike a deal: say 50% (or 70%, whatever) of the final submitted work must be the student’s own intellectual effort. We will track this by logging keystrokes, analyzing their patterns, and recording time on task.

On the theoretical front meanwhile, we want to test whether we can make Generative AI more reliable by supplementing it with the structures of ontologies or grammar-like meta-ontologies.

Pablo Contreras Kallens: Getting aligned: Exploring the effect of feedback on LLMs
The astonishing performance of Large Language Models (LLMs) on text prediction and generation has opened the door for using them to understand how humans learn and process language. However, in contrast to previous work using neural networks in psycholinguistics, researchers normally have little control over the architecture and training regime of the models they use. Thus, to explore the influence of specific factors in language, the strategy has been to appeal to model comparison. This has included, for example, models of different sizes [1] and amount of input during training [2].

An underexplored dimension of model variation is the technique that, along with the explosion in
model sizes, has been one of the drivers of the performance and utility of recent LLMs: Reinforcement Learning from Human Feedback (RLHF) [3]. In this type of training, after the model has been trained on next-token prediction on the corpus, it is fine-tuned using Reinforcement Learning. The reinforcement policy is derived from the ratings provided by humans on responses to queries. The network thus learns to produce output that maximizes the rating that it would get from the sample of raters used in the fine-tuning procedure. Thus, the model gets “aligned” with the preferences of humans through explicit feedback on the quality of its responses.

In this collaborative project, with Ross D. Kristensen-McLachlan (Aarhus University) and Morten H. Christiansen (Cornell University/Aarhus University), we explored the consequences that the RLHF procedure has on the behavior of LLMs when assessed in psycholinguistic experiments. This is important because the process of alignment of an individual human’s knowledge of language with the expectations of other humans in interactive contexts during learning has been a contentious topic in the language sciences. Specifically, generativists proposals assumed during the 20th century that feedback, particularly negative feedback, was minimal to nonexistent [4]. Therefore, the only source of input for the learning process would be the language input from others. And, because this stimulus was further assumed to be impoverished, it meant that knowledge of language had to be largely innate [5].
However, more recent research has suggested that explicit feedback, including negative feedback is common during language learning, particularly as reformulations of children’s productions [6]. This phenomenon could fill in the gap between relatively simple demonstrations of the power of statistical learning from the input and the acquisition of complex grammatical structures. For example, Frinsel et al. [7] recently found that feedback allows participants to learn more complex artificial grammars than what is possible with mere exposure. Our project proposes that a comparison between next-token prediction-trained models and equivalent models fine-tuned with RLHF can be taken as an implementation of claims about the effect of feedback during statistical-learning based acquisition.

The focus of our project is not on benchmark-driven differences in performance, but on assessing how closely LLMs capture phenomena of human language processing, and thus on how we can more adequately understand these models in cognitively relevant ways. This, in turn, allows us to delineate the limits of the use of LLMs as models of psycholinguistic processing. For this, we gathered three different experimental datasets that included detailed per-item information about both the stimuli presented to human participants and their responses. We then evaluated the output distribution of Vanilla and RLHF LLMs prompted with instructions that mirrored those presented to human participants as closely as possible. The datasets are a subject-verb agreement task [8], a sentence acceptability task [9], and an event knowledge task [10]. Then, we compared the differences in how each model captured the patterns of the participants.

When testing models of the GPT-3 family (davinci, or “Vanilla”, and text-davinci-003, or “Chat”), we found that the RLHF fine-tuning substantially changed the model’s behavior, making it more aligned with human responses. For example, in the Agreement task (Figure 1-A), we found differences in the performance of the models within each condition type. The Vanilla model struggles the most with trials including stimuli with a relative clause (“The advisor who directed the student”). In contrast, the Chat model has a substantially lower performance in trials with mismatching numbers where the first noun is singular, both in the prepositional (“The tire-eater in the carnival sideshows”) and the relative clause (“The actor who directed the films”) conditions. This difference exists despite marked improvements of the Chat model over the Vanilla model in all other trial types. This is precisely the type of trial in which human performance is also at its worst. That is, the Chat training reduced the performance of the model in the trial types in which human participants struggle the most. Another interesting finding is that in both the acceptability and the event knowledge tasks, the ratings obtained from the Chat model are positively correlated with measures of variance, in contrast to the Vanilla model (Figures 1-B and 1-C). This suggests that the distance between the ratings provided by the Chat model and those provided by humans grows as the differences between humans themselves increase.

The results obtained from the comparisons in our project have two important consequences. First, in the domain of psycholinguistics, they reinforce the idea that feedback from humans in a (quasi) interactive setting has substantial effects over and above the knowledge that can be derived from the input by tracking statistical patterns and dependencies. In this sense, the differences between Vanilla and Chat models could be interpreted as the effect, in the context of communication, of feedback on how the statistical patterns are used once they are acquired. Our project found that human behavior in these psycholinguistic tasks is approximated better by a model that includes a process of alignment with the expectations of other members of the linguistic community on top of accurately capturing the dependencies in the input.

Second, our project invites a reconsideration of what the most appropriate interpretation of LLMs in the context of language learning and processing is. Particularly, by analyzing the models’ output distribution, beyond their most-probable response to the input or the probabilities assigned to each stimulus, we found that the Chat models can capture sample-level variance, both in magnitude and direction. Thus, treating LLMs as models of individual language users might obscure the extent to which they capture significant patterns in language. This strategy risks labeling some observed behaviors as failures of the models to capture cognitive processes, when they might be successfully capturing the phenomenon at a higher level of organization. Instead, our project suggests that LLMs should be interpreted as models of the intersubjective cultural patterns that shape the trajectory of individual learning trajectories to different degrees, giving rise to the homogeneity-among-heterogeneity in language use within a language community [11]. That is, LLMs could be construed as capturing the community-wide cultural attractors of language processing and learning.

Kyle Mahowald: Language Models and Novel Content Generation
Today’s large language models generate coherent, grammatical text. This makes it easy, perhaps too easy, to see them as “thinking machines”, capable of performing tasks that require abstract knowledge and reasoning. I will draw a distinction between formal competence (knowledge of linguistic rules and patterns) and functional competence (understanding and using language in the world). I argue that language models have made huge progress in formal linguistic competence---which has important implications for linguistic theory. Even though they remain interestingly uneven at functional linguistic tasks, they can distinguish between grammatical and ungrammatical sentences in English, and between possible and impossible languages. As such, language models can be an important tool for linguistic theorizing.

In making this argument, I will draw on a project studying language models and constructions, specifically the A+Adjective+Numeral+Noun construction ("a beautiful five days in Austin"). I will show a series of experiments where train small language models on human-scale corpora, systematically manipulating the input corpus and pretraining models from scratch. I will discuss implications of these experiments for human language learning. More broadly, this line of work points toward a powerful method that today’s LLMs make available and that I am optimisitic will lead to insights for how humans generate text and produce language more generally: training models that learn interesting structure and controlled data and finding out what aspects of data are crucial for learning what phenomena.

While it is tempting to suggest that this work tells us only about models and not humans, I suggest that much work in linguistics theory has worked by focusing on claims about what can or can’t be learned, questions about whether some Structure X is underlyingly the same as Structure Y, or what kinds of representations are needed to learn Structure Z. These claims are often made by analyzing and comparing language data (corpora, human judgments, etc.), without necessarily linking these accounts to mechanistic cognitive or neurocognitive processes. Just as these works have been informative, it is possible that LLMs can be informative as well for understanding how humans produce and generate text.

In considering the workshop’s second question “What might happen in human writers' minds as they collaborate with AI (for example, ChatGPT) when producing texts?”, I turn to work that I did with Harvey Lederman in philosophy as to whether language models have beliefs. We start by defining bibliotechnism, the idea that LLMs are cultural technologies, like books or libraries, with no beliefs of their own. We extend and develop this idea showing how it can make sense of the fact they can still generate novel content, which get their meaning derivatively. We then pose another puzzle for bibliotechnism: the problem of novel reference, whereby a model is asked to refer to something (e.g., a diagram) that it just made up. This seems like a serious challenge for bibliotechnism, and we propose thinking about LLM beliefs using Dennett’s intentional stance. According to the intentional stance, a system has beliefs if and only if its behavior is well explained by the hypothesis that it has such beliefs. I will consider this position for LLMs and what it means for human agents working collaboratively with LLMs as writing or coding assistants.

Martin Pickering: Augmented language production
Imagine I have a really good LLM assistant to help me when I speak. I do not wish to replace myself, but rather have an assistant that helps me speak better, in real time. I might want to give someone time-critical instructions, for example how to perform a surgical operation that I am overseeing remotely. I need my assistant to prompt me with words I cannot access in time, to help me decide what to speak about next, to correct my mistakes and assist when I hesitate, and so on. If I had such an assistant, how would it change what goes on in my head as I speak? And how might we describe the human-assistant system?

Most theories explain language production in terms of a series of representations concerned with the meaning, grammar, and sound of words and sentences (e.g., Levelt, 1989). But they are primarily concerned with monologue, when there is no input from an addressee. Augmented language production involves an (artificial) co-producer, and is therefore a form of dialogue. Pickering and Garrod (2021) analyse dialogue in terms of a system in which interlocutors can make joint contributions (when I finish your utterance), can indicate failure to understand, or can suggest a revision or correction on-line, and this is successful because they predict what each other is going to do and align their representations both linguistically and in terms of their understanding of the situation under discussion. The representations that they construct are similar to those in monologue, but the input is partly provided by the addressee. For example, Levelt assumes that a speaker “self-monitors” what they are saying using comprehension and feeds any error signal back to their production system, but in dialogue the addressee “other-monitors” the speaker, so that the control loop involves both interlocutors. Other-monitoring is likely to be successful when the addressee has a fairly similar mental state to the speaker, because the addressee’s contribution will point out flaws in what the speaker has said, rather than indicate not knowing what the speaker is talking about. That is, dialogue is most successful when the interlocutors are fairly well aligned. In human-LLM dialogue, of course one interlocutor is replaced by an LLM, and dialogue will be most successful if the LLM develops similar representations to the human.

Our primary concern is with collaborative text production, and here the process of language production is quite different from speaking. The writer prepares a draft but can revise that draft as much as necessary before “publishing” it by making it public (as an article, email, SMS or whatever). During this process of drafting and redrafting, an co-author can contribute, and an LLM can do so as well. It can provide feedback about readability, ambiguity, structure of the argument, and so on. In addition, it can predict what I should talk about next (most likely, at the level of the next issue to discuss rather than just the next word) and make suggestions. This amounts to feed-forward monitoring (as discussed in Pickering & Garrod, 2013). In both cases, alignment of the discourse model is critical, so that it interprets and predicts in a similar way to the writer. Of course, this is only one way in which LLMs can be used in text production (an alternative is to get the LLM to generate an initial draft) – but it is important because conceptualization (i.e., message generation) remains with the human.

Co-composition requires alignment to be successful. Humans are aligned when they share relevant background knowledge and use of the language. Such alignment will of course be greater if they are both experts in the relevant domain. But they will also share linguistic representations, both at an item-specific level (e.g., possible meanings of a word) and at an architectural level – for example, conceptual representations involving features, an unordered representation of grammatical functions, or a representations of constituent structure but not lexical items. I propose that an LLM will aid composition better if it has the same representations as the writer, so that they are aligned. For example, if a human uses a homonym with a particular meaning, the LLM will align on that meaning (and won’t produce text in which the homonym has a different meaning).

In recent work, we investigated the extent to which LLMs (ChatGPT and Vicuna) behave similarly to humans when taking part in 12 well-known psycholinguistic experiments. We found that ChatGPT (and Vicuna to a slightly lesser extent) behaved similarly to humans in language (text) comprehension and production (Cai et al., 2023). They associated unfamiliar words with different meanings depending on their forms, continued to access recently encountered meanings of ambiguous words, reused recent sentence structures, attributed causality as a function of verb semantics, and accessed different meanings and retrieved different words depending on an interlocutor’s identity. ChatGPT (but not Vicuna) nonliterally interpreted implausible sentences that were likely to have been corrupted by noise, drew reasonable inferences, and overlooked semantic fallacies in a sentence. However unlike humans, neither model preferred using shorter words to convey less informative content, nor did they use context to resolve syntactic ambiguities. The results of these studies suggest that current models often construct similar representations to humans, but not always, and in turn that they would serve as reasonable, though not ideal, assistants for text production.

Why do we assume that these models have human-like representations at all? The answer is that we simply test them in just the same way that psycholinguists have tested humans for over half a century. To identify the representations in more detail, we have recently conducted a battery of structural priming studies to investigate the processing and representations of syntax in ChatGPT. Like humans, ChatGPT showed long-lasting abstract structural priming and a distance-sensitive lexical boost, didn’t fully suppress an incorrect syntactic analysis when completing a target preamble after a prime containing temporary syntactic ambiguity, repaired the syntax of implausible sentences to arrive at a plausible meaning, and showed within- and between-language structural priming, with translation-equivalent boost to between-language priming that was smaller than the lexical boost in within-language priming. Tense or aspect overlap did not enhance structural priming (like humans), though number overlap did enhance it (unlike humans). We can use these findings to determine the nature of linguistic representations in ChatGPT, just as we and others have done with humans. But our immediate concern is rather simpler – we just want to determine whether ChatGPT constructs the same representations as humans. If this is the case, then we are on the road to developing a good LLM assistant.
Returning to text production, the argument is that a human creates a draft and can re-draft it before publication. Such redrafting can be successful because the human understands what they have written and constructs the same levels of representation that they constructed during initial drafting. An LLM can also be successful to the extent that it does essentially the same thing as a human. In fact, the LLM is better being somewhat naïve – that is, not knowing too much about the writer’s conceptualization, as a reader will not know this either. Essentially, it should behave in the same way that the writer would if the writer did not know what specific message they were trying to convey.

What this all means is that we should try to establish LLMs as assistants in this manner – being able to comment on drafts, or “jump in” during composition and help at points of difficulty or when prompted to do so. We then need to develop LLMs that share the representations used by writers in general and by this writer in particular, so that their contribution to the process of co-composition is most beneficial.