Leveraging Large Language Models in Engineering Education Research: Methods and Applications
Plain Language Summary
This chapter offers engineering education researchers a guided tour of how large language models (LLMs, the AI systems behind tools like ChatGPT and Claude) can be woven into the research process, walking from early brainstorming through data collection, analysis, and writing up results. Rather than treating LLMs as a single, settled technology, we distinguish between the underlying model itself and the particular chatbot, app, or custom setup a researcher happens to be using, arguing that this difference shapes what a model can do and what a researcher can responsibly claim about it. We pair this practical walkthrough with attention to the ethical stakes of using these tools on human-subjects data, and we close by proposing that researchers keep a running, disclosed log of how, when, and why they used an LLM at each stage of a project.
Contribution
Rather than surveying LLM capabilities in the abstract, we contribute a phase-by-phase (planning and ideation, literature review, data collection and analysis, and writing and reporting) narrative synthesis for engineering education research specifically, organized around a "technical" versus "configured" LLM distinction and a Small-q/Big-Q/"Confused-q" framework for diagnosing epistemological mismatches in LLM-assisted qualitative analysis.
Research Questions
- How should engineering education researchers describe and disclose their use of LLMs, given that a model's behavior depends as much on its surrounding configuration (interface, prompts, pipeline) as on its underlying technical architecture?
- What is the range of viable applications, and accompanying ethical and methodological cautions, for incorporating LLMs across the phases of the EER research process, from ideation through writing?
- When engineering education researchers apply LLMs to qualitative data analysis, particularly frameworks rooted in Big-Q traditions such as reflexive thematic analysis, are the resulting procedures epistemologically consistent with the frameworks they claim to instantiate?
Methods
As a narrative, conceptual review rather than an empirical study, we synthesize literature spanning LLM technical operation, research methodology, and AI ethics, organizing our discussion around Baur's (2019) and Onwuegbuzie and Leech's (2005) phases of the research process (planning/ideation, literature review, data collection/analysis, writing/reporting). We ground our argument in an interactionist reading (Blumer, 1969; Charmaz et al., 2019) of how LLMs participate in human meaning-making, and in Maton's (2003) reflexive framework for tracking social, objectifying, and epistemic relations across a project. Throughout, we illustrate points with our own professional experience prompting LLMs (e.g., testing Claude Sonnet 4.5 on survey item refinement) and with close readings of specific published examples of LLM-assisted methods (e.g., CollabCoder's three-phase co-coding scheme).
Key Findings
We argue that the common shorthand of "using an LLM" collapses two things worth separating: the LLM as a technical object (a statistical model trained on a corpus) and the LLM as it is embedded within a particular configured system (a chatbot, an API call, a retrieval pipeline), with the latter carrying most of the methodological weight researchers must account for. Engaging directly with Jowsey et al.'s (2025) three objections to GAI in qualitative research, we contend that LLMs are not incapable of participating in meaning-making so much as they transform the human meaning-making systems they are embedded within, which shifts, rather than resolves, the ethical stakes of their use. Surveying LLM applications across ideation, literature review, survey design, data generation, transcription, interviewing, and quantitative and qualitative analysis (summarized in Table 1), we find qualitative analysis carries the most acute risk: much published LLM-assisted thematic analysis work invokes Big-Q frameworks like reflexive TA while operationalizing them with Small-q procedures (e.g., inter-rater reliability, systematic code generation), a mismatch we term "Confused-q" and argue is more accurately described as LLM-assisted content analysis.
Implications
We recommend that researchers maintain a running, disclosed audit trail of LLM touchpoints throughout a project (model and configuration details, prompting strategies, points of human verification) that can flex across the inconsistent GAI disclosure policies of different EER publication venues, journals, and funders. We call for the field to develop shared quality criteria for LLM-assisted qualitative work specifically, given how unevenly current studies align their procedural choices with the epistemological commitments of the frameworks they claim to use, and we flag unresolved equity concerns around computational access and the English-language dominance of widely used models as priorities for future methodological scholarship.
