Leveraging Large Language Models in Engineering Education Research: Methods and Applications
Plain Language Summary
Large language models (LLMs), the AI systems behind tools like ChatGPT, Claude, and Gemini, have become common in research work, yet many engineering education researchers lack a structured sense of where they fit and where they do not. This chapter walks through the research process one phase at a time, from early brainstorming and literature searching through survey design, transcription, interviewing, data analysis, and writing, describing what researchers are actually doing with these tools at each point and what can go wrong. A recurring concern is that the polished commercial products most people use hide the details, such as which model, which settings, and what information it was given, that reviewers and readers would ordinarily expect a methods section to report. We argue these tools can genuinely support research, though only when researchers treat how they used them as something to document and disclose rather than assume.
Contribution
Rather than cataloguing LLM capabilities in the abstract, we contribute a phase-by-phase narrative synthesis scoped to engineering education research, spanning planning and ideation, literature gathering and review, data collection and analysis, and writing and reporting, that treats prompt and context engineering as core methodological commitments rather than incidental implementation details and introduces a Small-q/Big-Q/"Confused-q" diagnostic for the epistemological mismatches now common in LLM-assisted qualitative analysis.
Research Questions
- How can engineering education researchers work with LLMs in ways that preserve transparency and reproducibility, given that widely used proprietary systems manage model details, configurations, and context on researchers' behalf?
- What applications, cautions, and recommendations accompany LLM use across the phases of the research process, from planning and ideation through writing and reporting?
- When researchers apply LLMs to qualitative data analysis, particularly frameworks rooted in Big-Q traditions such as reflexive thematic analysis, are the resulting procedures epistemically consistent with the frameworks they claim to instantiate?
Methods
This is a narrative, conceptual review rather than an empirical study; we synthesize literature about LLMs' technical operation, implementation, and results alongside commentary drawn from our own professional experience working with these systems. We organize the discussion around the broad and entwined phases of quantitative and qualitative research described by Baur (2019) and Onwuegbuzie and Leech (2005), namely planning and ideation, literature gathering and review, data collection and analysis, and writing and reporting, portraying them in Figure 1 as mutually feeding rather than strictly linear, with LLM applications "spinning off" at different moments. Throughout, we illustrate claims with worked examples, such as prompting Claude Sonnet 4.5 to diagnose double negatives and ambiguity in survey items, and with close readings of published LLM-assisted methods, such as CollabCoder's three-phase co-coding workflow, consolidating applications, cautions, and recommendations for each phase in Table 1.
Key Findings
We find the terms circulating in this space, "AI," "GenAI," "LM," and "LLM," obscure more than they specify, and that popular proprietary systems compound the problem by managing model details, configurations, and context on users' behalf; omitting those details, we argue, is akin to moving directly from Introduction to Discussion, leaving readers wondering about the provenance of things. Across phases, we locate three simultaneous relationships researchers hold with these systems, treating LLMs as objects of study, as tools in use, and as epistemic entities shaping what is known, and we contend the resulting positionality and reflexivity concerns are central features of LLM research requiring continuous appraisal rather than defects to be resolved. Qualitative analysis carries the most acute risk: much published LLM-assisted work invokes Big-Q frameworks such as reflexive thematic analysis while operationalizing them through Small-q means (inter-rater reliability benchmarking, systematic code generation, treating LLM-produced code groupings as thematic outputs), a mismatch we term "Confused-q" and suggest is more accurately characterized as LLM-assisted content analysis. We further observe this mixing is fugitive, since LLM outputs resemble qualitative reasoning while deriving from probabilistic token prediction, which allows model-generated codes to acquire the same epistemic weight as human-generated ones unless researchers articulate the two contributions separately.
Implications
We recommend researchers maintain a robust audit trail of LLM use, covering who used it, the model and its configurations, the intent of the interactions, the strategies employed, and how outputs were integrated, which can likely be tailored to the inconsistent disclosure policies of EER venues and federal agencies without requiring separate practices for each; we likewise suggest conferring with institutional review boards before and during LLM use, particularly given the sensitivity of educational data. Because scaling qualitative data analysis with LLMs scales the underlying research assumptions as well, we call for greater attentiveness to the alignment between the paradigmatic commitments of chosen frameworks and the procedural values embedded in their implementation, alongside normalized open-access repositories and published prompt and context engineering procedures. We flag the equity questions this raises, namely the uneven computational resources available to researchers and the language barriers of conducting EER with U.S.-trained, English-dominant models, as priorities for future methodological scholarship.
