What are the limitations of using LLMs as classifiers in qualitative research, and how can researchers mitigate these limitations?
The main limitations are sensitivity to prompt wording and example order, the constraint on the number of examples imposed by context size, and output variability caused by the model's stochastic nature. Researchers can mitigate these by running multiple executions, validating against manually labeled ground truth, treating LLM classifications as preselectors of relevant data rather than fixed categories, and combining or testing different LLMs and classification techniques.
The chapter identifies several limitations when using LLMs as classifiers in qualitative research. First, classification results can be highly sensitive to the exact wording of the prompt and the order in which examples are presented; small changes can lead to different outputs. Second, only a limited number of examples can be included because the model's context size restricts how much information can be provided. Third, because LLMs are stochastic, multiple executions of the same prompt can yield different classifications, which can compromise reliability and validity if unaddressed. To mitigate these limitations, the chapter recommends validation techniques adapted from supervised learning. Researchers can create a large labeled data sample as ground truth and measure how accurately different prompts classify data across several executions; this requires manually labeling a substantial amount of data, but it enables evaluation of the classification outputs. Multiple executions of prompts should be performed to generate reliable classifications. The authors also advise using LLMs as complementary tools rather than infallible classifiers. In particular, Garcia Quevedo et al. (2025) suggest using LLM classification capabilities as preselectors of relevant data when working with large datasets, so that researchers focus on the most relevant material instead of treating classifications as fixed categories. Further mitigation strategies include testing various classification techniques across different LLMs, establishing and validating fine-tuned LLMs for specific tasks such as sentiment analysis, integrating LLMs with other NLP techniques, and adopting a balanced approach that selects methods aligned with the specific research objectives and the nature of the data.
Key points
- Classification is sensitive to the exact wording of prompts and the order of examples.
- The number of examples is limited by the model's context size.
- Stochastic variation means different executions of the same prompt can produce different classifications.
- Use a large manually labeled sample as ground truth to measure classification accuracy across prompts and executions.
- Run multiple executions of prompts to improve reliability.
- Treat LLMs as complementary preselectors of relevant data, not as infallible fixed classifiers.
- Test different LLMs and classification techniques, including fine-tuned models, and align methods with research goals.
Related questions
AI for Qualitative Research: A Hands-On Guide for Management Scholars
Diana Garcia Quevedo
Palgrave Macmillan