AskReference
ExplanationIntroductory

What is the purpose of the synthetic dataset provided with the book, and why was it created?

The synthetic dataset was created to enable hands-on learning with LLMs for qualitative analysis while avoiding the ethical and data-protection risks of using real, private, or proprietary social media data. It consists of 1,115 fictional social media posts from ten imaginary personas and is intended exclusively for educational practice with the book's code examples.

The book provides a custom dataset of 1,115 synthetic social media posts representing ten imaginary personas. It was generated specifically for ethical and data-protection reasons, so readers can practice and explore LLM applications in qualitative analysis without exposure to real-world private or proprietary data. The dataset is meant for educational purposes only and can be downloaded and used as input to the code examples in the accompanying GitHub repository. It is protected by the same copyright agreement as the book and is not intended for commercial use.

Key points

  • The dataset contains 1,115 synthetic social media posts from ten imaginary personas.
  • It was created to avoid the ethical and data-protection risks of using real private or proprietary data.
  • Its purpose is exclusively educational: to allow practice and exploration of LLM applications in qualitative analysis.
  • The dataset is used with sample code from the book's dedicated GitHub repository.
  • It is protected by the book's copyright and not intended for commercial use.
Source:AI for Qualitative Research: A Hands-On Guide for Management Scholars· Overview of Artificial Intelligence, Machine Learning, Natural Language Processing, and Large Language Models· p. 6–14
Cover of AI for Qualitative Research: A Hands-On Guide for Management Scholars

AI for Qualitative Research: A Hands-On Guide for Management Scholars

Diana Garcia Quevedo

Palgrave Macmillan

View this ebook