According to the chapter, what are the limitations of large language models (LLMs) in the context of cybersecurity, and why does the author believe threat hunting remains a uniquely human endeavor?
LLMs are limited because they are trained on language data to complete sentences, not to reason generally; they reproduce training data and cannot generate truly original thought or generalize well beyond that data. In cybersecurity, high-quality datasets for known attacks are scarce and zero-day or evolving attacks have no training data, so human background knowledge, judgment, intuition, innovation, and adaptability still cannot be replicated. Threat hunting remains uniquely human because it means constantly looking for nonobvious or new attacks using experience-based intuition and deep knowledge not available to train models.
The chapter argues that LLMs are not few-shot threat hunters. Their capabilities are domain-specific to language, and even there they have clear limitations. LLMs are trained to fill in missing words and complete sentences, not to act as general-purpose problem solvers or reasoners. They can reproduce their training data and direct analogs of it, but they should not be expected to produce truly original thought, reasoning, or intelligence. Because they depend on large, high-quality training datasets, they perform well only on tasks within the bounds of the well-known and do not generalize to situations that deviate far from their training data or fall outside the tasks they were trained to perform. In the cybersecurity and threat-hunting space, obtaining a high-quality dataset with known attacks is very difficult, and training data for new attack types such as zero-day exploits does not exist. Threat actors continuously advance their attack strategies to bypass detection, which makes previously collected training data outdated. Researchers also find it very hard to imitate with AI the human traits that cybersecurity requires: background knowledge, sense of judgment, innovation, and adaptability. This difficulty arises partly from the lack of appropriate training data and training methods that do not revolve around solving cybersecurity problems. Threat hunting exemplifies problems beyond automated solution. Threat-hunting professionals continuously search for nonobvious exploitations or new attacks, which requires in-depth background knowledge that is not available to train models and intuition that only develops with experience. Therefore, the author believes threat hunting remains, for the foreseeable future, a uniquely human endeavor.
Key points
- LLMs are language-focused deep neural networks trained to predict and complete text, not to reason generally.
- LLMs can only reproduce training data and direct analogs, so they do not produce truly original thought or robust reasoning.
- Cybersecurity lacks high-quality training data for known attacks and has no training data for zero-day or evolving attacks.
- Human traits such as background knowledge, judgment, innovation, and adaptability are very hard for AI to imitate.
- Threat hunting requires continuous search for nonobvious or novel attacks using experiential intuition not available to train models.
- The author concludes that LLMs may assist human threat hunters but cannot yet perform the entire job independently.
AI for Cybersecurity_ Research and Practice
Unknown
John Wiley & Sons, Inc.