A recent pilot study conducted at Brigham Young University indicates that individuals struggle to distinguish between phishing text messages generated by AI and those crafted by humans, particularly when the messages are personalized with work-related details. The study involved 25 volunteers who were presented with a mix of AI-generated and human-written text messages, tailored to their personal information.
Each participant provided details about their job, workplace, hobbies, city, and recent online posts. This information was used to create prompts for GPT-4, which generated six messages per person. An additional six messages per person were written by undergraduate students enrolled in a deception course. Participants were then asked to sort all twelve messages, identifying which ones they would be most likely to click and which they believed were AI-generated.
The study found that messages generated by GPT-4 were placed above the "click line" 28% of the time, compared to 21.3% for human-written messages. While this difference of 6.7 percentage points might suggest AI superiority, the small sample size means the study cannot definitively conclude whether AI messages are more effective than human ones. However, the researchers noted that the AI messages were produced with a single, short prompt, while the human-written messages were created by students who had received phishing instruction and whose output was screened by cybersecurity professors.
A significant finding was the heightened danger of work-related messages. Job-themed messages prompted a hypothetical click 38% of the time, substantially higher than hobby-themed messages (19%) and social media-themed messages (17%). This was the only comparison in the study that maintained statistical significance.
Despite the personalization, errors in specific details sometimes undermined the messages. For example, a participant dismissed a message referencing a non-existent coworker named "Mike," and another recognized that a named individual was not part of their university's instructional design team. Such inaccuracies were more detrimental than a lack of personalization, as they immediately raised suspicion.
Participants were largely unable to identify AI-generated messages. Across 300 judgments, their accuracy was 52%, barely better than random chance. Volunteers offered various theories, such as AI writing being too formal or too generic, or perfect grammar indicating a machine. Conversely, some interpreted typos as proof of human authorship. Interestingly, emojis were a strong indicator of AI-generated content (66% of GPT-4 messages had them versus 2% of human ones), yet only five participants mentioned them, and their interpretations varied.
In contrast to human intuition, a machine learning classifier trained by the researchers successfully distinguished between AI and human-written messages with 88.7% balanced accuracy. This classifier was tested under strict conditions, including standardization of URLs, removal of emojis, and case/digit/punctuation flattening. However, the researchers cautioned that this classifier was trained on a specific dataset and its ability to generalize to other models or prompt designs is unproven. Furthermore, research suggests that paraphrasing AI-generated text can significantly reduce the effectiveness of such detectors.
The study's limitations include the small sample size and the fact that messages were presented as printed cards, lacking the real-world context of a buzzing phone or live links. The human comparison group consisted of novice students, not professional social engineers. The exact GPT-4 snapshot and API logs used for message generation were also not recorded, preventing replication of the AI generation process.
The practical advice derived from the study emphasizes vigilance regarding sender identity, communication channel, link destinations, and the nature of the request. The study strongly suggests that attempting to discern whether a message "sounds like a robot" is an unreliable defense mechanism.






