It’s just going toward sunset, sitting on my boat watching the gators chase the mullet, and I’m thinking about my company’s staff. Just after dinner, when executives are checking their phones before spending time with their families, is the perfect time to send an email that hits all the marks: something they know, something important requiring action, and a lure set deep with misdirection. Now add to that synthetic vishing from the help desk or impersonation from a vendor, and I’m left considering a job as an outboard mechanic.
I keep wondering how long until somebody creates an entire synthetic company for phishing, vishing, and generally attacking companies. It would have a website, phone numbers, email, and everything, maybe even be registered as a company. When you interact with it, you are only interacting with AI. The CEO? AI persona. Help desk? CISO? Sales people? All AI personas could appear on video, take phone calls, and deliver the same general talk track needed to build trust. That may be in the future, but it is how I came up with the idea.
Threat Model
The adversary has not learned a new trick. Social engineering is the same con it has always been: establish trust, create urgency, extract value. What changed is the cost of running it and defending against it.
IBM X-Force put a number on the shift in a 2023 experiment, and IBM’s 2025 Cost of a Data Breach Report repeated it. A convincing phishing email that took a skilled operator sixteen hours to research, draft, and customize now takes five minutes with a generative AI tool. That change collapses the economics of deception. The attacker who could afford to target three executives per day can now target three hundred, and each message reads like it came from someone who knows the recipient’s job, their current projects, and the name of their boss.
Researchers at Harvard, including Fred Heiding and Bruce Schneier, tested fully automated AI spear phishing on human subjects in 2024. The AI emails drew a 54% click rate, the same as emails written by human experts, against 12% for a control group of generic phishing. The AI produced each one for about four cents. KnowBe4 found that 82.6% of the phishing emails it analyzed between September 2024 and February 2025 showed some use of AI. Cofense tracked malicious emails bypassing secure email gateways at a rate of one every 19 seconds through 2025, more than double the rate from the prior year.
But email is only the setup. The real threat model for 2026 is multimodal.
McAfee researchers produced a voice clone with an 85% match from about three seconds of sample audio using freely available tools. Every earnings call your CFO has ever done, every conference keynote your CEO recorded, every podcast appearance anyone on your leadership team sat for is training data sitting on the open internet. CrowdStrike’s 2025 Global Threat Report tracked a 442% surge in vishing attacks during the second half of 2024. By Q1 2025, security firm Right-Hand reported that deepfake vishing attempts had jumped 1,633% over the prior quarter. In a December 2025 piece published by Fortune, deepfake researcher Siwei Lyu argued that voice cloning had crossed what he called the “indistinguishable threshold,” the point at which the average listener cannot reliably tell a cloned voice from a real one.
Video adds the final channel. In January 2024, a finance employee in the Hong Kong office of UK engineering firm Arup joined a video conference with people he believed were his CFO and several senior colleagues. Every face on the call was an AI-generated deepfake. Every voice was synthetic. He authorized 15 wire transfers totaling $25.6 million before anyone realized the entire meeting had been fabricated. That case set the operational template that attackers refined through 2025 and into 2026. In a Gartner survey of 302 cybersecurity leaders published in September 2025, 62% of organizations reported a deepfake attack in the prior year, 43% reported at least one deepfake audio call incident, and 37% reported deepfakes in video calls.
The attack is no longer a single phishing email. It is a coordinated campaign across email, voice, and video that builds synthetic trust through multiple channels simultaneously. Palo Alto’s Unit 42 found that 36% of its incident response cases between May 2024 and May 2025 began with social engineering, and more than a third of those used techniques other than phishing, including SEO poisoning, fake system prompts, and help desk manipulation.
And the synthetic company I was sitting on my boat thinking about is not as far off as it sounds. The pieces are already in use. In May 2024 the U.S. Department of Justice charged a scheme in which North Korean IT workers used more than 60 stolen or borrowed American identities to land remote jobs at more than 300 U.S. companies, generating at least $6.8 million. Gartner projects that by 2028, one in four candidate profiles worldwide will be fake. The infrastructure to build an entire synthetic company capable of conducting long-term social engineering against your vendor management process or your accounts payable team already exists. It just has not been packaged and pointed at your organization yet.
What the attacker needs to run this operation is remarkably modest. Access to a commercial LLM or one of the criminal variants like GhostGPT or FraudGPT, which run about $150 to $200 per month. A voice cloning tool, several of which are available as commercial APIs. A deepfake video service, increasingly sold as Deepfake-as-a-Service. Open source intelligence on the target company, which LinkedIn, corporate websites, press releases, and SEC filings provide for free. And a domain that looks close enough to pass a quick glance.
What the attacker produces that a defender could observe is equally specific. The email originates from a newly registered domain. The linguistic patterns in the message do not match the supposed sender’s historical writing. The communication request breaks the normal pattern of who contacts whom, when, and about what. The urgency in the message does not map to any known business process. And the verification channel the attacker suggests is one the attacker controls.
Those observables exist. The question is whether your defenses are looking for them.
Traditional Tools
Most small and mid-sized companies have built their social engineering defenses around three pillars: a secure email gateway, a security awareness training program, and a set of verification procedures for sensitive actions like wire transfers.
The secure email gateway is the workhorse. Products from Microsoft (Defender for Office 365), Proofpoint, Mimecast, Barracuda, and Cisco sit between the internet and your users’ inboxes. They scan inbound email for known malicious URLs, suspicious attachments, sender reputation anomalies, and content that matches patterns in threat intelligence databases. They check SPF, DKIM, and DMARC records to verify that the sending domain authorized the sending server. In my experience, enterprise gateways stop the overwhelming majority of bulk phishing by volume. Mimecast reported flagging more than 9.13 billion threats in the first nine months of 2025 alone.
That block rate is real, and the tools deserve credit for it, but the risk sits in what gets through. When Cofense measured what got through, they found that one malicious email reached a protected inbox every 19 seconds in 2025. Cofense found that conversational attacks, text-only messages with no links or attachments, made up 18% of the malicious emails it saw in 2025, and that these messages bypass most security controls. The volume of email is so large that even a fractional miss rate delivers a steady stream of well-crafted attacks to your users’ inboxes every day.
Security awareness training is the second pillar. KnowBe4, Proofpoint Security Awareness, Cofense PhishMe, and similar platforms send simulated phishing emails to employees, track who clicks, and deliver training modules to the people who fall for them. The results are measurable. KnowBe4’s 2025 benchmarking study of 67.7 million simulated phishing tests across 14.5 million users and 62,400 organizations showed that the average baseline phish-prone percentage starts at 33.1% and drops to 4.1% after 12 months of consistent training. That is an 86% reduction, so training works.
What training teaches people to look for is the problem. The standard curriculum was built around visual indicators: misspelled sender addresses, poor grammar, mismatched URLs, suspicious formatting, unexpected attachments. These indicators presume that the attacker is constrained by language fluency, time, and the cost of producing a polished message. When the email is grammatically perfect, contextually appropriate, written in the recipient’s native language with their industry’s terminology, and references a real project the recipient is working on, the visual indicators that training taught people to spot simply are not there.
The third pillar is verification procedures. Dual-approval requirements for wire transfers. Callback policies that require a phone call to a known number before executing a payment change. Manager approval workflows for sensitive access requests. These controls are often the most effective part of the stack because they force a human decision point outside the attacker’s channel. The policy says: before you wire money, call the person who requested it on a number you already have, not a number from the email.
DMARC enforcement rounds out the technical controls. When configured at p=reject, DMARC prevents attackers from sending email that spoofs your exact domain. EasyDMARC’s 2026 analysis of the top 1.8 million domains found that fewer than a quarter enforce DMARC at quarantine or reject. The rest remain open to exact-domain spoofing, which means an attacker can send email that genuinely appears to come from your CEO’s address.
These tools were not badly designed. They were designed for a specific adversary.
The Gap
Every control in the traditional stack rests on a set of assumptions about the attacker’s operational constraints. Those constraints were valid for twenty years. They stopped being valid sometime around 2024.
The secure email gateway assumes that malicious emails contain detectable artifacts: a known bad URL, a suspicious attachment, content patterns that match threat intelligence, or sender authentication failures. This worked when the attacker’s constraint was effort. Crafting a phishing email that contained none of these artifacts required a skilled operator investing significant time per target. AI removed that constraint entirely. An attacker can now generate a text-only email from a freshly registered domain, with no URLs, no attachments, and no content matches in any threat intelligence database, that reads like a routine business communication from a trusted contact. The gateway scans it, finds nothing suspicious, and delivers it. A February 2026 Abnormal AI blog post on email threats bypassing secure email gateways made the architectural point directly: SEGs protect infrastructure, not people. Attacks that carry no malicious payload, only persuasive text from a seemingly legitimate sender, pass through by design.
Business email compromise is the clearest expression of this gap. A BEC message contains nothing technically malicious. It carries no malware and no exploit, only a well-written email from what appears to be a known contact, requesting a routine business action that happens to route money to the wrong account. The FBI’s 2024 IC3 report logged $2.77 billion in BEC losses across 21,442 incidents. The gateway cannot catch this because there is nothing for it to catch.
Security awareness training assumed that the attacker’s constraint was quality at volume. A human attacker could write a convincing email or a sloppy one, but not a thousand convincing ones in an afternoon. The training program was designed to help employees spot the sloppy ones, which were the overwhelming majority. When the attacker’s production constraint disappears, every email arrives polished. The visual tells that training taught people to find do not exist in the message. An employee who correctly applies everything the training program taught and sees no red flags will still click, because by every observable standard the email looks legitimate.
Verification procedures assumed the attacker operated through a single channel. The callback policy works when the attacker’s only reach is email. You get a suspicious wire request, you pick up the phone, you call the CFO at a number you already have, and the CFO says, “I never sent that.” But when the attacker can also clone the CFO’s voice and answer that callback with a synthetic voice that sounds exactly right, or join a video conference as a deepfake that looks and sounds like three of your colleagues, the verification step itself becomes the attack surface. The Arup case proved this. The employee was suspicious and did what he was supposed to do by getting on a video call to verify. Everyone on the call was synthetic, and the verification step became the mechanism that overcame his skepticism.
The underlying constraint that all of these tools relied on was the cost of producing trust. Trust used to be expensive for an attacker. Researching a target took time. Writing convincingly in the target’s language required skill. Sustaining a multi-channel deception required real infrastructure and human labor over weeks or months. Those costs served as a natural rate limiter on the volume and sophistication of social engineering attacks.
AI collapsed them simultaneously. An attacker can now pull LinkedIn profiles, company announcements, and SEC filings through OSINT scrapers into a structured target brief in minutes. The same attacker writes in any language with native fluency, scales from one target to a thousand without adding staff, and sustains a synthetic relationship across email, phone, and video without deploying a single human being. The total cost is about what a streaming subscription runs.
Your gateway was calibrated for an adversary who left detectable artifacts because the effort was expensive. Your training was calibrated for an adversary who made visible mistakes because maintaining quality at scale was impossible. Your verification procedures were calibrated for an adversary who operated through a single channel because multi-channel deception required too many people. None of those calibrations hold in 2026.
AI Augmentation
The defender’s response has to operate at the same layer as the attack. The attack has shifted from content to behavior, and the defense has to follow.
Communication Baseline and Behavioral Anomaly Detection
The most immediately useful thing you can build with AI and existing infrastructure is a behavioral baseline of your organization’s communication patterns. This does not require a new platform. It requires access to your email logs and an LLM you can query via API.
Your mail server already records metadata for every message that enters and leaves your organization: sender, recipient, timestamp, subject line, and for Microsoft 365 or Google Workspace environments, the message body itself is available through management APIs. Start by pulling 90 days of email metadata for your highest-risk users, which means anyone who can authorize payments, approve access requests, or change vendor banking information. Build a profile for each user that captures who they normally communicate with, when those communications typically happen, what subject patterns are normal, and how frequently they exchange messages with each contact.
When a new inbound email arrives for one of those users, run the message metadata against the baseline. An email from a contact who has never written to this user before, on a topic that does not match their normal work, arriving at an unusual time, requesting an action involving money or access, and using urgent language should generate an alert that goes to the user and their manager before the user acts on it.
The build needs a Python script pulling from the Microsoft Graph API or Google Workspace Admin SDK, storing metadata in a local database, and scoring new messages against baseline distributions. The LLM component adds the ability to assess the text of the message itself against the supposed sender’s historical writing patterns. Feed the model ten emails from the real sender and ask whether the new message matches their typical vocabulary, sentence structure, and tone. The comparison produces a confidence score you can act on.
The main dependency risk is alert fatigue. If this system generates false positives at high volume, users will learn to ignore it. Set the alerting threshold high enough that analysts only see messages with multiple behavioral anomalies, not just one. If the LLM component goes down or produces unreliable results, the metadata baseline still functions independently. If an attacker learns about the system and attempts to match the real sender’s communication patterns precisely, the metadata layer catches anomalies that content-matching cannot, because the attacker still has to send from an unusual source at some point in the attack chain. The manual fallback for a total system failure is a standing policy: any request involving money, access changes, or sensitive data that arrives via email requires a verbal confirmation on a pre-established phone number, not a number provided in the message. This policy should exist whether or not the AI system is running.
Inbound Communication Linguistic Analysis
A second layer examines the content of inbound messages for indicators that the text was generated or templated rather than written by the claimed sender. This is distinct from the behavioral baseline. The baseline asks “is this communication pattern normal?” The linguistic analysis asks “does this text match how this person writes?”
This works best on correspondents your organization has history with. For known vendors, partners, and internal colleagues, maintain a writing sample corpus of their past messages. When a new message arrives from that correspondent, use the LLM to compare the new message against the corpus for consistency in word choice, sentence length distribution, formality level, and characteristic language habits. Most people have writing fingerprints they are not aware of, specific phrases they reuse, punctuation habits, paragraph structures they default to. A message that deviates significantly from the historical pattern warrants a flag.
For unknown correspondents, the analysis shifts to detecting AI-generated content. This is harder and less reliable, but it adds signal. LLM-detection scoring tools can estimate the probability that a text was machine-generated. The score alone is not actionable, but combined with other behavioral indicators, it contributes to a risk assessment. An email from a first-time contact, requesting a financial action, written in text that scores high on AI-generation probability, arriving on a Friday afternoon, warrants more scrutiny than any one of those indicators alone.
The dependency risk here is that the LLM’s assessment of writing style may be wrong, producing both false positives and false negatives. This layer should never be the sole basis for blocking a message. It feeds into a composite risk score alongside the behavioral baseline, and the composite score determines whether the message gets an analyst review or a user warning. The manual fallback is the same callback procedure. When in doubt about any communication, pick up the phone and call a known number.
Real-Time Verification Workflow Automation
The third component targets the verification gap directly. Instead of relying on users to remember and execute callback procedures under pressure, automate the trigger.
When the behavioral baseline and linguistic analysis together flag a message above a defined risk threshold, the system should automatically initiate a verification workflow. This means sending a notification to the claimed sender through a separate, pre-authenticated channel: a Slack message, a Teams notification, or an automated text to a registered phone number. The notification asks a simple question: “Did you send an email to [recipient] about [subject] at [time]?” If the claimed sender confirms, the message is cleared. If they deny it or do not respond within a defined window, the message is quarantined and an analyst is notified.
This removes the cognitive burden from the target. The employee does not have to remember the callback policy, does not have to decide whether the email feels suspicious, and does not have to take action under the emotional pressure the attacker is deliberately creating. The system does the verification automatically, through a channel the attacker does not control.
Building this requires connecting your email security layer, your messaging platform, and a simple workflow engine. A webhook from the email scoring system to a Slack or Teams bot, a database to track verification requests and responses, and a quarantine action in your email system. For Microsoft 365 environments, this connects through the Microsoft Graph and Defender for Office 365 APIs. For Google Workspace, through the Gmail API’s modify and trash functions, or admin quarantine rules in the Workspace console.
The dependency risk is that an attacker who compromises the verification channel defeats the system. If the attacker has access to the target’s Slack account as well as their email, the verification request goes to an account the attacker controls. Mitigate this by requiring verification through a channel that uses a different authentication factor than the one the message arrived on. If the email came through the corporate account, send the verification to a personal mobile number registered during onboarding. The manual fallback for a verification system failure is the same as the standing policy: direct phone call to a number from the corporate directory, not from the message.
Voice and Video Channel Monitoring
The hardest part of this threat model to defend is the voice and video channel, and the honest answer is that small teams cannot yet deploy reliable real-time deepfake detection on live calls. The technology exists in research labs and in some enterprise products from vendors like Pindrop and GetReal Security, but it is not something a two-person security team is going to stand up in a week.
What a small team can do is change the process around high-risk voice and video interactions. Establish a policy that no financial transaction, access change, or sensitive data transfer may be authorized solely on the basis of a phone call or video conference, regardless of who appears to be on the call. Every such request must also be confirmed through a text-based channel with strong authentication. This means that even if an attacker produces a perfect deepfake of your CEO on a video call requesting an emergency wire transfer, the transfer does not execute until the CEO also confirms via a digitally signed message or an authenticated request in your financial approval system.
This is a process control, not a technical control. It works because it forces the attacker to compromise two independent systems with different authentication mechanisms, which substantially increases the cost and complexity of the attack. It is also the only control that remains fully effective when every other system fails.
The dependency risk for process controls is human compliance. People under pressure from someone they believe is their boss will skip the process. The mitigation is making the process frictionless for legitimate requests, so that following it is faster than bypassing it, and making clear that the policy exists because of documented attacks where the process would have prevented millions in losses. The Arup case is the teaching example. The employee did everything right except require confirmation through an independent authenticated channel. That one missing step cost $25.6 million.
The work here is unglamorous: scripts and APIs and metadata databases and workflow automations built on tools you already own. But it operates at the layer where the attack now lives, which is the layer of trust and behavior rather than the layer of malicious content. Your gateway will keep catching the bulk phishing. Your training will keep teaching people to report suspicious messages. What these additions provide is the ability to catch the attacks that produce no artifacts your existing tools were designed to find, because the adversary who produces those attacks is no longer constrained by the costs that made them rare.
Works Cited
Abnormal AI. (2026, February 1). 12 email threats bypassing SEGs in 2026: What your gateway doesn’t see. https://abnormal.ai/blog/email-threats-bypassing-segs
Abnormal Security. (2025, January). GhostGPT: Uncensored AI chatbot used by cybercriminals. https://abnormalsecurity.com/blog/ghostgpt-uncensored-ai-chatbot
Cofense. (2026, February 4). Cofense report reveals AI-powered phishing accelerated to one attack every 19 seconds [Press release]. https://cofense.com/Blog/Cofense-Report-Reveals-AI-Powered-Phishing-Accelerated-to-One-Attack-Every-19-Seconds
CrowdStrike. (2025). CrowdStrike 2025 global threat report: Executive summary. https://www.crowdstrike.com/en-au/resources/reports/global-threat-report-executive-summary-2025
EasyDMARC. (2026). DMARC adoption and enforcement report 2026. https://easydmarc.com/blog/ebook/dmarc-adoption-report-2026/
Federal Bureau of Investigation, Internet Crime Complaint Center. (2025). 2024 internet crime report. https://www.ic3.gov/AnnualReport/Reports/2024_IC3Report.pdf
Gartner. (2025, July 31). Gartner survey shows just 26% of job applicants trust AI will fairly evaluate them [Press release]. https://www.gartner.com/en/newsroom/press-releases/2025-07-31-gartner-survey-shows-just-26-percent-of-job-applicants-trust-ai-will-fairly-evaluate-them
Gartner. (2025, September 2). Why CIOs can’t ignore the rising tide of deepfake attacks [Press release]. https://www.gartner.com/en/newsroom/press-releases/2025-09-02-why-cios-cannot-ignore-the-rising-tide-of-deepfake-attacks
Heiding, F., Lermen, S., Kao, A., Schneier, B., & Vishwanath, A. (2024). Evaluating large language models’ capability to launch fully automated spear phishing campaigns: Validated on human subjects (arXiv:2412.00586). arXiv. https://arxiv.org/html/2412.00586v1
IBM. (2025). Cost of a data breach report 2025. https://www.ibm.com/reports/data-breach
KnowBe4. (2025). New KnowBe4 report reveals a spike in ransomware payloads and AI-powered polymorphic phishing campaigns [Press release]. https://www.knowbe4.com/press/new-knowbe4-report-reveals-a-spike-in-ransomware-payloads-and-ai-powered-polymorphic-phishing-campaigns
KnowBe4. (2025, May 13). KnowBe4 report reveals security training reduces global phishing click rates by 86% [Press release]. https://www.businesswire.com/news/home/20250513295204/en/KnowBe4-Report-Reveals-Security-Training-Reduces-Global-Phishing-Click-Rates-by-86
Lyu, S. (2025, December 27). 2026 will be the year you get fooled by a deepfake, researcher says. Fortune. https://www.fortune.com/2025/12/27/2026-deepfakes-outlook-forecast
McAfee. (2023). Beware the artificial impostor. https://www.mcafee.com/content/dam/consumer/en-us/resources/cybersecurity/artificial-intelligence/rp-beware-the-artificial-impostor-report.pdf
Mimecast. (2025). Global threat intelligence report 2025: January to September 2025. https://www.mimecast.com/resources/ebooks/threat-intelligence-january-june-2025/
Mitrade. (2025, September 4). Deepfake voice phishing drives $20M+ in losses as crypto execs targeted [Reporting Right-Hand Cybersecurity data]. https://www.mitrade.com/insights/news/live-news/article-3-1092996-20250904
Nair, P. (2023, July 26). Criminals are flocking to a malicious generative AI tool. GovInfoSecurity. https://ciso2ciso.com/criminals-are-flocking-to-a-malicious-generative-ai-tool-source-www-govinfosecurity-com
Palo Alto Networks Unit 42. (2025). 2025 Unit 42 global incident response report: Social engineering edition. https://unit42.paloaltonetworks.com/2025-unit-42-global-incident-response-report-social-engineering-edition/
South China Morning Post. (2024, May 17). UK multinational Arup confirmed as victim of HK$200 million deepfake scam. https://scmp.com/news/hong-kong/law-and-crime/article/3263151/uk-multinational-arup-confirmed-victim-hk200-million-deepfake-scam-used-digital-version-cfo-dupe
The Associated Press. (2024, May 16). North Koreans stole American identities and took remote work tech jobs. Fortune. https://www.fortune.com/2024/05/16/north-koreans-stole-american-identities-and-took-remote-work-tech-jobs