When I was a tenured professor at Purdue in West Lafayette, I used to teach one of my passions: malware forensics. It was never so much the malware itself as it was dealing with pseudo-machine code and getting back to what drove me toward computer science in the first place. I always wanted to know how computers worked at the electron level. It took being in the Marines to finally learn that lesson, which is a strange place for it, but there I was in 29 Palms, California, at the Basic Electronics Course, learning how electrons make transistors work. A lesson I carried forward to sitting in front of a class of graduate students, wondering why I was staring off into space.
Now, sitting at the navigation station of my sailboat on my lunch break with 30 minutes to type this, I have to think that malware generated by AI is nuts. It is something I cannot even imagine trying to forensicate, and then again, is it even worth forensicating at this point? The attack cycle is so fast, so compressed, and now I am left wondering if there is some new direction defense has to go. I keep thinking this is where we put the human in the loop differently, where you get the speed of the machine and the unique insights of the human working the same problem at the same time. That is not easy.
Threat Model
The adversary in this chapter is not writing malware from scratch for every target. The adversary uses AI to rewrite existing malware continuously so no two copies look alike to a signature engine, a behavioral sandbox, or a static analysis tool.
In June 2025, Google’s Threat Intelligence Group identified a VBScript dropper called PROMPTFLUX that used a hardcoded API key to query Google’s Gemini model. The dropper’s “Thinking Robot” component queried Gemini for new VBScript obfuscation and evasion code. A later variant replaced it with a function that instructed Gemini to act as an expert VBScript obfuscator and rewrite the dropper’s entire source code every hour. The regenerated code was written to the Windows Startup folder for persistence and copied to removable drives and network shares for propagation. Each rewrite preserved the original payload, the API key, and the regeneration logic itself, creating a recursive mutation loop. GTIG assessed PROMPTFLUX as experimental, still in development, with some features commented out. Google disabled the associated assets. GTIG found no evidence it had compromised a victim network, but the architecture was complete.
A month later, in July 2025, Ukraine’s CERT identified LAMEHUG, an infostealer attributed with moderate confidence to APT28. LAMEHUG was built in Python, packaged with PyInstaller, and delivered through spear phishing emails impersonating Ukrainian ministry officials. What made LAMEHUG different from APT28’s prior tooling was that it contained no hardcoded commands for data collection. Instead, it called the Qwen 2.5 Coder 32B Instruct model through the Hugging Face API and asked the model to generate the system commands it needed at runtime. The commands it generated were for gathering system information and searching specific folders for documents. Because the commands were generated fresh by the LLM on each execution, no static command strings existed in the binary for a signature engine to match.
These two samples represent different approaches to the same operational problem. PROMPTFLUX mutates its own code to defeat static analysis. LAMEHUG generates its operational commands at runtime to defeat behavioral pattern matching. Both eliminate something that detection tools have relied on since antivirus was invented: the assumption that malicious code contains identifiable, repeatable patterns.
The concept is not new. HYAS Labs demonstrated it in 2023 with BlackMamba, a proof-of-concept keylogger that used OpenAI’s API to generate polymorphic keylogging code at runtime. The generated code ran entirely in memory through Python’s exec() function and was never written to disk. HYAS tested BlackMamba against what they described as an industry leading EDR product, repeatedly, and recorded zero alerts or detections. The keylogger collected credentials and exfiltrated them through Microsoft Teams, blending its network traffic with legitimate collaboration traffic.
What changed between 2023 and 2025 is that the concept moved from a security researcher proof of concept to state-sponsored operational deployment. APT28 did not publish a paper about LAMEHUG. They used it against Ukrainian government targets during a war.
For the CISO at a 1000-person company running CrowdStrike or SentinelOne or Defender for Endpoint, the threat model looks like this: an attacker writes or obtains a payload that works. Before delivering it, the attacker runs it through an LLM with instructions to rewrite the code while preserving functionality, changing variable names, restructuring control flow, altering string encoding, and modifying the execution sequence. The attacker tests the output against the same detection tools the target runs. If the payload triggers a detection, the attacker feeds the detection output back to the LLM and asks for another rewrite. This loop runs until the payload clears. The whole process takes minutes, not days. The attacker needs an LLM API key, a copy of the detection tool or a reasonable proxy for it, and a basic script to automate the loop.
The infrastructure cost is negligible. A Hugging Face API key is free for low volume use. Running a model locally on consumer hardware is feasible with quantized versions of open weight models. The attacker does not need to be a skilled malware developer. The attacker needs to be a competent prompt engineer with a working payload to start from.
Traditional Tools
Most organizations in the 250 to 1500 employee range have some version of the following stack deployed against malware.
Endpoint protection is the centerpiece. For companies that have invested in modern tooling, this means an EDR agent from CrowdStrike Falcon, SentinelOne, Microsoft Defender for Endpoint, Sophos Intercept X, or one of their competitors. These products combine signature matching against known malware hashes with behavioral analysis that monitors process execution chains, file system activity, registry modifications, network connections, and memory operations. The behavioral component is what separates EDR from the old antivirus model. When a process spawns PowerShell, which then reaches out to an external IP and downloads a binary that injects into a running service, the EDR agent is trained to flag that chain regardless of whether it recognizes the specific binary involved.
Signature updates come from the vendor’s threat intelligence feed, typically refreshed multiple times per day. When a new malware sample is identified in the wild, the vendor generates a signature and pushes it to all deployed agents. The speed of this update cycle varies by vendor, but even the fastest operate on a timeline measured in hours between identification and global deployment.
Sandboxing sits behind the email gateway or the web proxy at many organizations. When an attachment or download arrives, the sandbox executes it in an isolated environment and watches what happens. Does it reach out to a command-and-control server? Does it attempt privilege escalation? Does it modify registry keys associated with persistence? Sandbox engines from vendors like Palo Alto WildFire, FireEye (now Trellix), or the sandbox components built into Microsoft Defender apply a combination of live execution and static analysis to render a verdict.
Email security gateways from Proofpoint, Mimecast, Microsoft Defender for Office 365, or Barracuda filter inbound mail for known malicious attachments and links. They maintain their own signature databases, and some apply basic behavioral analysis to attachments before delivery.
SIEM platforms collect logs from all of the above plus network devices, identity systems, and cloud services. At a well-run mid-sized company, the SIEM correlates EDR alerts with network flow data, authentication events, and DNS queries to provide context around detections.
This stack catches a large volume of commodity malware. It catches known attack tools. It catches payloads that reuse infrastructure or code from previous campaigns. It catches attackers who use techniques that are common enough to appear in the vendor’s training data. The tools were not designed badly. They were designed for a threat environment where creating a new, unique malware variant required skilled human labor and measurable time.
The Gap
The constraint that made all of this work was the cost of novelty.
Writing a new malware variant that evaded a specific detection engine used to require a developer who understood both the target detection logic and the low-level mechanics of the operating system well enough to restructure the payload without breaking it. Polymorphic malware, metamorphic engines, and server-side polymorphism all existed before AI, but each of these required custom engineering. The attacker had to understand packing, encryption routines, code permutation, and the specific heuristics that the target EDR used to identify suspicious behavior. That work took days to weeks per variant, and the skill required limited the number of attackers who could do it.
Signature-based detection assumed that malware reuse was inevitable because rewriting was expensive. Behavioral analysis assumed that the behavioral patterns themselves would remain recognizable across variants because restructuring behavior, not just code, was even more expensive than restructuring syntax. Sandboxing assumed that the payload would exhibit its malicious behavior within the observation window because attackers could not cheaply test against the sandbox beforehand to calibrate their evasion timing.
AI collapsed all three costs at once. An attacker with a working payload and access to an LLM can produce a functionally identical variant with different code structure, different string signatures, and different execution flow in the time it takes to run an API call. The attacker does not need to understand the detection logic. The LLM handles the restructuring. The attacker does not need to understand packing or encryption. The LLM generates novel obfuscation on each pass. The attacker does not need days of testing against a sandbox. The attacker can run dozens of variants through a local copy of the target EDR in an afternoon and keep the ones that clear.
SonicWall documented over 210,000 never-before-seen malware variants in 2024 alone, which SonicWall put at 637 a day. The CrowdStrike 2026 Global Threat Report found that 82 percent of detections in 2025 were malware-free, driven by credential theft and living off the land techniques that bypass file-based detection entirely. The Sophos 2025 Active Adversary Report documented a 126 percent year-over-year increase in the abuse of living-off-the-land binaries. When attackers moved from their own tools to the tools already on the box, it was partly because those tools do not trigger signature hits, and partly because AI made it cheap to find and chain the right native binaries for each target environment.
CISA validated this gap directly. In a 2024 red team assessment of a U.S. critical infrastructure organization, the agency found that the target’s EDR solutions detected only a few of the red team’s payloads across both Windows and Linux environments. In one case, the EDR did block an initial phishing payload and generated an alert, but the network defenders did not read or respond to the alert. The red team succeeded by avoiding known bad detections and, in one technique, inflating file sizes above the EDR’s upload threshold so the files were never sent to the cloud for analysis. A legacy environment in the same organization had no EDR at all, and the red team persisted there undetected for months.
The Akira ransomware group demonstrated a different form of this gap. When EDR quarantined Akira’s initial payload on a Windows endpoint, the attackers pivoted to a Linux-based webcam on the same network, a device that no EDR agent could run on. From the webcam, they mounted SMB shares and encrypted the network from a device that was invisible to the endpoint detection stack. In February 2026, Symantec and Carbon Black researchers reported that the Reynolds ransomware embedded a vulnerable signed driver directly in the ransomware payload, using the bring-your-own-vulnerable-driver technique to terminate EDR processes from CrowdStrike, Cortex XDR, Sophos, and Symantec before the ransomware began encrypting.
The gap is not that EDR is broken. The gap is that EDR was engineered around the assumption that generating novel, evasion-tested payloads required enough time and skill to keep the volume manageable. That assumption held for twenty years and no longer does.
AI Augmentation
The defender’s response to AI-accelerated evasion is to use the same AI capability the attacker uses but turn it against your own detection stack before the attacker does.
There are two things a small team can build on top of what they already have.
Anomaly detection that does not depend on known patterns. The signature and behavioral models in your EDR were trained on data the vendor collected. When a payload is novel enough that it falls outside that training data, the EDR misses it. What the EDR still provides, even when it misses the payload itself, is telemetry. Process creation events, network connections, file writes, registry modifications, and authentication events all still flow into your SIEM or your EDR’s cloud console.
The augmentation is a script or a set of API calls that pulls that telemetry and runs anomaly detection against your environment’s own baseline rather than against the vendor’s global model. You are not looking for “this matches a known attack.” You are looking for “this has never happened before in our environment.” That means flagging a process that has never previously spawned from a particular parent, or a user account authenticating from a host it has never touched before. It means noticing when a binary writes to a directory where nothing has written before, or when an endpoint opens a network connection to a destination no machine in your environment has ever contacted.
The build starts with your existing log infrastructure. Export a rolling window of process creation events, network connections, and authentication events from your EDR or SIEM into a data store you control. A simple time series database or even a structured set of CSV files will work for a prototype. Write a script that calls an LLM API or runs a local model to compare each new event against the historical baseline and flag statistical outliers. The model does not need to know what malware looks like. It needs to know what your environment looks like on a normal day, and it needs to tell you when something deviates from that pattern.
The output is a ranked list of anomalies delivered to a Slack channel, an email, or a dashboard each morning. The analyst reviews the list. Most will be benign: a new software installation, a configuration change, an employee working unusual hours. Some will be worth investigating. This detection method does not care whether the payload was polymorphic, novel, or generated by AI. It cares whether the payload’s behavior produced events that are unusual for your specific environment.
If the anomaly detection relies on an external LLM API, the attacker could potentially deny service to it by triggering rate limits, or the API provider could have an outage at the worst time. The mitigation is to run the baseline comparison against a local model that does not depend on external connectivity. Quantized open-weight models run on a single GPU and can handle the comparison workload for a mid-sized environment. The manual fallback is a scheduled script that runs the same baseline comparison using statistical methods without an LLM, looking for standard deviations from the mean on event counts per process, per host, and per user. It will produce more false positives than the LLM driven version, but it will still catch gross anomalies.
Adversarial testing of your own detection stack. The second augmentation inverts the attacker’s workflow. Instead of waiting for an attacker to test payloads against your defenses, you do it yourself, continuously.
Take a known malware sample from a public repository like MalwareBazaar or VirusTotal. Feed it to an LLM with instructions to rewrite the code while preserving functionality, using the same kinds of prompts an attacker would use: change variable names, restructure control flow, alter string encoding, modify the execution sequence. Generate ten variants. Run each variant through your detection stack in an isolated test environment. Record which variants your EDR catches and which it misses.
For the variants that clear your detection, analyze what changed between the original and the variant. Did the LLM restructure the process execution chain enough to escape the behavioral model? Did it change the network callback pattern? Did it modify the persistence mechanism? The answers tell you exactly where your detection has blind spots, and they tell you before an attacker finds those same blind spots.
This requires a test environment that mirrors your production endpoint configuration. A single virtual machine running your EDR agent with the same policy set as production is enough to start. Write a script that automates the variant generation by calling an LLM API, drops each variant into the test VM, waits for the EDR verdict, and logs the result. Run this on a weekly cadence. When you find variants that bypass detection, write custom detection rules in your SIEM or EDR that target the specific behavioral gap the variant exploited, and add those variants to your next test cycle to verify the new rules catch them.
The dependency risk here is that the test environment must accurately mirror production. If your test VM runs a different EDR policy version, different OS patch level, or different agent version than production, the test results will not transfer. The mitigation is to automate the test environment build from the same configuration management tools that build your production endpoints. The manual fallback is to submit the generated variants to your EDR vendor’s sandbox through their API and compare the vendor’s verdict against your local EDR agent’s verdict. Discrepancies between the two tell you something about your local configuration even without a dedicated test environment.
Running both of these augmentations together creates a feedback loop. The adversarial testing tells you which evasion patterns your detection misses. The anomaly detection catches those patterns in production by looking for behavior that deviates from your environment’s baseline rather than matching a known signature. When the adversarial testing finds a new blind spot, you tune the anomaly detection to weight that class of behavior more heavily. When the anomaly detection flags something in production that the adversarial testing did not anticipate, you add that pattern to your next round of variant generation.
Neither augmentation requires a new platform purchase. Both run on API calls, scripts, and existing log infrastructure. A single engineer can prototype the anomaly detection pipeline in a week using Python, an LLM API, and the export functions already built into whatever SIEM or EDR the organization runs. The adversarial testing requires more care in the test environment setup, but the actual variant generation and testing loop is a script that runs unattended once configured.
The malware forensics I taught at Purdue assumed you had time to sit with a sample, disassemble it, trace its execution path, and understand what it did. That assumption was already strained before AI entered the picture. With AI-generated variants appearing at machine speed, the forensic question shifts from “what did this specific sample do” to “what is happening in my environment that should not be happening.” The tools to answer that second question are within reach of every CISO reading this, and they do not require understanding malware at the electron level to deploy.
Works Cited
Cato Networks. (2025). Cato CTRL threat research: Analyzing LAMEHUG, first known LLM-powered malware with links to APT28 (Fancy Bear). https://www.catonetworks.com/blog/cato-ctrl-threat-research-analyzing-lamehug/
Cybersecurity and Infrastructure Security Agency. (2024, November 21). Enhancing cyber resilience: Insights from CISA red team assessment of a U.S. critical infrastructure sector organization (AA24-326A). https://cisa.gov/news-events/cybersecurity-advisories/aa24-326a
Dark Reading. (2023). AI-powered ‘BlackMamba’ keylogging attack evades modern EDR security. https://www.darkreading.com/endpoint-security/ai-blackmamba-keylogging-edr-security
Google Threat Intelligence Group. (2025, November). GTIG AI threat tracker: Advances in threat actor usage of AI tools. https://cloud.google.com/blog/topics/threat-intelligence/threat-actor-usage-of-ai-tools
BleepingComputer. (2025, March). Akira ransomware encrypted network from a webcam to bypass EDR. https://www.bleepingcomputer.com/news/security/akira-ransomware-encrypted-network-from-a-webcam-to-bypass-edr/
Sophos. (2025). 2025 Sophos active adversary report. https://www.sophos.com/en-us/blog/2025-sophos-active-adversary-report
SonicWall. (2025). Executive summary: 2025 SonicWall cyber threat report. https://sonicwall.com/resources/white-papers/executive-summary-2025-sonicwall-cyber-threat-report
Symantec and Carbon Black Threat Hunter Team. (2026, February). Reynolds: Defense evasion capability embedded in ransomware payload. Broadcom. https://security.com/threat-intelligence/black-basta-ransomware-byovdThe Hacker News. (2025, July 25). CERT-UA discovers LAMEHUG malware linked to APT28, using LLM for phishing campaign. https://thehackernews.com/2025/07/cert-ua-discovers-lamehug-malware.html