It was a Tuesday. It was snowing outside the gray building in a gray town, and I was sitting in a room without windows. Contemplating the idea that prisoners at least got to see the sun, I realized that my before sunrise and after sunset schedule meant another day working toward seasonal affective disorder, and I thought about quitting. Fortunately, one of my favorite people was walking through the door in a few minutes.
What saved my day? Traditional tools used by the enterprise were filtering out all the “noise” from attacks on the outside of the network because everybody everywhere was always attacking, so why worry about it was the question. My favorite person had used machine learning to gather all that noise and then created a set of scripts that identified attack patterns and could accurately predict the next attack based on the adversary’s precursor activity. They could track the work factor and work schedules, giving us a good idea of where they were working. Most importantly, they could learn the infrastructure that might be used for secondary or tertiary targeting.
Now I’m sitting here on my sailboat with a Blue Moon beer in my hand, typing all of this and smiling because, damn, that day was amazing. I’m wondering, given the structure, what that newly minted PhD could have done with AI to enable his analysis. Unfortunately, all of the attack vectors are now AI-enabled, so the task is even more extreme. A dribble of condensation falls on my knee, reminding me that beer is meant to be drunk. I’m still smiling.
Threat Model
The adversary doing reconnaissance against your company in 2020 was a human being sitting at a keyboard running tools one at a time. They would start with your domain name in crt.sh, pull every certificate ever issued, and extract the subdomains from the Subject Alternative Name fields. Then they would run those subdomains through DNS resolution to see what was live. Then they would run the live hosts through Shodan or Censys to see what ports were open. Then they would check your job postings on LinkedIn and Indeed to figure out your technology stack. Then they would search breach databases to see which of your employees had reused passwords. Then they would look at your GitHub repositories for hardcoded API keys and connection strings. Then they would map your vendor relationships from press releases and partnership announcements.
Each of those steps took time. Not a lot of time for any single step, but the total chain of correlation, the work of tying the DNS results to the breach data to the job postings to the certificate history into a coherent picture of where you are weak, that took a human researcher hours or days. For a company with 200 employees and a handful of external services, a thorough passive reconnaissance effort might take a skilled operator two full working days.
That time constraint was foundational to your defense. Your security program was designed around it even if nobody wrote it down. A human researcher working at human speed could run that chain against a limited number of targets. The economics of reconnaissance meant that your 1000-person Aerospace company or your 1500-person logistics firm was probably not worth the effort for most adversaries unless someone was specifically targeting you. The sheer labor cost of building a comprehensive attack plan from public sources created a natural threshold that filtered out a large portion of potential attackers. AI collapsed that labor cost.
An adversary can now feed your company name to an agent that pulls certificate transparency logs, resolves DNS, queries Shodan, scrapes LinkedIn, checks breach databases, reads your GitHub, and correlates the results into a structured attack plan. The same reconnaissance that took a human two days produces output in minutes. The fidelity of that output varies, and it still requires human judgment to prioritize, but the collection and initial correlation work is no longer the bottleneck.
What makes this different from simply running faster scans is the correlation layer. A certificate transparency query returns a list of subdomains. A Shodan query returns a list of open ports. A LinkedIn scrape returns a list of employee names and their listed skills. Individually, none of that is new intelligence. What the AI component produces is the synthesis: this staging server at staging.internal.yourcompany.com is running an outdated version of Jenkins (found via Shodan), and the engineer who manages it (found via LinkedIn) has an email address that appears in three separate breaches (found via Have I Been Pwned), and the server shares infrastructure with your production CI/CD pipeline (inferred from DNS records and certificate issuance patterns). That synthesis is what used to take the human researcher the longest. The AI produces it as part of the default output.
The constraint that changed is not the availability of the data. All of these sources have been public for years. Certificate transparency logs exist because the industry built them to detect fraudulent certificates. Shodan exists because researchers built it to index internet-connected devices. LinkedIn exists because professionals built profiles to advance their careers. Breach databases exist because someone aggregated the results of data breaches that already happened. The constraint that changed is the cost of turning all of that data into an actionable attack plan against a specific target.
The numbers tell the story of how much raw material is available. GitGuardian’s 2026 State of Secrets Sprawl report found 28.65 million new hardcoded secrets in public GitHub commits in 2025 alone, a 34 percent increase over the prior year. The Verizon 2026 DBIR found that vulnerability exploitation is now the leading initial access vector at 31 percent of breaches, up from 20 percent the year before, overtaking credential abuse for the first time in 19 years. Forrester found that attack surface management tools initially discover about 30 percent more cloud assets than security teams know they have, and EASM vendors report their own customer figures between 30 and 43 percent. All of that data is sitting in public sources waiting for someone, or something, to correlate it.
The attacker using AI for reconnaissance is not a nation-state with unlimited resources. This is not an exotic threat. This is a mid-tier criminal group that bought API access to an LLM and pointed it at your digital footprint. They are not smarter than the human researcher who came before them. They are faster, and they can run the same process against a thousand targets in the time the human spent on one.
Traditional Tools
A competent security program built five years ago and maintained on a moderate budget would have several tools that address parts of the reconnaissance problem.
External attack surface management is the category that comes closest to the defender’s version of this threat. Products from vendors like Censys, Palo Alto Cortex Xpanse, Microsoft Defender EASM, and CrowdStrike Falcon Surface continuously discover internet-facing assets associated with your organization. They find subdomains, cloud instances, exposed services, and expired certificates. When configured correctly, they provide an outside-in view of what your organization looks like from the internet. The better products will identify shadow IT and assets your security team did not know existed. For a company that deploys one of these, the tool does what it says: it discovers and inventories your external attack surface.
Vulnerability scanning is the workhorse. Products from Tenable, Qualys, Rapid7, and their open-source equivalents scan your known assets for known vulnerabilities and produce reports ranked by severity. Most companies run authenticated scans against internal assets on a regular schedule, often weekly or monthly, and unauthenticated scans against external assets on a similar cadence. The scanners compare what they find against the CVE database and produce a prioritized list. They are effective at finding what they are designed to find: known vulnerabilities on known assets.
Penetration testing, usually outsourced to a third-party firm, happens on a cycle. For most companies in the 500 to 1500 employee range, that cycle is annual or quarterly, sometimes driven by compliance requirements like PCI DSS or CMMC rather than by a risk assessment that called for that specific frequency. The pentest firm sends a team that spends a defined window, usually one to three weeks, attempting to find and exploit vulnerabilities. The team does their own reconnaissance as part of the engagement, writes a report, and delivers it. The findings get fed into the remediation pipeline.
SIEM platforms, whether Splunk, Microsoft Sentinel, Elastic, or one of the managed alternatives, collect and correlate log data. In most deployments at this scale, the SIEM ingests authentication logs, endpoint detection data, and some subset of network traffic. Firewall logs often make it into the SIEM, but the deny logs, the ones that show blocked connection attempts from external sources, frequently get summarized or sampled to control ingestion costs. A Sentinel deployment on pay-as-you-go pricing at $4.30 to $5.59 per gigabyte, depending on US region, has real incentive to reduce the volume of firewall traffic it ingests.
DNS logging sits in an odd place. External DNS queries, the ones that show who is looking up your domain names and subdomains, often exist only in the logs of your DNS provider or registrar. Internal DNS logs from Active Directory are more commonly collected, but even those are sometimes excluded from the SIEM to manage volume. Passive DNS services exist that record historical DNS resolution data, but few companies in this size range subscribe to them for their own monitoring.
These tools work for the problems they were built to solve: the EASM product finds unknown assets, the scanner finds known vulnerabilities, the pentest finds exploitable paths, and the SIEM correlates events across sources. Each tool carries an assumption about how fast the adversary moves through the reconnaissance phase, and until recently, that assumption was accurate.
The Gap
Every tool in the traditional inventory relies on one foundational assumption: the adversary’s reconnaissance takes longer than your detection and response cycle. Your quarterly pentest assumes the adversary spent enough time in the reconnaissance phase that the quarterly check is frequent enough to catch changes before they get exploited. Your EASM platform assumes that its daily or weekly scan cadence will discover new assets before an adversary maps them. Your vulnerability scanner assumes that the CVE-based detection window, the time between a vulnerability’s disclosure and your next scan, is shorter than the time an adversary needs to find that vulnerability in your environment and build an exploit path that uses it.
That assumption held when reconnaissance was labor-intensive. A human researcher taking two days to map your external footprint meant your daily EASM scan had a reasonable chance of finding the same things before the attacker finished their work. A human researcher taking a week to correlate your breach data with your LinkedIn profiles with your certificate history meant your annual pentest had a reasonable chance of testing the same attack paths before someone used them.
AI compressed the reconnaissance timeline from days to minutes, and the tools did not get worse. The assumption they were built on stopped being true.
Consider what happens to your EASM deployment. The product discovers your internet-facing assets based on seed data you provide: your primary domain, your IP ranges, your known subsidiaries. It works outward from those seeds. It is good at finding things connected to what it already knows about. What it does not do is correlate your EASM results with your employees’ breach history, your job postings’ technology disclosures, your GitHub repositories’ accidentally committed secrets, and your certificate issuance patterns into a single targeting package. The EASM product answers one question well: what is exposed? It does not answer the question an AI-equipped adversary is asking: given everything publicly knowable about this organization across all sources, where is the highest-probability path to initial access?
Your vulnerability scanner has the same structural problem. It runs on a cadence. Between scans, new services come online, configurations change, and new CVEs get published. The Verizon 2026 DBIR reports that only 26 percent of CISA Known Exploited Vulnerabilities were fully remediated in 2025, down from 38 percent the prior year, and the median time to full remediation rose from 32 days to 43 days. Those numbers describe a patching pipeline that was already struggling to keep pace with human-speed exploitation. Against AI-compressed reconnaissance that identifies the vulnerability, correlates it with the specific service in your environment, and packages it into an attack plan within hours of disclosure, the gap between scan cycles becomes the attacker’s window.
The penetration test is no different. The pentest team does a few days of reconnaissance as part of their engagement scope. They find what a skilled human can find in that window and they test it. But their reconnaissance is bounded by the same human constraints the old adversary operated under. They are not correlating your certificate transparency history with your employees’ credential exposure, your job postings’ tech stack disclosures, and your DNS patterns. They are testing a subset of the attack surface using the same tools and the same time constraints your adversary has now escaped.
The gap is not that your tools stopped working. The gap is that your tools were designed around a constraint that AI removed. Reconnaissance used to be the most labor-intensive phase of the kill chain. It was the phase where the economics of attack and defense were closest to balanced because both sides were spending human time. AI tipped those economics by making the attacker’s reconnaissance cost approach zero while the defender’s detection infrastructure stayed on its pre-existing cadence.
Your firewall deny logs are a perfect illustration. Every day, your edge firewall blocks thousands of connection attempts from sources scanning your perimeter. Most security teams treat those deny logs as noise because the firewall did its job. The connections were blocked. The conventional wisdom is that scanning is so constant and so undifferentiated that analyzing the blocked traffic produces no actionable intelligence. That conventional wisdom was formed when scanning was generic and reconnaissance was manual. When an adversary uses AI to automate targeted reconnaissance, the scanning that precedes it looks different. It comes from specific ranges, targets specific ports that match the services your job postings disclosed, and follows patterns that correlate with your recently issued certificates. The intelligence is in the deny logs. Almost nobody is looking at it.
AI Augmentation
The defensive play is straightforward to describe and feasible to build without a new platform purchase. You run the same AI-powered correlation against your own footprint that an adversary would run against you, and you do it on a continuous cycle. Then you use the data sources your organization already generates but currently discards to detect when someone else is running the same process against you.
The first piece is offensive self-reconnaissance on a continuous cycle. Build a script or a set of API calls that replicates what an adversary’s AI reconnaissance chain would do against your organization. Pull your certificate transparency history from crt.sh. Resolve every subdomain. Query Shodan or Censys for open ports and service banners on your IP ranges. Scrape your own job postings for technology stack disclosures. Check your domains against breach databases. Search public GitHub for any repositories that reference your domain, your internal hostnames, or your employee email addresses. Run this weekly, automatically. Feed the results to an LLM with a prompt that asks it to correlate the findings and produce a prioritized list of likely attack paths, ranked by probability of success based on the combination of exposure, credential availability, and service vulnerability.
The output is an attack plan written from the adversary’s perspective, using only the information the adversary would have. Compare it against what your EASM product found. Compare it against your last pentest report. The delta between what your tools found and what the self-reconnaissance correlation produced is your actual gap. That delta is almost always nonempty because no single existing tool does the cross-source correlation.
The implementation can start simple. A Python script that calls the crt.sh API, runs DNS resolution with dnspython, queries the Shodan API, and feeds the consolidated results to the Anthropic or OpenAI API with a structured prompt asking for attack path analysis. For a company with a single primary domain, a few dozen subdomains, and a small IP range, the whole thing runs in under ten minutes and costs a few dollars per execution in API calls. You can run it from a cron job on an existing server.
The second piece turns throwaway logs into reconnaissance detection. It uses data your organization already generates. External DNS query logs show who is resolving your subdomains. Firewall deny logs show who is scanning your perimeter. Failed authentication logs show who is testing credentials against your external services. Web server access logs show who is crawling your public-facing applications.
Most organizations either discard this data, sample it to reduce SIEM costs, or collect it but never analyze it because the volume is too high for manual review. AI changes the economics of analysis the same way it changed the economics of reconnaissance.
Build a pipeline that ingests your firewall deny logs, your external DNS query logs (available from most DNS providers via API), and your failed authentication logs from externally facing services. Feed batches to an LLM or a purpose-built clustering model with a prompt that asks it to identify patterns: Are any source IPs targeting ports that match the services disclosed in your job postings? Are any source IPs querying subdomains that map to your recently issued certificates? Are any source IPs attempting authentication against services that appear in breach databases with your employees’ credentials? Is there temporal correlation between the DNS lookups and the subsequent port scanning that suggests coordinated reconnaissance rather than background noise?
The LLM does not need to see the individual log lines in real time. You can batch the data daily. Export your firewall denies, your DNS queries, and your failed auth logs to a JSON file. Run them through a script that clusters by source IP and computes basic statistics: how many distinct ports did each source IP target, how many distinct subdomains did each source IP resolve, does the timing suggest sequential discovery or parallel scanning. Feed the clusters to the LLM and ask it to score them on a likelihood scale from background noise to targeted reconnaissance. The output is a short list of source IPs and IP ranges that warrant further investigation, with the reasoning visible in the response.
It puts to work data your SIEM is probably not ingesting and your EDR never sees because the activity happens entirely outside your network boundary. It turns logs that were considered disposable into an early warning system for targeted reconnaissance.
Now for what breaks when this fails. The self-reconnaissance pipeline depends on API access to external services (crt.sh, Shodan, breach databases) and API access to an LLM. If Shodan changes its API terms, you lose port visibility. If the breach database you query goes offline, you lose credential exposure data. If the LLM provider is unavailable, you lose the correlation layer. None of these failures are catastrophic because the underlying data sources remain accessible through manual queries. The manual fallback is a checklist version of the same process: a security analyst runs the same queries by hand, documents the results in a spreadsheet, and does the correlation manually. It takes a day instead of ten minutes, but it produces the same output. The automation makes it continuous. The manual process makes it survivable.
An adversary who knows you are running self-reconnaissance could attempt to poison the data sources. They could register subdomains on your domain that point to honeypots designed to waste your analysis time. They could inject false entries into breach databases (though the major aggregators have validation processes that make this difficult at scale). They could generate high volumes of scanning traffic from diverse source IPs to flood your reconnaissance detection with false positives, pushing real targeted reconnaissance below the noise floor. The mitigation for data poisoning is the same discipline that applies to any intelligence analysis: corroborate across sources, be skeptical of single-source findings, and calibrate your model against ground truth by periodically running the correlation manually and comparing it to the automated output.
The dependency risk you need to name explicitly is the LLM itself. If your correlation pipeline produces a false negative if it tells you everything looks clean when an adversary has already mapped your attack surface you miss the warning. The mitigation is to never treat the LLM’s output as a definitive security assessment. Treat it as a lead generator. Every output should route to a human who validates the top findings before they become action items or are filed as clean. The LLM speeds up the analysis, while the human keeps the judgment.
The person who walked through my door on that Tuesday built something from first principles, using the tools and data available to them. They looked at what everyone else was discarding and found the signal. The tools were primitive by current standards, the data was noisier, and the compute was slower. But the fundamental insight was sound: the adversary’s reconnaissance generates observable artifacts, and those artifacts contain patterns you can learn if you are willing to look at what everyone else treats as noise. The AI did not change the insight. It changed the speed at which both sides execute on it.
Works Cited
McDaniel, D. (2026). The state of secrets sprawl 2026: AI-service leaks surge 81% and 29M secrets hit public GitHub. GitGuardian. https://dev.to/gitguardian/the-state-of-secrets-sprawl-2026-ai-service-leaks-surge-81-and-29m-secrets-hit-public-github-2bgj
Palo Alto Networks. (2022, February 14). Forrester: ASM for complex cloud [Summarizing Forrester, Find and cover your assets with attack surface management]. https://www.paloaltonetworks.com/blog/?p=153605
Security Boulevard. (2026, May). Microsoft Sentinel pricing explained: How to cut costs. https://securityboulevard.com/2026/05/microsoft-sentinel-pricing-explained-how-to-cut-costs/
TechTarget. (2026). Verizon 2026 DBIR: 6 key takeaways for CISOs. https://www.techtarget.com/cybersecurity/news/366643420/Verizon-2026-DBIR-6-key-takeaways-for-CISOs
Verizon. (2026). 2026 data breach investigations report. https://www.verizon.com/business/resources/reports/dbir/