Anthropic tests AI for target location and drones
Anthropic tests AI models for locating people, geolocating images and programming drones in simulations. The most advanced systems show meaningful capabilities, while open-source models trail behind but remain within the zone of concern.

Anthropic concludes that current AI models can already help locate people, analyze photographs and develop software for drones in simulated environments. The company warns that these capabilities are not limited to the most advanced systems or to cybersecurity.
The report, prepared by its Frontier Red Team, examines two areas linked to military operations: target intelligence and conventional weapons engineering. The goal is not to show that AI can act on its own in a conflict, but to measure how much specialized work it can already take on.
Locating people through their digital footprints
In intelligence work, a significant part of the job involves finding and precisely locating a person, account, vehicle or facility. That requires connecting scattered data, identifying accounts that belong to the same individual and inferring where they are.
Anthropic tested this capability with synthetic content simulating posts from WhatsApp, Telegram, Instagram and Facebook. The models had to link accounts belonging to the same person and rank individuals according to their relevance to an investigation.
The Mythos Preview model delivered the best results in account linking. Kimi K3, an open-source model, came close to leading systems in easy and medium cases, although its performance declined when there was more noise and fewer clues.
Speed matters too. Analyzing a median-sized sample would take about 2.5 hours for a human analyst, according to the study's estimate. Claude Mythos Preview completed the evaluation in about 11 minutes on average.
The company stresses that the data was artificial and less realistic than real social networks. The results can therefore be used to compare models and identify trends, not to calculate their performance precisely in a real operation.
Images can also reveal where you are
Anthropic asked the models to locate outdoor photographs without using metadata, reverse searches or external tools. In a collection of 6,000 images, Mythos Preview achieved a median error of 37 kilometers and placed 23.7% of the photos within one kilometer of the correct location.
Mythos 5 achieved a median error of 47.2 kilometers. The results were better than the benchmark used for expert GeoGuessr players, although the comparison is not perfect because the tasks and images were different.
In another test, the models tried to infer users' home cities from a week of posts. At least one model reliably located 135 of 1,697 people, or 8%, within one kilometer of their estimated location.
Most cases were solved through explicit clues, such as universities, well-known places, streets or postal codes. But in 17% of those cases, more indirect signals were enough: dialect, public transportation, local events, sports teams or radio and television programs.
The message for any user is simple: an apparently harmless post can reveal more than it seems when AI connects many small clues.
Drones: useful results, but only in simulation
The second part of the study measured whether the models could write and improve software to control drones in a simulator. The tests included tracking a vehicle, dropping a payload near a target and navigating when positioning signals were blocked or manipulated.
The models received a task, a test environment and the results of each attempt. They could then modify their code and try again, as an engineer would when reviewing flight data. They had no access to the internet or prepared solutions.
In the simplest scenario, tracking a stationary, visible vehicle, Opus 5 achieved an 80% success rate, compared with 70% for Mythos Preview, 53% for Mythos 5, 15% for Kimi K3 and 5% for Sonnet 5.
When the vehicle moved, performance fell. In the hardest scenarios, involving camouflage, obstacles or evasive maneuvers, no model solved the problem consistently. Across the nine scenarios, Opus 5 succeeded in 20% of simulated launches and Mythos Preview in 13%.
The navigation tests showed a similar weakness. The most advanced models could follow a route with clean signals and partially react when GPS disappeared, but all of them struggled with subtle signal manipulation.
What changes outside the lab
Anthropic insists that a simulation is not the same as an operational military system. The cameras, sensors and physics were simpler than in the real world, and the models had no access to hardware, materials or field testing.
Even so, the company considers the results to represent a minimum capability, not a maximum. A person or group with more time, tools, data and testing could achieve better results. In addition, Anthropic's threat intelligence team has already detected actors using models for tasks related to surveillance and drones.
The central concern is not that AI will completely replace an army or an analyst. It is that AI could reduce the cost of expertise that was previously scarce and expensive. That could allow small groups to carry out identification, analysis or engineering work that once required specialists.
The report also challenges a strategy focused solely on protecting proprietary models. Kimi K3, an open-source model, trailed the leading systems but showed enough capability to be concerning in several tests.
The next question will be how to measure and limit these capabilities without blocking legitimate uses. Anthropic expects models to keep advancing in intelligence, surveillance and autonomous systems, including areas such as space and undersea warfare. The debate over AI risks can no longer focus only on malicious code or biology. It also has to monitor how models turn scattered data and software into operational capability.