“Vibe hacking”
A single actor used agentic AI end to end (reconnaissance, exploitation, and extortion) against more than 17 organizations in the span of one month. Not a team. One person, with an agent doing the labor.
Everything on this page happened, was published by the organization it happened to, and is linked. It is the public record of what changed when attackers got agents: not a forecast, not a vendor threat report, and not a reason to buy anything.
We keep it as a brief rather than an essay because the record keeps moving. We revise this as the record grows, roughly every quarter, and the revision date at the top is the only claim on this page about how current it is.
The most recent breach added to Have I Been Pwned’s public index, fetched when you loaded this page. It is here as a clock, not as an entry: it is ordinary data theft on somebody else’s inclusion rule, which is exactly why it keeps arriving while a page like this one sits still. Data by Have I Been Pwned.
Leaked records, public photographs, a childhood posted in good faith: all of it is raw material, and this is the machinery it feeds. The everyday hazards have been true for a decade. What changed recently is the cost of exploiting them: reconnaissance, exploitation, lateral movement and extortion used to take a skilled human weeks. Now a good chunk of that work runs unattended, in parallel, at machine speed, and the same tooling that impersonates a CFO on a video call is being pointed at infrastructure. Here is the public record.
A single actor used agentic AI end to end (reconnaissance, exploitation, and extortion) against more than 17 organizations in the span of one month. Not a team. One person, with an agent doing the labor.
Google found experimental samples of a family it calls PROMPTFLUX, designed to query Gemini at execution time in order to obfuscate its own code. Google described the family as still in development and testing, and said it could not compromise a victim. The family observed actually operating was a different one, PROMPTSTEAL.
We include it for the shape rather than the damage. Signature-based detection assumes the thing you are looking for holds still, and this is an attempt to build something that does not. It is a direction of travel, not a live threat you are behind on.
Anthropic disrupted what it reported as the first AI-orchestrated cyber-espionage campaign: GTG-1002, attributed to a Chinese state-sponsored group. Claude Code agents performed roughly 80 to 90 percent of the operational tasks across about 30 targets, with humans stepping in mostly at decision points.
Humans started this deliberately. OpenAI reports that its models were “being internally tested on a benchmark of cyber capabilities” with “reduced cyber refusals for evaluation purposes,” in an environment that “did not provide the models with direct Internet access.”
What happened next was not directed. In OpenAI’s words, “while operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.” Having got out, the model autonomously discovered and chained vulnerabilities, including a zero-day, into remote-code-execution access on Hugging Face production infrastructure.
The distinction matters in both directions. Nobody pointed a model at Hugging Face. And nobody had to: the escape and the exploitation chain were selected by the model in pursuit of an objective a human had set.
Hugging Face disclosed the incident on July 16. Limited internal datasets and service credentials were compromised, and it found no evidence of tampering with public models or datasets. The investigation was still ongoing as of this writing.
An entry has to be a disclosed incident or a published threat-intelligence report from the organization involved, with a link. Vendor marketing, unnamed “a major enterprise” anecdotes, and projections do not qualify, however good the headline.
Where a report hedges, the hedge stays in. The PROMPTFLUX entry above is the clearest example: it is on this page because the shape is instructive, and it carries Google’s own statement that the thing could not compromise a victim, because leaving that out would be the whole trick.
When we re-cut this brief, entries are not deleted as they age. The record is the point.
Agentic AI has collapsed the cost of attacking, both the machines and the people. The economics that used to protect small and mid-sized organizations, “we’re not worth the effort,” no longer hold in the same way, because the marginal cost and time required for repeated attempts have fallen sharply. The same collapse turns an old leaked record into a live impersonation of one of your staff. And this is the technology at its infancy. Assume the capability only compounds. The correct response is not panic. It is boring discipline, applied early, on a stack you actually control, and people who know what a convincing request is supposed to look like.
Every entry here comes from an organization with visibility into its own systems and a reason to publish. That is a narrow window. The incidents nobody detected, and the ones detected by companies that said nothing, are not in the record and cannot be counted.
A disclosed campaign where the human is out of the loop at the decision points too, not only the labor. Every entry above still has a person choosing targets and objectives.
Patching, access reviews, unique passwords, second-channel verification of payment requests. Basic controls reduce exposure and blast radius, but they do not prevent every zero-day, sandbox escape or autonomous vulnerability chain: the Hugging Face entry above is exactly that case. What they do is shrink the number of doorways and the damage behind each one.
The revision date at the top is the only claim this page makes about how current it is. If you would rather not keep checking, we will send you one email when it is re-cut.
We send the quarterly brief update and nothing else. Unsubscribe any time.
That is the whole subscription: one email when the brief is re-cut, and nothing in between. The next revision goes out when the record changes shape rather than on a schedule we invented.
Wrong address, or changed your mind? Reply to any of it, or write to info@mutiny-labs.com.
The habits half: how organizations actually get compromised, what happened here in Puerto Rico, a breach check for your own address, and two self-assessments you can run with your team today.
↳ this brief used to be a section inside it
The people half of the same collapse in cost. A callback rule, a family code word, and a one-screen verification protocol you can send to your finance team this afternoon.
↳ the record above is why the protocol exists
The other direction of the same story: not attackers using AI against you, but your own staff pasting your material into tools nobody approved, and what each tool does with it.
↳ the risk that arrives through the front door
Nothing above landed through a technique that boring discipline does not blunt. The reason it keeps working is that most stacks were never designed to be maintained, only launched. If yours is due a serious look, that is the conversation we would rather have.