PROMPTFLUX: What Google’s Report Says About AI-Assisted Malware

Google reported an experimental VBScript dropper that queried Gemini to rewrite and obfuscate its code. The sample was still in development, and Google did not report a successful device or network compromise.
PROMPTFLUX: What Google’s Report Says About AI-Assisted Malware

Malware developers are testing AI services not only while writing code, but also as a component that a malicious program can call while it runs. Google Threat Intelligence Group (GTIG) documented PROMPTFLUX, an experimental VBScript dropper discovered in June 2025, as one example.
The finding matters, but its boundary matters too: Google described PROMPTFLUX as being in development or testing and did not report a successful compromise of a device or network by the sample.
The Discovery: PROMPTFLUX and AI-Driven Obfuscation
GTIG found that PROMPTFLUX was designed to call the Gemini API during execution and ask for a more obfuscated version of its own VBScript code.
This does not mean the sample independently selected goals or learned from an environment. It followed attacker-written instructions that asked a model to produce another code variation.
How It Works
The analyzed sample attempted several specific actions:
- Runtime rewriting: It queried Gemini for another version of its VBScript code.
- Obfuscation request: Its prompt asked for changes intended to make antivirus detection less likely.
- Repeated update attempt: One version instructed the script to rewrite itself hourly.
- Preserved components: The requested rewrite retained the decoy payload, API key, and self-modification logic.
What the Report Did Not Establish
GTIG assessed PROMPTFLUX as a development or testing sample. Some functions were incomplete, Gemini API use was rate-limited, and the analyzed version did not demonstrate the ability to compromise a device or network.
Google also disabled resources connected with the activity. The report therefore does not support a claim that PROMPTFLUX had already infected organizations successfully.
This distinction separates an observed technical experiment from a confirmed campaign outcome.
Defensive Implications
The experiment shows why defenders should monitor program behavior and outbound service calls alongside static code signatures.
Signature Checks Are Only One Layer
Changing code can weaken a signature-only control. Behavioral monitoring remains useful because a suspicious script may still call an external AI service, download or rewrite code, run on a schedule, and request unusual permissions.
AI-Service Calls Create Observable Activity
External API dependencies can create network, credential, rate-limit, and service-disruption signals. They are potential detection points as well as attacker capabilities.
Detection Questions
Depending on the environment, teams can check for:
- script execution that does not match an approved application;
- unexpected calls to model APIs or other code-generation services;
- scheduled rewriting, downloading, or execution behavior; and
- exposed or unfamiliar API credentials on an endpoint.
Technical Limitations—For Now
The reported sample had practical limitations:
- External dependency: Runtime rewriting depended on reaching an external service.
- Credentials: The workflow needed an API key that could be detected, revoked, or disabled.
- Rate limits: Google reported that API limits constrained the tested behavior.
- Incomplete capability: The analysed version did not demonstrate a successful compromise.
These constraints may change as services and attacker methods change, so they should be monitored rather than assumed permanent.
Possible Future Directions, Not Observed PROMPTFLUX Behavior
Security teams may reasonably test scenarios in which malicious software gains more adaptive capabilities. The following are possibilities, not behaviors that Google verified for PROMPTFLUX:
- More adaptive task selection within attacker-defined boundaries
- Use of deployment feedback to revise later variants
- Coordination through shared attacker infrastructure
- Faster production of code variations
What This Means for Organizations
For businesses and security teams, the report supports several practical checks:
Enhanced Monitoring
Relevant monitoring may include:
- Unusual patterns of code execution
- Unexpected API calls to AI services
- script behavior that does not match an approved workflow
- Network traffic patterns consistent with AI model queries
Defense in Depth
Select layers from the organisation's threat model rather than from the "AI-powered" label. Examples include restricted script execution, least-privilege access, endpoint telemetry, network controls, credential protection, and tested recovery.
Incident Response Evolution
An exercise can test whether responders can identify the process, revoke its API credential, block the external dependency, preserve evidence, and contain the endpoint. PROMPTFLUX itself did not demonstrate adaptation to containment efforts.
The Broader Context
The same report discusses a wider pattern of threat actors using AI tools. The following scenarios should be treated as risks to test, not as capabilities demonstrated by PROMPTFLUX:
- More sophisticated social engineering attacks using AI-generated content
- Automated vulnerability discovery and exploit generation
- AI-powered reconnaissance and target selection
- Deepfake-enhanced phishing and fraud
Evaluation Checklist
Defenders can respond without treating an experimental sample as a proven autonomous threat:
- separate behavior observed by Google from future scenarios;
- verify whether model-service traffic is visible in existing endpoint and network logs;
- control where script interpreters and API credentials are allowed;
- add the observed runtime-rewriting pattern to a contained incident exercise; and
- recheck the primary report when new samples or provider findings are published.
Conclusion
PROMPTFLUX does not prove that self-learning malware is already operating autonomously at scale. It does show that an attacker tested a live AI service as part of runtime code rewriting.
The appropriate response is evidence-led monitoring: control script execution, protect API credentials, inspect unusual outbound calls, combine behavioral and endpoint signals, and include this technique in incident-response exercises.
For Oyu Intelligence, the relevant lesson is to separate an observed capability from a projected one and to design controls around behavior, connections, credentials, and permissions rather than fear-based labels.

Read Next
Related notes

Adversarial Poetry and AI Safety: What a 25-Model Study Found
In a preprint, 20 hand-written poetic prompts produced a 62% average attack-success rate across 25 models; a larger 1,200-prompt conversion test averaged about 43%.

A Design Question for Multi-Agent Systems: When Is Language the Wrong Interface?
An engineering essay on where free-form agent dialogue can lose intent, and when schemas, state machines, or typed messages may be easier to inspect.

Automation, AI Workflows, and Agents: A Practical Boundary
A concise way to choose between fixed rules, model-assisted steps, and systems that select among tools—without calling every workflow an agent.