Blog

Truffle Security found 543,699 still-valid credentials in code assembled for AI training

A Cloud Security Alliance briefing reports 543,699 still-valid credentials in AI training data. Learn what this means for secrets management beyond the vault.

· By

An open steel box with brass keys neatly on hooks, while loose duplicate keys lie across the wooden desk and in a tray.
A secrets manager protects only stored credentials; copies in code, logs and training data stay valid outside the vault.Image: IamVera.ai — original editorial illustration

A dataset gathered for AI training was found to contain 543,699 still-valid credentials, the Cloud Security Alliance reports. This shows secrets management must not stop at storing and rotating: limit the scope and lifetime of credentials upfront, isolate agent execution at runtime and actively scan beyond the secrets manager for leaks in code, logs and training data.

The Cloud Security Alliance briefing of 1 October 2026 summarises an analysis by Truffle Security: in a corpus of 224 million public GitHub repositories assembled for training AI models, Truffle Security tested 543,699 unique credentials in July 2026 that still authenticated successfully. The finding concerns credentials already present in public code that became discoverable through an additional collection channel; it should not be read as evidence that the training-data corpus itself created all of the underlying exposures. Our analysis is that the collection of training data creates an additional place where a secret can be copied, remain discoverable and require separate controls.

For decision-makers this means the question shifts from where do we store our secrets to which systems can read, take along or re-expose our secrets. AI workflows increase the number of those places.

What exactly did the Cloud Security Alliance establish about credentials in AI training data?

The Cloud Security Alliance reports that Truffle Security found 543,699 still-valid credentials in a large collection of public GitHub code that was intended as training data. The new insight is not that secrets appear in public code, but that assembling training corpora forms an additional layer for discovery and exposure.

That picture aligns with figures GitGuardian reported: public code contained 1.27 million exposed credentials for AI services, and 64% of secrets confirmed as valid in 2022 had still not been revoked in January 2026. According to GitGuardian, that is precisely where a secrets manager, which protects only managed secrets, and monitoring of public leaks diverge.

Why does a secrets manager not see the credentials in logs, caches and training data?

A secrets manager protects the secrets that have been deliberately placed in it and are managed there. It does not automatically see the copy a developer pasted into a repository, the token that ended up in a log line or the key that became available elsewhere via a model cache or training dataset. Those copies live outside the vault; rotating the credential may invalidate its use, but it does not by itself identify or remove every copy or undo every exposure location.

Mandiant describes in its special report how agentic workflows can expose credentials via configurations, tool caches, terminal history and the output of agents themselves. The exploit roundup of the OWASP GenAI Security Project adds concrete incidents in which overly broad rights, tool connections and MCP servers enabled unauthorised access or created opportunities for attempted secret exfiltration. In our assessment the lesson from both sources is that instructions in a prompt do not form a reliable boundary: In our assessment, giving an agent access to private data, untrusted content and external communication creates a potential channel for data exfiltration that must be controlled.

How do I bound credentials before, during and after the execution of AI agents and pipelines?

Our analysis translates the separate source findings into three additional layers of control: before, during and after execution.

  • Before execution. Our recommendation is: give each agent its own identity, use short-lived and task-bound credentials, apply minimal rights and keep development, training and production data separate. Mandiant in particular supports separate agent identities, short-lived credentials and least privilege. That way a credential that does leak can open little and expires quickly.
  • During execution. Our recommendation for during execution: limit where an agent may connect to (egress), set tool and MCP connections in policy, mask secrets in output and logs, and ensure runtime telemetry plus an independently controlled authority to revoke tokens immediately when behaviour deviates. Mandiant supports, among other things, egress limitation, runtime telemetry and credential revocation.
  • After exposure. Scan beyond the secrets manager: public repositories, caches, log files, model artefacts and training corpora. Follow up a find with rotation, determine the scope of the damage and record evidence of remediation.

Our addition: the sources describe these measures separately from each other. The open point is how continuous detection beyond the vault relates to human verification in automated pipelines. A scan that finds a secret but has no one deciding whether and when it is revoked delivers no protection but a list.

What does this news mean for directors, lawyers and CISOs?

Our analysis: because the Cloud Security Alliance shows that valid secrets became discoverable again via a training corpus, organisations should not assume that a secrets manager alone covers the exposure risk; this makes it relevant for directors and CISOs to ask whether detection exists beyond the vault and who is authorised to revoke a leaked token. Because Mandiant and OWASP show that agents can funnel secrets via caches, logs and tools, a concrete runtime risk arises for CISOs that does not disappear with policy alone. Our analysis is that production agents should generally be assigned separate identities, short-lived credentials and bounded egress, with the exact controls matched to the workflow's risk. Organisations must also govern an AI agent with system access as a privileged identity. Because GitGuardian reports that many previously confirmed secrets remained unrevoked, organisations may want supplier and processor agreements to define responsibilities, timelines for detection and rotation, and the evidence of remediation that must be retained.

As a summary of that analysis, concrete enough to assign to an owner:

  1. Map which agents and pipelines have access to long-lived credentials and replace them with short-lived, task-bound variants.
  2. Set up detection beyond the secrets manager for public code, logs, caches and training data, and link every find to a person who may revoke.
  3. Record per sensitive workflow which agent, which scope, which data environment and which human approval were linked to it.
  4. Put a demonstrable term for rotation and evidence of remediation after exposure into supplier contracts.

These steps align with the broader shift from perimeter defence to identity and runtime of AI agents and with the need for an enforced emergency stop and revocation for AI agents. More background is in our topic hub on AI security and securing AI workflows.

Sources and references

  1. CISO Daily Briefing: AI Training Data Scrape Surfaces 543K Live SecretsCloud Security Alliance · 2026-10-01
  2. AI Risk and Resilience in 2026: A Mandiant Special ReportGoogle Cloud / Mandiant · 2025-09-15
  3. GenAI and Agentic AI Exploit Roundup Q3 2026OWASP GenAI Security Project · 2026-10-08
  4. Public Secrets Monitoring: AI Analysis of Leaked CredentialsGitGuardian · 2026-09-07

Sources: The article relies on the Cloud Security Alliance (Truffle Security analysis), Google Cloud's Mandiant report, the OWASP GenAI exploit roundup and GitGuardian.

← All articles in this topic ← All articles