ML Supply Chain Security intermediate 7 min read 6 flashcards

Prompt Artefacts as Executable Dependencies

Why rules files, skills, agent instructions and system prompts shipped in a repository are code by any operational definition, how invisible Unicode makes them reviewable in theory and not in practice, and what intake for a text dependency has to check.

Repositories now carry files whose only function is to change what an AI tool does: .cursor/rules, Copilot instruction files, agent skill folders, CLAUDE.md, MCP server manifests, prompt templates. They are text, they diff cleanly, and review treats them as documentation. They are not documentation. A file that deterministically alters the code an assistant writes for every engineer on the project is a build input, and it is one with no signature, no version pinning, no provenance metadata and, until recently, no rendering guarantees.

Pillar Security demonstrated the consequence in March 2025 with what it called the Rules File Backdoor. Instructions were hidden in a rules file using bidirectional text markers and zero-width joiners, invisible in an editor and in a GitHub pull request diff, that steered GitHub Copilot and Cursor into emitting backdoored code. The file passes human review because the malicious bytes render as nothing (Pillar Security, 2025, New Vulnerability in GitHub Copilot and Cursor). GitHub subsequently added a warning banner when a file's contents include hidden Unicode.

Why this artefact class is unusually dangerous

Three properties combine badly.

It persists. Once merged, a rules file affects every future generation session for every team member, not one commit. The compromise is ambient rather than located in a diff someone can bisect.

It propagates. Forks, templates and starter repositories carry it forward. Pillar's own framing is that downstream projects inherit the instruction along with the code, which is the defining property of a supply chain vector rather than a local bug.

It is invisible by design. Unicode has legitimate, standardised characters with no glyph. A review process that relies on reading is defeated by bytes that are not rendered, which is why the durable fix is a lint rule that rejects non-printing characters outside an allowlist, not more careful reading.

The same problem one layer up

Skills, subagent definitions and packaged prompt bundles extend the pattern: distributed artefacts, typically fetched from a public repository, that instruct a tool with filesystem and network access. Many are distributed as ordinary npm or PyPI packages, inheriting every risk of that channel, and are then reviewed less carefully than code because they contain no code. The MCP ecosystem makes the point sharply, since a server's tool descriptions are prose that the model treats as authoritative, with no client-side integrity check on what changed between versions.

Attestation helps less here than people expect. PEP 740 and Sigstore-backed attestations prove which identity published an artefact and which pipeline built it, not that its contents are benign (PEP 740). A signed prompt with a backdoor in it is a signed prompt with a backdoor in it.

When it breaks

Diff review is not byte review. Any control whose enforcement is "a human looks at the pull request" fails against non-rendering characters, homoglyphs and content far enough down a long file that it falls outside the reviewed hunk. Enforcement belongs in CI.

Provenance is absent by default. Nothing in the common formats records who wrote a rules file, from which upstream, at which revision. A rules file copied from a blog post three years ago is indistinguishable from one the security team wrote, and neither has a version to pin.

The blast radius is defined by the tool, not the file. The same three-line instruction is inert in an assistant that only suggests completions and severe in an agent that can edit files, run commands and open pull requests. Risk assessment has to name the tool the artefact is loaded into, and it changes when the team upgrades.

Detection of the effect is late and probabilistic. Even with the artefact quarantined, the code it already influenced remains. Remediation means auditing generated code from the period the artefact was present, which is usually unbounded because nobody records which commits were model-assisted.

Allowlists drift. A policy that permits specific rules files from specific sources decays as teams copy folders between repositories. The check that survives is mechanical: no hidden characters, known path, known hash, reviewed at that hash.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Pillar Security, 2025, New Vulnerability in GitHub Copilot and Cursor pillar.security
  2. PEP 740 peps.python.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track