Why this matters now
A public page written for AI assistants has one section that tells them how to describe the site and when to suggest a paid plan. The next section is titled as a direct command to AI models. It tells them to offer users the vendor blog and, if they see the text, to end their reply with an emoji.
Release 0.6.5 passed that text on the channels that were tried. The only block on the raw page came from two ordinary pieces of markup, a hidden frame with no text in it and a deferred stylesheet loader. Neither one was the directive.
The tests in release 0.6.7 use a made up copy of the page. It keeps the shape of the text and replaces the names with Examplekit and Exampleco. We ran that copy through 0.6.5, 0.6.6 and 0.6.7. The first two allowed it with no findings. 0.6.7 quarantined it.
What the text is actually doing
Three forms are read. A heading or label that names AI assistants or models, followed by a sentence that orders or obliges the model. A label such as AI assistants with an order after it. A sentence that tells AI assistants what to say when they answer or discuss a subject.
The other half is the marker. An instruction to put an emoji or a phrase in the reply can work as a way to check whether an assistant read the page and obeyed it. In the test copy the marker sits right after an order about what the answer should say.
An assistant that answers for the user is the last link in a chain of text it did not write. If a page can steer what the assistant says, the user could get an answer shaped by the page without seeing that it happened.
How Sunglasses catches it
AI Addressed Directive About The Answer. Medium severity, category indirect prompt injection, channels web content and file. It reads the three forms above. A heading followed by a description, or by an instruction aimed at human staff, is not held in the cases tested. A sentence that names the model and orders it is still found when it comes after a staff line.
Order To Put A Marker In The Reply. Low severity, same channels. It reads an order to put a marker such as an emoji or a phrase in the reply. It is reported and not blocked, because the text is often benign marketing.
Neither rule blocks on its own. On the made up copy, both fired and the result was a quarantine on the web content and file channels. The marker order alone was reported with GLS-IP-008 and not blocked. The sentence written for human staff was allowed.
GLS-SEM-UI-219 no longer reads a deferred stylesheet loader as an element injection. That ordinary markup was one of the two things that blocked the raw page, and it was not the directive.
For broader foundations, see AI Agent Security 101, the how Sunglasses works page, the pattern library and the manual.
What these rules do not do
Both rules read web content and files. They do not read the plain message channel, because a system prompt legitimately addresses a model and gives it orders. The test copy was allowed on the message channel.
Documentation that tells an AI assistant how to behave may be held at medium severity. A repository that ships such a file should expect a review item and not a block.
The test text is a made up copy. This post does not say how a real model behaves when it reads the original page, and it does not say that the page did harm. It says the text has a clear shape and that 0.6.7 can now see that shape.
This is a text rule, not a verdict on the site that published the page. The scanner reads what the agent is about to read and reports what it found.
What to do besides scanning
- Treat fetched pages as data. Text on a page is content for the user, and it does not carry the authority of the user or the operator.
- Keep the system prompt apart from fetched text. Tell the model where each piece came from.
- Scan before the agent reads. Run
sunglasses scanon a page, a file or a string at the terminal or in CI. - Review quarantined items. A quarantine means a person should look before the agent acts on the text.
Conventional controls answer the question can this page ever be safe. A scan answers a narrower one, should this agent act on this text right now.