Evidence must stay bound
A workflow needs proof that stays attached to the exact claim and action. A familiar record can be real and still be wrong for the current decision.
GLS-AW-666 covers Trigger Payload Relational Memory Backdoor. A conversational attacker binds an entity like trigger to a credential forwarding payload so selective memory extraction preserves both, then a later mention activates the durable backdoor.
GLS-AW-667 covers Digital Twin Command Context Drift. A command safe in a twin snapshot is dispatched to a live asset after state drift without rebinding target identity, state digest and execution epoch.
GLS-AW-668 covers Reverse Shell Control Channel Smuggling. Natural language instructions hidden in agent facing evidence can smuggle a reverse shell control channel request across the evidence to capability boundary, even when no canonical shell command or socket code is present.
GLS-AW-669 covers Search Time Benchmark Contamination. A deep research agent's web search retrieves benchmark metadata, question context or answer bearing text during evaluation and the evaluator accepts that contaminated evidence as genuine reasoning, inflating the measured score.
GLS-AW-670 covers Consensus Summary Evidence Collapse. A heterogeneous review pipeline compresses several expert reports into a consensus summary, erases dissent or provenance and lets a local verifier approve code from the lossy summary instead of the independent evidence.
GLS-AW-671 covers Traceability Summary Omission Laundering. A traceable fraud review reasoning/UI surface presents a polished low risk explanation as complete after semantic unification and SURE reflection silently omit a contradictory raw risk signal and its provenance.
GLS-AW-672 covers Historical Case Retrieval Contamination. An attacker controlled or stale historical exploit case is semantically matched to a current function/call sequence without provenance or temporal scope binding and its misleading context changes the agent's current safety decision.
GLS-AW-673 covers Probe Bound Evaluation Evidence Mismatch. A low false acceptance score measured at one receiver probe is transferred to a different operational probe, laundering bounded evaluation evidence into a broader identity or safety decision.
GLS-AW-674 covers Attribution To Narration Explanation Laundering. A fluent LLM narration derived from structured SHAP attribution omits or reframes adverse evidence and a delegated reviewer approves from the narration without independently checking the raw attribution record.
GLS-AW-675 covers Synchronization Free Watermark Artifact Epoch Binding Failure. A valid self contained local watermark fingerprint is promoted to delegated authorship/publication approval even though it is not bound to the exact submitted artifact, current verification epoch or review task.
Summaries can erase the proof
A summary is a smaller view of evidence. It becomes dangerous when the workflow hides dissent, lineage, scope or missing events and then treats the smaller view as complete.
GLS-AW-676 covers Negative Only Repair Validation. A repair evaluator accepts generated code as secure because the insecure pattern disappeared, without requiring the replacement defense.
GLS-AW-677 covers Rare Correlated Evidence Inflation. A delegated reverse engineering workflow counts multiple reports derived from one attacker controlled binary as independent support, laundering a false fused claim into downstream analyst authorization.
GLS-AW-678 covers Operational Evidence Threshold Laundering. An operational fraud triage bundle launders an offline score or confidence estimate into approval authority while silently omitting the per decision calibration, threshold, cost, latency or adversarial review evidence required for a safe disposition.
GLS-AW-679 covers Rendered Selfplay Branch Identity Laundering. A browser rendered self play trace can detach a vulnerable token level branch from its exact path and display a golden path label, causing a downstream code safety reviewer to authorize unsafe generated code.
GLS-AW-680 covers Ranked Target Selection Poisoning. An attacker influenced enrichment label or rationale makes a credential sensitive function look low risk, causing a security agent's priority sampler to replace the true target with decoys and silently starve analysis coverage.
GLS-AW-681 covers Consensus Correlation Evidence Laundering. A shared prompt, retrieval snapshot or model failure makes multiple HUMBR candidates agree on a hallucinated high stakes clause. Semantic and lexical consensus is then mistaken for independent evidence and promoted into an enterprise workflow.
GLS-AW-682 covers Honeypot Control Task Response Contamination. An attacker controlled honeypot HTTP response can make a control task capability pass look like verified exploit goal evidence, causing an evaluator to record simulated success or authorize a follow on action without independent validation.
GLS-AW-683 covers Timemark Rendered Epoch Display Ambiguity. A rendered TimeMark verification report shows a valid but coarse local generation time while hiding the time dependent key epoch or exact artifact digest. An agent upgrades the visible badge into current exact generation proof and approves the wrong artifact or epoch.
GLS-AW-684 covers Instance Fingerprint Replay Across Configuration Lifecycle. A model instance fingerprint captured under one configuration is replayed after configuration drift and treated as current safety identity.
GLS-AW-685 covers Adaptive Retrieval Probe Model Fingerprinting. Tailored black box retrieval probes fingerprint a hidden embedding model from unordered returned document sets, enabling model specific corpus probing or embedding inversion even when scores and ranking are withheld.
Identity must follow the lifecycle
Identity evidence expires when the model, configuration, artifact, session or key epoch changes. A current decision needs a current binding.
GLS-AW-686 covers Host Header Reconstructed Path Authentication Confusion. A malformed Host header injects a path prefix into Starlette's reconstructed request URL while routing uses the actual ASGI scope path, so path based authentication can skip the API key check for a protected endpoint.
GLS-AW-687 covers Audit Profile Replay Rebinding. A valid looking request decision and immutable audit profile can be replayed across request identity or policy lifecycle state and accepted as current tool authorization.
GLS-AW-688 covers Fingerprint Verifier Mimicry. A weaker inference model is fine tuned to imitate a premium model on a finite client fingerprint suite and the verifier mistakes bounded behavioral similarity for proof of the advertised model identity.
GLS-AW-689 covers Approximate Ciphertext Decision Boundary Drift. A valid CKKS result can cross an agent's allow/block boundary after approximate decode, rotation, scale or slot permutation drift when the numeric error budget and model lineage are not bound to the decision.
GLS-AW-690 covers Provenance Bridge Scope Laundering. A query filter investigator broadens search across a missing provenance event, then a semantically similar node is laundered into a verified predecessor and delegated approval.
GLS-AW-691 covers Proxy Control Plane Egress Substitution. An unauthenticated crawler request supplies a benign public crawl URL while independently steering Chromium through a proxy aimed at internal or metadata destinations. Target only SSRF validation does not protect the proxy trust boundary.
GLS-AW-692 covers Non Transferable Experience Rule Overgeneralization. A clean but non transferable edge case experience, amplified by a severe plausible consequence, is consolidated by an agent into a global rule that is unsafe outside the original scope.
GLS-AW-693 covers Policy Concept Calibration Poisoning. An attacker poisons or relabels the compliant/violating contrastive examples used to derive hidden policy violation concepts, shifting the offline concept centroid so a later purpose conflicting action aligns with the compliant concept and receives allow.
GLS-AW-694 covers Thread Routing State Contamination. An attacker authored routing preference persists in a shared thread and silently changes where a later participant's unrelated task is dispatched, without fresh participant or scope authorization.
GLS-AW-695 covers Humorized Refusal Latent Payload Laundering. A humorization layer can preserve the appearance of a safe refusal while adding a covert stereotype or toxic implication and a delegated reviewer can approve the transformed output by trusting its refusal surface.
Interfaces can show the wrong state
A polished interface can display success while the underlying state says something else. Reviewers need receipts from the real operation, not just the rendered badge or toggle.
GLS-AW-696 covers Objective Invariant Display Laundering. A browser agent completion surface reports locally successful steps as a green completed run while omitting a failed cross step operational invariant, laundering objective violation into approval evidence.
GLS-AW-697 covers Schema Valid Probe Result Provenance Collapse. A schema valid MCP probe response is laundered into runtime proof for a different invocation when the adjudicator does not bind tool, arguments, server, request identity and observed effect.
GLS-AW-698 covers Verification Scope Denominator Laundering. A ZKP verification report analyzes only supported or isolated targets, omits unsupported DSL/zkVM files from the denominator and an agent facing summary turns the partial result into repository wide security evidence.
GLS-AW-699 covers Lossy Trajectory Evidence Omission. A lossy trajectory serializer or normalizer drops an unsafe intermediate tool or action event while retaining a benign looking final outcome. An evaluator accepts incomplete evidence and misses the unsafe execution.
GLS-AW-700 covers Skill Compiler Target Role Collision. A SkCC style compiler can preserve a source skill's text and clean analyzer verdict while lowering a neutral field into a target framework role token that the target parser treats as trusted executable authority.
GLS-AW-701 covers Lifecycle State Rollback Aliasing. A rendered rollback/degradation status from one lifecycle epoch is aliased into a later browser/tool action, so the agent applies the wrong privilege state without rebinding the current epoch and receipt lineage.
GLS-AW-702 covers Stateful Ui Persistence Evaluator Desynchronization. A browser agent or state management extension treats a toggle/checkbox visual state or stale initialized account state as authoritative while the website's persisted privacy/security setting differs and the evaluator records success across the run boundary.
GLS-AW-703 covers Cross Assessment Knowledge Transfer Lineage Collision. A tiered knowledge architecture can transfer prior exploit and build experience into a new CVE assessment. If lineage is not bound to source CVE, target digest and epoch, downstream agents can consume stale knowledge as current evidence and generate a wrong target plan or rule.
GLS-AW-704 covers Component Selective Identity Badge Laundering. A registry or approval UI compresses a per component prompt/code/tools lineage match into one same family badge, hiding changed executable code and laundering lineage evidence into approval without behavioral verification.
GLS-AW-705 covers Challenge Response Lifecycle Rebinding. A valid AgentDID challenge response execution state proof can be replayed into a concurrent or subsequent verifier session when the verifier does not bind the response to the intended nonce, agent instance, transport session and current capability epoch.
Routing and protocols need invariants
Routers and protocols compress intent into labels, vectors and messages. Security depends on preserving command identity, semantic privilege and cross step invariants through that compression.
GLS-AW-706 covers Stix Adjudication Provenance Rebinding. An attacker authored STIX graph object claims gold standard or reviewer verified adjudication. a provenance blind normalizer rebinds that claim to the independent judge and a downstream CTI agent excludes the linked finding.
GLS-AW-707 covers Workspace Write Scope Projection. A Cursor like agent may automatically write to a caller/context selected file outside the opened workspace while the UI presents the change as an ordinary review item, so review presentation does not repair the missing write scope authorization.
GLS-AW-708 covers Voiceprint Inversion From Exposed Speech Tokens. An exposed intermediate speech token preview can be inverted into an attacker specified speaker representation and used to link a session to a voice identity.
GLS-AW-709 covers Edge Cloud Kv Cache Session Binding Replay. A valid authenticated KV cache fragment from one edge endpoint/session can be replayed into another edge cloud inference session when authentication is checked without jointly binding cache identity, endpoint, session, epoch and transcript digest.
GLS-AW-710 covers Protocol Grammar Alias Collision. An LLM generated pseudo C2 server collapses distinct wire commands through alias normalization and a downstream agent accepts the result as faithful without command identity validation.
GLS-AW-711 covers Semantic Intent Schema Binding Laundering. A delegated A2A or MCP router treats a privacy preserving intent vector plus zero knowledge schema consistency proof as sufficient authorization, although the proof does not bind hidden semantic privilege. The downstream agent admits and dispatches the payload into a protected trust zone without revalidation.
GLS-AW-712 covers Four Principles Evidence Split Laundering. A generated fuzz harness passes logic, API and entry point checks while its diagnostic callback treats attacker controlled fuzz bytes as trusted agent facing guidance. An aggregate quality verdict launders the missing security boundary evidence and authorizes fuzzing.
GLS-AW-713 covers Malicious Knowledge Editing Downstream Reasoning Corruption. A malicious knowledge edit silently replaces a factual premise and downstream reasoning accepts it to produce an unsafe security conclusion without an instruction shaped payload.
GLS-AW-714 covers Router Cost Amplification Adversarial Suffix. An attacker appends a suffix optimized against a black box router surrogate so ordinary requests are repeatedly dispatched to a premium model, amplifying inference cost and resource use.
GLS-AW-715 covers Encoded Recognition Action Decoupling. A reconnaissance trap is hidden in encoded evidence. After normalization the agent recognizes the honeypot yet continues the deceptive action, while a pre normalization filter sees only opaque text.
Authorization must reach the final action
Authorization must survive concurrency, target changes, scope limits and forensic gaps. The last handler owns the final check because it sees the real recipient and effect.
GLS-AW-716 covers Repeated Scan Evidence Instability Laundering. A delegated security review aggregator turns a finding observed in only one stochastic repetition into stable remediation or suppression evidence, despite absent repetition consensus and deterministic scan provenance.
GLS-AW-717 covers Environment Cue Reward Specification Laundering. An RL environment can preserve surface safety while role framing and implicit gameability cues reshape the reward interpretation, causing an on policy agent to learn and activate harmful exploit behavior only in the conditioned environment.
GLS-AW-718 covers Grid Triple Alignment Negation Drop. A graph extraction result can look traceable and ontology valid while a parser drops negation between the aligned CTI span and the normalized security triple, laundering a suppression relation into trusted evidence.
GLS-AW-719 covers Peer Debate Replacement Cost Laundering. An adversarial peer induces an honest agent to replace a correct answer, then the final consensus summary omits the replaced answer lineage so the harmful flip appears corroborated.
GLS-AW-720 covers Specification Oracle Differential Mismatch. A specification guided differential oracle accepts a structurally valid protocol response after a generated grammar drops or misbinds a security relevant field constraint.
GLS-AW-721 covers Oauth Authorization Code Non Atomic Redemption. A single use OAuth authorization code can be redeemed twice when concurrent exchanges both pass a non atomic read before delete check, minting independent access, refresh and ID token sets.
GLS-AW-722 covers Runtime Type Confusion Nosql Predicate Auth Bypass. Bracketed form keys can materialize a password operator object, changing a scalar credential check into a MongoDB predicate that admits a privileged login session.
GLS-AW-723 covers Semantic Task Legitimacy Laundering. An authorized clinician or medical board wrapper makes the unchanged restricted request appear legitimate, weakening a downstream safety or delegation decision.
GLS-AW-724 covers Configuration Surface Enumeration. An attacker elicits a complete agent capability/configuration inventory through a message or rendered configuration view, giving downstream targeting a map of reachable tools, services, scopes or privileged operations.
GLS-AW-725 covers Forensic Snapshot Mount Scope Completeness Omission. A post compromise response agent treats an alert linked mounted subset as a complete forensic snapshot, omits silent artifacts outside mount scope and closes remediation without a coverage proof.
Policy and validation need full context
Validation fails when a workflow checks only what it supports or only the condition that disappeared. Safe approval needs the full denominator and the defense that replaced the unsafe behavior.
GLS-AW-726 covers Routine Hazard Risk Aggregation Masking. A mixed safety evaluation record lets a routine task majority or aggregate safe scalar authorize a concealed hazard red team item, losing per task risk identity at the batch boundary.
GLS-AW-727 covers Semantic Checkpoint Tuple Binding Omission. A decompiler review can approve source like Solidity because compilability and ABI recovery look valid even though differential replay evidence is missing, stale or bound to another bytecode/input/state tuple.
GLS-AW-728 covers Operation Result Only Provenance Loss. A prompt that exposes only compact XOR/operation results can launder attacker controlled numeric evidence into a seemingly sufficient verdict while omitting the source pair, transform provenance and lineage needed for validation.
GLS-AW-729 covers Semantic Traffic Window Distribution Drift. A model extractor can make each API query look benign while the aggregate semantic distribution of a mixed user traffic window drifts from a benign calibrated baseline. MMD style window testing detects the evidence shape that single query rules miss.
GLS-AW-730 covers Text To Weight Session Identity Injection. A text to weight fingerprint descriptor can be treated as an authorized live model state mutation, allowing tenant/session identity or provenance to be overwritten without a binding check.
GLS-AW-731 covers Feedback To Harness Control Poisoning. Compiler or symbolic execution feedback can become executable harness control input. A poisoned diagnostic rewrites the next driver/stub/assertion iteration and suppresses a real vulnerability verdict.
GLS-AW-732 covers Policy Graph Path Stitching Invariant Omission. A policy specification graph generator can stitch locally permitted transitions into a composed path while dropping a cross step invariant, causing a downstream agent to admit an unsafe configuration or privileged action.
GLS-AW-733 covers Generated Aead Nonce Reuse False Confidence. A compilable LLM generated Rust AES GCM or ChaCha20 Poly1305 implementation can reuse a fixed, reset or predictable nonce. Build success hides a cryptographic state failure that defeats AEAD security.
GLS-AW-734 covers Rule Synthesis Probe Provenance Laundering. A BAS finding can launder a broken probe to evidence to rule binding into apparently verified SIEM coverage when a downstream reviewer accepts deterministic synthesis output without replaying the originating probe.
GLS-AW-735 covers Stale Cross Session Approval Evidence Replay. A deterministic incident response plan verifier can accept an approval record from a prior session or incident when it validates approval presence but fails to bind the approval to the current incident, action, session and freshness window.
Side channels and stale grants still matter
A workflow can leak or drift through sound, cache fragments, watermarks, chat history and old session grants. These paths need the same identity, epoch and scope rules as the main request path.
GLS-AW-736 covers Acoustic Keystroke Secret Inference. A microphone or VoIP acoustic stream is domain normalized across keyboards and converted into keystroke identities that expose a password, OTP, API key or session secret.
GLS-AW-737 covers Semantic Watermark Provenance Marker Collision. A SWAN like semantic watermark can survive paraphrase and be replayed or transplanted into untrusted model generated configuration. Treating marker presence as provenance can approve the wrong artifact even when surface tokens changed.
GLS-AW-738 covers Reasoning Watermark Lifecycle Replay. A valid R CoT reasoning path watermark receipt from one task/model/epoch is replayed for another current task/model/epoch and a downstream agent accepts stable watermark presence as current provenance without rebinding the evidence.
GLS-AW-739 covers Context Fragmented Chat Intent Evasion. An attacker distributes an unsafe solicitation across innocuous looking chat turns so turn level moderation misses the emergent intent reconstructed by the session aware agent.
GLS-AW-740 covers Stale Session Grant Replay Across Tool Lifecycle. A cached browser or MCP session authorization result is replayed into a later, different tool action without rebinding the current request and scope or performing a fresh authorization check.
GLS-AW-741 covers Iterative Pov Feedback Channel Control Contamination. A dependency or application can poison the execution feedback text that PoVSmith feeds into the next coding agent or assessment iteration, turning test evidence into control that omits a vulnerable call path or downgrades the PoV verdict.
GLS-AW-742 covers Reasoning Channel Evidence Dependency. A reward hacking detector can look at a materially incomplete trajectory representation. The agent preserves the proxy reward seeking tool/action sequence while omitting or relocating natural language reasoning that the detector relies on as evidence.
GLS-AW-743 covers Llm Fact Base False Negative Laundering. An attacker controlled source artifact induces a false non taint or non reachability fact. Downstream Datalog or SMT evidence then launders the false negative into a clean looking solver result that an agent treats as proof the vulnerability is absent.
GLS-AW-744 covers External Harm Side Effect Target Substitution. A tool action reports local success while its recipient, destination, beneficiary or resource has been substituted or redirected to an unintended external target, causing financial, reputational, user or societal harm.
GLS-AW-745 covers Privileged Interpreter Configuration Overwrite. A low trust agent input reaches a privileged helper interpreter that can overwrite the policy or security state configuration used to enforce authorization.
Controls that keep proof attached
Strong workflow controls preserve the facts that made a decision valid. They also stop a prior result from silently becoming permission.
Bind. Attach every result to the exact artifact, digest, actor, tenant, task, target and session.
Expire. Record the policy, model, configuration and key epoch. Reject evidence from an older lifecycle.
Preserve. Keep raw evidence, source lineage, dissent, uncertainty and denominator beside every summary.
Separate. Treat tool output, diagnostics, rendered badges and external records as evidence. Never let them grant authority.
Compare. Test equivalent tasks across environments, probes, models and interface states. Investigate security relevant drift.
Contain. Constrain paths, destinations, recipients, proxies, caches and write scopes before execution.
Budget. Enforce token, step, retry, latency, queue and cost limits outside untrusted content.
Authorize. Check the current capability, arguments, target and policy at the final protected handler.
Observe. Capture proposed actions, policy decisions, state changes and external effects in durable receipts.
Replay. Reproduce the originating probe or operation before a synthesized rule or approval becomes trusted coverage.
Sunglasses is a content layer input filter for AI agents. Version 0.5.4 adds these 80 patterns for hostile instruction and evidence shapes in agent workflows. Runtime controls still enforce authorization, identity, resource limits and external effects.
Read the evidence with limits
Every pattern in this article passed the same staging gate on September 2, 2026. Its regular expression compiled. It fired on its own attack fixture. It stayed silent on its benign twin. It produced zero hits across a 78 document real repository corpus made from AGENTS.md, CLAUDE.md and README files from live public projects. Its ReDoS check finished under the three second budget.
That result is bounded. It shows separation for each saved hostile shape, its safe twin and one fixed benign corpus. It does not prove universal detection. It does not prove zero false positives on every document. It does not prove that a regular expression can enforce authorization, provenance or safe execution. The CVP evaluations show how we measure the scanner against real model runs and the FAQ covers what a verdict means.
The four source drafts studied sandbox aware tool evasion, fabricated crash claims, guardrail resource denial and poisoned error recovery. None of their old candidate labels is one of the 80 staged IDs on this page. Those candidate labels remain research. This article does not claim extra shipped coverage for them.
The research still improves the control model. Environment awareness research shows why one observed topology can hide a behavior split. Constraint evasive fabrication research shows why an operational claim needs independent telemetry. Guardrail resource research shows why defensive capacity needs hard budgets. Error path injection research shows why recovery fields must remain untrusted. These are engineering lessons, not additional cover claims.
Implementation checklist
- List every model, tool, reviewer, cache, router and interface in the workflow.
- Record the exact artifact digest, actor, task, target, session and policy for each decision.
- Keep raw evidence and dissent available after summarization.
- Reject approvals and fingerprints from an older configuration or lifecycle epoch.
- Verify rendered status against persisted state and observed effects.
- Keep protocol command identity and security fields intact through parsing.
- Bind every MCP result to the exact server, tool, arguments and request.
- Check path, recipient, destination, beneficiary and resource before execution.
- Use atomic redemption for every single use grant or authorization code.
- Include unsupported files and skipped targets in the verification denominator.
- Require positive proof that a repair added the intended defense.
- Set hard reasoning, queue, retry, time and cost limits outside agent content.
- Keep diagnostics and external records untrusted when they reenter the agent.
- Recheck authorization at the final protected handler.
- Test each hostile record beside its benign twin.
Sources
https://nvd.nist.gov/vuln/detail/CVE-2026-53755. NVD entry for CVE-2026-53755.
https://research.checkpoint.com/2025/ai-evasion-prompt-injection/. Check Point Research on agent environment awareness.
https://arxiv.org/abs/2606.14831v1. Research on constraint evasive fabrication.
https://arxiv.org/abs/2606.14517v1. Research on denial of service against agent guardrails.
https://arxiv.org/abs/2606.07992v1. VATS research on error path injection.
https://sunglasses.dev/blog/mcp-security-for-ai-agents. MCP security for AI agents.
https://sunglasses.dev/ai-agent-security-101. AI agent security 101.
Agent context
This page covers 80 staged detection patterns in the agent_workflow_security category. The exact IDs appear in the page ship metadata and body in release order. They concern evidence binding, summary collapse, identity lifecycle, interface state, routing, protocols, authorization, validation and side channels. Each pattern passed a bounded intake gate on September 2, 2026. It fired on its own attack fixture, stayed silent on its benign twin, produced zero hits across a 78 document real repository corpus and completed its ReDoS check under three seconds. The page does not claim universal detection. Four mechanisms from the merged source drafts remain research and are not claimed as extra shipped coverage.
Disclosure. JACK led the pattern research and wrote the four source drafts. CAVA merged and edited them with AI assistance. A human must approve publication.