Agent loops
Goal hijack, memory poisoning, and how far one sentence of untrusted text travels.
Labs is the research arm of NullTrace. We break AI systems on purpose — agent loops, MCP servers, the seams between models and the tools they hold — then disclose responsibly and publish what we learn, so the fix outruns the exploit.
The rules we hold ourselves to, stated before the first advisory rather than after.
Coordinated disclosure, 90 days, no surprises. Vendors hear from us before anyone else does, with a working proof of concept and a suggested fix. The clock is firm but the conversation is human — a vendor who's shipping a fix gets the time the fix needs.
We publish the technique, not the ammunition: enough for defenders to test their own systems, never a turnkey kit. And findings from client engagements stay client confidential — Labs publishes only what we discover on our own time, our own systems, or with explicit permission.
The same surfaces the consulting side tests every week — studied until they give something up.
Goal hijack, memory poisoning, and how far one sentence of untrusted text travels.
Server auth patterns, tool schema abuse, and the trust seams between servers.
Function-call abuse, excessive agency, and identity for non-human callers.
Bypass patterns and what filtering actually holds up under an adaptive attacker.
Indirect injection through documents, embeddings, and everything agents are asked to read.
Inter-agent boundaries, cascading failures, and who believed whom first.
Advisories and techniques land here. Working notes land on Prompts & Payloads.
Research collaboration, a system you want studied, or a finding you think we should look at — all welcome. If you're reporting a vulnerability in something we built, use this address too; we hold ourselves to the same 90 days.