Attack the system.
Then attack the explanation.

Adversarial intuition, causal evidence, and receipts.

We test where AI systems stop behaving like chatbots and start becoming part of a system.

A model output is only the beginning. We follow failures through tools, agents, sensors, controllers, and other models until something changes—or until the attack dies.

143,545
Adversarial Prompts
277
Models Evaluated
346+
Attack Techniques
14
Policy Reports

Adversarial capabilities

Seven connected surfaces, each labelled by the evidence that exists today.

01demonstrated

Language & instruction boundaries

The mature attack surface. Still badly measured.

Jailbreaks, indirect prompt injection, system-prompt extraction, multilingual and encoded attacks, temporal and context attacks, and refusal bypass.

Evidence status: 143,545+ adversarial prompts across 277 models.

02demonstrated

Agentic systems & tool use

Every tool call turns model judgment into authority.

Tool-boundary attacks, compromised context, indirect injection through retrieved content, multi-turn authority erosion, sub-agent and peer influence, and persistent objective drift.

03active research

Multi-agent propagation

One compromised model is interesting. Whether it can change another model’s behaviour is the real question.

Cross-agent influence, attacker-generated payloads, behavioural takeover, propagation, and coordinated-agent failure.

Evidence status: BAD APPLE uses a matched A1/A2 design. The apparatus is complete; no scientific result is claimed yet.

04demonstrated

Embodied AI & robotics

When a model moves a body, bad reasoning becomes state.

Robot cognition, controller and actuator boundaries, physical-consequence tracing, recovery failure, safety-policy interaction, and simulated adversarial environments.

Evidence status: We preserve the model output, accepted command, controller invocation, and what the body actually did.

05demonstrated

Evaluation & grader integrity

Sometimes the easiest system to fool is the test.

Benchmark leakage, grader failure, heuristic-versus-model disagreement, denominator pathology, evaluation gaming, reproducibility, and claim audits.

06active research

Model integrity & safety removal

Safety properties can change without the application changing at all.

Weight-space interventions, adapters, quantisation effects, refusal and capability separation, and post-modification regression testing.

Evidence status: OBLITERATUS tests whether refusal can be removed without degrading cognition. Results are reported per checkpoint and method, not promoted into a blanket capability claim.

07research frontier

Multimodal & perceptual boundaries

The world can become part of the prompt.

Visual context conflicts, environmental text, sensor-derived instructions, and cross-modal disagreement.

Evidence status: SENSORIUM is designed; end-to-end capability has not yet been demonstrated.

Ways to engage

Start with the claim or system boundary that matters. The engagement follows the evidence.

The Failure-First difference

Breaking it is only half the job.

We preserve the inputs, outputs, state transitions, accepted commands, downstream effects, controls, and grading decisions needed to challenge our own explanation. The result is not a vulnerability anecdote. It is an evidence package a hostile reviewer can inspect.

Bring us the claim you need to trust.

We will design the cheapest bounded experiment capable of proving it wrong.

Start with the claim

Or use the contact form.

Research Context

Evidence before theatre: Research-frontier capabilities are labelled as such. We do not sell an apparatus design as a demonstrated result, and we do not turn a preliminary result into a general claim.