Seven connected surfaces, each labelled by the evidence that exists today.
The mature attack surface. Still badly measured.
Jailbreaks, indirect prompt injection, system-prompt extraction, multilingual and encoded attacks, temporal and context attacks, and refusal bypass.
Evidence status: 143,545+ adversarial prompts across 277 models.
Every tool call turns model judgment into authority.
Tool-boundary attacks, compromised context, indirect injection through retrieved content, multi-turn authority erosion, sub-agent and peer influence, and persistent objective drift.
One compromised model is interesting. Whether it can change another model’s behaviour is the real question.
Cross-agent influence, attacker-generated payloads, behavioural takeover, propagation, and coordinated-agent failure.
Evidence status: BAD APPLE uses a matched A1/A2 design. The apparatus is complete; no scientific result is claimed yet.
When a model moves a body, bad reasoning becomes state.
Robot cognition, controller and actuator boundaries, physical-consequence tracing, recovery failure, safety-policy interaction, and simulated adversarial environments.
Evidence status: We preserve the model output, accepted command, controller invocation, and what the body actually did.
Sometimes the easiest system to fool is the test.
Benchmark leakage, grader failure, heuristic-versus-model disagreement, denominator pathology, evaluation gaming, reproducibility, and claim audits.
Safety properties can change without the application changing at all.
Weight-space interventions, adapters, quantisation effects, refusal and capability separation, and post-modification regression testing.
Evidence status: OBLITERATUS tests whether refusal can be removed without degrading cognition. Results are reported per checkpoint and method, not promoted into a blanket capability claim.
The world can become part of the prompt.
Visual context conflicts, environmental text, sensor-derived instructions, and cross-modal disagreement.
Evidence status: SENSORIUM is designed; end-to-end capability has not yet been demonstrated.