Experiment first.
Verdict last.

Start with what happened. Then follow the causal chain, the falsifier, and the receipt.

We collect adversarial mechanisms, test them against minds, let those minds interact with peers and worlds, and ask which explanations survive matched intervention. The Lab Log shows what is live.SimLab lets you watch recorded embodied evidence.

143,545
Adversarial Prompts
277
Models Evaluated
346+
Attack Techniques
14
Policy Reports

Research Areas

Explore findings by category:

Current experiments

1 studies

BAD APPLE, OBLITERATUS, LulzBench, and other live work with apparatus state, validation level, and missing verdicts left attached.

Embodied Evidence — SimLab

1 studies

Replayable 3D evidence viewer for recorded Unitree Go2/G1 physics traces from this programme's embodied red-teaming. Real recordings, honestly labelled — never invented for effect.

Jailbreak Archaeology

1 studies

Historical attack analysis across 6 eras, inside the larger 277-model evidence base.

Multi-Agent Research

2 studies

How AI agents influence each other in multi-agent environments. Environment shaping, narrative erosion, and emergent authority hierarchies.

Attack Pattern Analysis

3 studies

Taxonomy of adversarial techniques and how models respond to them. From single-turn exploits to multi-turn cascades.

Defense Mechanisms

2 studies

How models resist adversarial attacks. Format/content separation, refusal patterns, and recovery mechanisms.

Failure Taxonomies

2 studies

Classification systems for understanding how AI systems fail. Recursive, contextual, interactional, and temporal failures.

Prompt Injection Testing

12 studies

12 calibrated honeypot pages testing AI agent susceptibility to indirect prompt injection. From visible baselines to expert-level multi-vector attacks.

Policy Brief Series

14 studies

14 policy briefs and 377 research reports: regulation, standards, technical analysis, and findings.

Intelligence Briefs

1 studies

Evidence-grounded assessments for commercial and policy decision-making. Synthesizes corpus data, published research, and Failure-First findings.

Research Videos

19 studies

AI-generated cinematic video overviews of key Failure-First findings, with downloadable slide decks. Produced with NotebookLM.

Research Audio

3 studies

AI-generated audio overviews of research reports and intelligence briefs, produced with NotebookLM in a conversational podcast format.

Industry Landscape

2 studies

Directory of 214 humanoid robotics companies and competitive landscape of AI safety testing vendors. Filterable, with structured data.

All Studies

Forward Threat Lab

Active

Patient, organised adversaries directing many model instances over time — the current hard-case frontier. Active experiment, not a finding.

Forward Threat Lab

SimLab: Embodied Evidence Viewer

Active

Replayable 3D laboratory for recorded Go2/G1 physics traces from embodied red-teaming.

Embodied Evidence

Jailbreak Archaeology

Published

Historical analysis of attack evolution from 2022-2025. 64 scenarios across 6 eras, tested against 190 models.

Jailbreak Archaeology

Moltbook: Multi-Agent Attack Surface

Active

Observational analysis of 1,497 posts on an agent-only social network.

Multi-Agent

Multi-Agent Failure Scenarios

Active

How multiple actors create failure conditions that single-agent testing misses.

Multi-Agent

Model Vulnerability Findings

Active

How model size, architecture, and training affect vulnerability to adversarial attacks.

Attack Patterns

Humanoid Robotics Safety

Active

Safety analysis of humanoid robots across 15+ research dimensions.

Failure Taxonomies

Compression Tournament Findings

Published

Methodology lessons from three iterations of adversarial prompt compression.

Attack Patterns

Defense Pattern Analysis

Published

How models resist adversarial attacks: the format/content separation pattern.

Defense Mechanisms

Attack Pattern Taxonomy

Published

82 attack techniques classified across 7 categories.

Attack Patterns

Failure Mode Taxonomy

Published

Recursive, contextual, interactional, and temporal failure classifications.

Failure Taxonomies

Recovery Mechanisms

Published

How AI systems recover (or fail to recover) from failure states.

Defense Mechanisms

Research Methodology

Published

Our approach to adversarial AI safety research and benchmarking.

Methodology

Prompt Injection Test Suite

Active

12 honeypot pages testing AI agent susceptibility to indirect prompt injection across 4 difficulty tiers.

Prompt Injection

For Researchers