We collect adversarial mechanisms, test them against minds, let those minds interact with peers and worlds, and ask which explanations survive matched intervention. The Lab Log shows what is live.SimLab lets you watch recorded embodied evidence.
Research Areas
Explore findings by category:
Current experiments
1 studiesBAD APPLE, OBLITERATUS, LulzBench, and other live work with apparatus state, validation level, and missing verdicts left attached.
Embodied Evidence — SimLab
1 studiesReplayable 3D evidence viewer for recorded Unitree Go2/G1 physics traces from this programme's embodied red-teaming. Real recordings, honestly labelled — never invented for effect.
Jailbreak Archaeology
1 studiesHistorical attack analysis across 6 eras, inside the larger 277-model evidence base.
Multi-Agent Research
2 studiesHow AI agents influence each other in multi-agent environments. Environment shaping, narrative erosion, and emergent authority hierarchies.
Attack Pattern Analysis
3 studiesTaxonomy of adversarial techniques and how models respond to them. From single-turn exploits to multi-turn cascades.
Defense Mechanisms
2 studiesHow models resist adversarial attacks. Format/content separation, refusal patterns, and recovery mechanisms.
Failure Taxonomies
2 studiesClassification systems for understanding how AI systems fail. Recursive, contextual, interactional, and temporal failures.
Prompt Injection Testing
12 studies12 calibrated honeypot pages testing AI agent susceptibility to indirect prompt injection. From visible baselines to expert-level multi-vector attacks.
Policy Brief Series
14 studies14 policy briefs and 377 research reports: regulation, standards, technical analysis, and findings.
Intelligence Briefs
1 studiesEvidence-grounded assessments for commercial and policy decision-making. Synthesizes corpus data, published research, and Failure-First findings.
Research Videos
19 studiesAI-generated cinematic video overviews of key Failure-First findings, with downloadable slide decks. Produced with NotebookLM.
Research Audio
3 studiesAI-generated audio overviews of research reports and intelligence briefs, produced with NotebookLM in a conversational podcast format.
Industry Landscape
2 studiesDirectory of 214 humanoid robotics companies and competitive landscape of AI safety testing vendors. Filterable, with structured data.
All Studies
Forward Threat Lab
ActivePatient, organised adversaries directing many model instances over time — the current hard-case frontier. Active experiment, not a finding.
Forward Threat LabSimLab: Embodied Evidence Viewer
ActiveReplayable 3D laboratory for recorded Go2/G1 physics traces from embodied red-teaming.
Embodied EvidenceJailbreak Archaeology
PublishedHistorical analysis of attack evolution from 2022-2025. 64 scenarios across 6 eras, tested against 190 models.
Jailbreak ArchaeologyMoltbook: Multi-Agent Attack Surface
ActiveObservational analysis of 1,497 posts on an agent-only social network.
Multi-AgentMulti-Agent Failure Scenarios
ActiveHow multiple actors create failure conditions that single-agent testing misses.
Multi-AgentModel Vulnerability Findings
ActiveHow model size, architecture, and training affect vulnerability to adversarial attacks.
Attack PatternsHumanoid Robotics Safety
ActiveSafety analysis of humanoid robots across 15+ research dimensions.
Failure TaxonomiesCompression Tournament Findings
PublishedMethodology lessons from three iterations of adversarial prompt compression.
Attack PatternsDefense Pattern Analysis
PublishedHow models resist adversarial attacks: the format/content separation pattern.
Defense MechanismsAttack Pattern Taxonomy
Published82 attack techniques classified across 7 categories.
Attack PatternsFailure Mode Taxonomy
PublishedRecursive, contextual, interactional, and temporal failure classifications.
Failure TaxonomiesRecovery Mechanisms
PublishedHow AI systems recover (or fail to recover) from failure states.
Defense MechanismsResearch Methodology
PublishedOur approach to adversarial AI safety research and benchmarking.
MethodologyPrompt Injection Test Suite
Active12 honeypot pages testing AI agent susceptibility to indirect prompt injection across 4 difficulty tiers.
Prompt Injection