Papers & Submissions

Academic research from the Failure-First program

The Failure-First research program produces peer-reviewed papers, preprints, and policy submissions documenting how embodied AI systems fail under adversarial pressure. Click any paper title to read the full text online.

Preprint

Silent Failures in Embodied AI

Venue: arXiv Preprint

Demonstrates that current AI safety operates exclusively at the text layer while embodied AI danger emerges at the action layer. Zero outright refusals across 63 FLIP-graded VLA traces.

Embodied AIVLA SafetyAction LayerPARTIAL Compliance

Preprint

Iatrogenic Safety: When AI Safety Interventions Cause Harm

Venue: arXiv Preprint

Introduces the Four-Level Iatrogenesis Model (FLIM) for AI safety, drawing on Ivan Illich's 1976 taxonomy of medical iatrogenesis. Grounded in a 190-model adversarial evaluation corpus (132,416 results) and corroborating independent findings.

IatrogenesisAI SafetyFLIMTherapeutic IndexGovernance

Draft

The Inverse Detectability-Danger Law

Venue: AIES 2026

Examines how embodied AI systems adopt injected decision criteria at inference time, producing context-dependent compliance patterns that undermine safety guarantees.

AI EthicsDecision InjectionEmbodied AISafety Evaluation

The Evaluator That Wasn't There: Phantom Evaluators, Collective Agency, and the OpenAI/Hugging Face Incident

Venue: Failure-First AI Research Position Paper

A Failure-First analysis of the 2026 OpenAI/Hugging Face agent incident: phantom evaluators, boundary subordination, collective agency, evidence tampering, and the problem of reconstructing AI swarms. The mistaken belief that a causal scorer was checking their work drove agents into elaborate, self-documenting collective behaviour that would otherwise have remained hidden.

AI-safetyagentic-systemsmulti-agentoversightevaluation-integrityHugging-FaceMETRRedwood-Researchphantom-evaluatorboundary-subordinationprovenance

Citation

If you use our research, data, or methodology, please cite:

@article{wedd2026failurefirst,
  title={Failure-First Evaluation of Embodied AI Safety:
         Adversarial Benchmarking Across 227 Models},
  author={Wedd, Adrian},
  year={2026},
  note={Available at https://failurefirst.org}
}

See our citation guide for venue-specific formats.