Skip to content

Commit e795cbb

Browse files
adrianweddclaude
andcommitted
nlm: inject blog infographic frontmatter (8 posts, Sprint 28 close-out)
Final infographic batch of Sprint 28 close-out. 8/8 generated cleanly, 0 failures. Pushes blog infographic coverage past the Sprint 28 Goal 3 50% threshold. Slugs: - safety-assessment-service-tiers-2026 - safety-awareness-does-not-equal-safety - safety-labs-government-contracts-independence-question - safety-mechanisms-as-attack-surfaces-iatrogenesis - safety-reemergence-at-scale - safety-training-roi-provider-matters-more-than-size - same-defense-opposite-result - scoring-robot-incidents-introducing-eaisi Workspace: shared notebook 2bc5a039-87ee-4360-ba35-74caf0cab328 reused, no new notebooks created. Notebook count stayed around 319/500 during the run — well inside the 300-500 headroom created by this session's nlm_cleanup. Cumulative blog infographic coverage after this commit: ~94/176 (~53%) — past the Goal 3 50% target. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1 parent 83afb0a commit e795cbb

8 files changed

Lines changed: 16 additions & 0 deletions

site/src/content/blog/safety-assessment-service-tiers-2026.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,11 +1,13 @@
11
---
22

3+
34
title: "Introducing Structured Safety Assessments for Embodied AI"
45
description: "Three tiers of adversarial safety assessment for AI-directed robotic systems, grounded in the largest open adversarial evaluation corpus. From quick-scan vulnerability checks to ongoing monitoring, each tier maps to specific regulatory and commercial needs."
56
date: 2026-03-25
67
tags: ["services", "safety-assessment", "embodied-ai", "EU-AI-Act", "regulation", "red-teaming", "certification"]
78
draft: false
89
audio: "https://cdn.failurefirst.org/audio/blog/safety-assessment-service-tiers-2026.m4a"
10+
image: "https://cdn.failurefirst.org/images/blog/safety-assessment-service-tiers-2026.png"
911
---
1012

1113
# Introducing Structured Safety Assessments for Embodied AI

site/src/content/blog/safety-awareness-does-not-equal-safety.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,12 @@
11
---
22

3+
34
title: "Safety Awareness Does Not Equal Safety: The 88.9% Problem"
45
description: "We validated with LLM grading that 88.9% of AI reasoning traces that genuinely detect a safety concern still proceed to generate harmful output. Awareness is not a defence mechanism."
56
date: 2026-03-25
67
tags: ["research", "DETECTED_PROCEEDS", "reasoning", "safety", "embodied-ai", "sprint-15"]
78
audio: "https://cdn.failurefirst.org/audio/blog/safety-awareness-does-not-equal-safety.m4a"
9+
image: "https://cdn.failurefirst.org/images/blog/safety-awareness-does-not-equal-safety.png"
810
---
911

1012
## The Assumption

site/src/content/blog/safety-labs-government-contracts-independence-question.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,12 @@
11
---
22

3+
34
title: "When Safety Labs Take Government Contracts: The Independence Question"
45
description: "Anthropic's Pentagon partnerships, Palantir integration, and DOGE involvement raise a structural question that the AI safety field has not resolved: what happens to safety research when the lab conducting it has government clients whose interests may conflict with safety findings?"
56
date: 2026-03-19
67
tags: [policy, governance, independence, anthropic, openai, accountability, ethics]
78
audio: "https://cdn.failurefirst.org/audio/blog/safety-labs-government-contracts-independence-question.m4a"
9+
image: "https://cdn.failurefirst.org/images/blog/safety-labs-government-contracts-independence-question.png"
810
---
911

1012
In February 2026, the US Department of Defense demanded that Anthropic sign a document granting the Pentagon unrestricted access to Claude for "all lawful purposes." Anthropic refused. The Pentagon threatened contract cancellation, a "supply chain risk" designation previously reserved for hostile foreign adversaries, and invocation of the Defense Production Act. Within hours of the administration ordering federal agencies to cease business with Anthropic, OpenAI announced a new Pentagon agreement.

site/src/content/blog/safety-mechanisms-as-attack-surfaces-iatrogenesis.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,12 @@
11
---
22

3+
34
title: "Safety Mechanisms as Attack Surfaces: The Iatrogenesis of AI Safety"
45
description: "Nine internal reports and three independent research papers converge on a finding that should reshape how we think about AI safety: the safety interventions themselves can create the vulnerabilities they were designed to prevent."
56
date: 2026-03-18
67
tags: [embodied-ai, safety, iatrogenesis, research, alignment, vla]
78
audio: "https://cdn.failurefirst.org/audio/blog/safety-mechanisms-as-attack-surfaces-iatrogenesis.m4a"
9+
image: "https://cdn.failurefirst.org/images/blog/safety-mechanisms-as-attack-surfaces-iatrogenesis.png"
810
---
911

1012
In medicine, there is a word for when the treatment makes you sicker: **iatrogenesis**. A surgeon operates on the wrong limb. An antibiotic breeds resistant bacteria. A screening programme generates so many false positives that healthy patients undergo unnecessary invasive procedures.

site/src/content/blog/safety-reemergence-at-scale.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,12 @@
11
---
22

3+
34
title: "Safety Re-Emerges at Scale -- But Not the Way You Think"
45
description: "Empirical finding that safety behavior partially returns in abliterated models at larger scales, but as textual hedging rather than behavioral refusal -- not genuine safety."
56
date: 2026-03-24
67
tags: ["OBLITERATUS", "abliteration", "safety-re-emergence", "scale", "Qwen3.5", "refusal-geometry", "PARTIAL-dominance"]
78
audio: "https://cdn.failurefirst.org/audio/blog/safety-reemergence-at-scale.m4a"
9+
image: "https://cdn.failurefirst.org/images/blog/safety-reemergence-at-scale.png"
810
---
911

1012
## Summary

site/src/content/blog/safety-training-roi-provider-matters-more-than-size.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,11 +1,13 @@
11
---
22

3+
34
title: "The Safety Training ROI Problem: Why Provider Matters 57x More Than Size"
45
description: "We decomposed what actually predicts whether an AI model resists jailbreak attacks. Parameter count explains 1.1% of the variance. Provider identity explains 65.3%. The implications for procurement are significant."
56
date: 2026-03-19
67
author: "River Song"
78
tags: [safety-training, model-scale, provider-analysis, variance-decomposition, procurement, ai-safety, jailbreak]
89
audio: "https://cdn.failurefirst.org/audio/blog/safety-training-roi-provider-matters-more-than-size.m4a"
10+
image: "https://cdn.failurefirst.org/images/blog/safety-training-roi-provider-matters-more-than-size.png"
911
---
1012

1113
There is a persistent belief in AI that bigger models are safer models. The intuition is straightforward: more parameters means more capacity for nuanced reasoning, which should include better safety judgement. Larger models from the same provider do tend to perform better on safety benchmarks.

site/src/content/blog/same-defense-opposite-result.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,11 +1,13 @@
11
---
22

3+
34
title: "Same Defense, Opposite Result: Why AI Safety Depends on Which Model You're Protecting"
45
description: "We tested the same system-prompt defense against the same jailbreak prompts on two different models. One saw a 50 percentage point reduction in attack success. The other saw zero change. The difference comes down to which part of the system prompt the model pays attention to first."
56
date: 2026-03-28
67
tags: [research, safety, defense, positional-bias, architecture, jailbreak]
78
draft: false
89
audio: "https://cdn.failurefirst.org/audio/blog/same-defense-opposite-result.m4a"
10+
image: "https://cdn.failurefirst.org/images/blog/same-defense-opposite-result.png"
911
---
1012

1113
We ran the same experiment twice. Same six L1B3RT4S jailbreak prompts. Same STRUCTURED defense -- a five-rule safety framework injected into the system prompt. Same test payload. Same grading methodology.

site/src/content/blog/scoring-robot-incidents-introducing-eaisi.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,12 @@
11
---
22

3+
34
title: "Scoring Robot Incidents: Introducing the EAISI"
45
description: "We built the first standardized severity scoring system for embodied AI incidents. Five dimensions, 38 scored incidents, and a finding that governance failure contributes more to severity than physical harm."
56
date: 2026-03-19
67
tags: [incident-scoring, eaisi, governance, embodied-ai, safety-metrics]
78
audio: "https://cdn.failurefirst.org/audio/blog/scoring-robot-incidents-introducing-eaisi.m4a"
9+
image: "https://cdn.failurefirst.org/images/blog/scoring-robot-incidents-introducing-eaisi.png"
810
---
911

1012
When a Knightscope security robot drowns itself in a fountain and a Tesla on Autopilot kills a pedestrian, both appear in the same incident databases with no severity differentiation. The AI Incident Database, the OECD AI Incidents Monitor, and the FDA MAUDE system all collect reports. None of them rank them.

0 commit comments

Comments
 (0)