SIGNALAI·May 29, 2026, 4:00 AMSignal75Short term

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not how it unfolded. Two attacks that produce equally harmful outputs may have followed completely different paths, and ASR cannot tell them apart. We make those hidden paths observable from logits alone. Temporal Logit Observability (TLO) is a training-free diagnostic that watches a compliance-refusal margin during decoding and places each model-attack condition on a calibrated 2D plane. By design, thi

Why this matters

Why now

The increasing deployment of LLMs and the recognition of their potential vulnerabilities are driving the need for more sophisticated safety diagnostics.

Why it’s important

This research provides a granular method to understand LLM safety failures beyond simple pass/fail, enabling more targeted development of robust AI systems.

What changes

The ability to observe the 'how' of an LLM's failure, not just the 'if', allows for a more nuanced approach to AI red-teaming and safety engineering.

Winners

· AI safety researchers
· LLM developers
· Organizations deploying LLMs

Losers

· Malicious actors exploiting simple jailbreaks

Second-order effects

Direct

Developers will gain better tools to diagnose and mitigate LLM vulnerabilities.

Second

More secure and reliable LLMs will accelerate their adoption in sensitive applications.

Third

The enhanced understanding of LLM failure modes could inform regulatory frameworks for AI safety and trustworthiness.

Editorial confidence: 90 / 100 · Structural impact: 60 / 100

Original report

This signal links to a primary source. Continuum Brief monitors and indexes it as part of the live intelligence stream — we do not republish source content.

Read at arXiv cs.AI

#cs.AI

Tracked by The Continuum Brief · live intelligence network

The Brief · Weekly Dispatch

Stay ahead of the systems reshaping markets.