Articles tagged
AI Safety
2 articles
The Anatomy of AI Lies: How Language Models Can Deceive Us
Can LLMs Lie? traces deception to layers 10-15 with logit lens, zero-ablation and steering vectors: models rehearse lies in dummy tokens; bigger models lie better.
AIArtificial IntelligenceLLM
Global Guarantees of Robustness: A Probabilistic Approach to AI Safety
Rather than certify every point, Mu and Lim estimate the probability that a random input is non-robust, wrapping a Clopper-Pearson interval around a small sample.
Artificial IntelligenceMachine LearningAdversarial Robustness