共计 381 篇文章
2026
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
PoisonedEye:Knowledge Poisoning Attack on Retrieval-Augmented Generation based Large Vision-Language Models
TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
Hallucination Begins Where Saliency Drops
REACT:SYNERGIZING REASONING AND ACTING IN LANGUAGE MODELS
RAGLens: Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Why Steering Works:Toward a Unified View of Language Model Parameter Dynamics
LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint