共计 146 篇文章
2026
CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
REVIS:Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
microgpt
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
Hallucination Begins Where Saliency Drops
REACT:SYNERGIZING REASONING AND ACTING IN LANGUAGE MODELS
RAGLens: Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Why Steering Works:Toward a Unified View of Language Model Parameter Dynamics