共计 381 篇文章
2026
Trap: Mitigating Poisoning-Based Backdoor Attacks by Treating Poison With Poison
Hallucination as Exploit: Evidence-Carrying Multimodal Agents
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
IN-CONTEXT SHARPNESS AS ALERTS: AN INNER REPRESENTATION PERSPECTIVE FOR HALLUCINATION MITIGATION
CYBER-ZERO: Training Cybersecurity Agents without Runtime
S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
Where Does Reasoning Break? Step-Level Hallucination Detection via Hidden-State Transport Geometry
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks