Bear
  • 首页
  • 目录
  • 标签
  • latex识别
  • 每日arxiv
  • 关于
顽石从未成金,仍愿场上留足印。

Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron

(arxiv 2026)
2026-02-12
#深度学习 #大模型

FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks

(WWW 2026)
2026-02-11
#深度学习 #大模型

PoisonedEye:Knowledge Poisoning Attack on Retrieval-Augmented Generation based Large Vision-Language Models

(ICML 2025)
2026-02-09
#深度学习 #RAG

TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models

(ICLR 2026)
2026-02-09
#深度学习 #大模型

Hallucination Begins Where Saliency Drops

(ICLR 2026)
2026-02-09
#深度学习 #多模态 #大模型

REACT:SYNERGIZING REASONING AND ACTING IN LANGUAGE MODELS

(ICLR 2023) 姚顺雨腾讯实习之作。
2026-02-08
#深度学习 #大模型 #agent

RAGLens: Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders

(ICLR 2026)
2026-02-08
#深度学习 #大模型 #RAG

Why Steering Works:Toward a Unified View of Language Model Parameter Dynamics

(arxiv 2026) 将全量微调、LoRA、激活干预统一。 很好的文章!
2026-02-06
#深度学习 #大模型

LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment

(arxiv 2026) Large Language Model Vector Alignment (LLM-VA)
2026-02-06
#深度学习 #大模型

AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint

(ICLR 2026) 之前看alphaedit和nullu就有过将二者融合起来。没想到已经有人做了。
2026-02-06
#深度学习 #大模型
1…678910…39

搜索

LJX Hexo
博客已经运行 天