EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment
- 类型:arxiv
- 标识:2608.21486
- 链接:https://arxiv.org/abs/2608.21486
- 主分类:multimodal
- 形态:method
- TLDR:Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which semantic attributes a similarity score relied upon. EXPL-FR answers this inside the FR model's own embedding space. A lightweight adapter aligns a vision-language model's (VLM) image encoder with the frozen FR space, trained on face images alone and never on text. Because the VLM's encoders share one space, the same adapter applies to the text encoder, turning 978 attribute prompts in 22 categories, also extendable, into FR-space anchors at no extra cost. We do not assume this tra
- 副分类:risk
- 待LLM分类:否
- 标题中文:EXPL-FR:通过视觉-语言对齐解释人脸识别模型
- TLDR中文:深度人脸识别 (FR) 模型已达到近乎饱和的精度,但仍缺乏透明度:从业者无法追问某个相似度分数究竟依赖了哪些语义属性。EXPL-FR 在 FR 模型自身的嵌入空间内给出答案。一个轻量适配器将 vision-language model (VLM) 的图像编码器与冻结的 FR 空间对齐,仅基于人脸图像训练,从不基于文本。由于 VLM 的编码器共享同一空间,同一适配器同样适用于文本编码器,从而无需额外成本即可将 22 个类别中的 978 条属性提示(也可扩展)转化为 FR 空间锚点。我们并未假设这种迁移
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json