一项基于任务向量迁移的实验:在计算了价值偏好方向对应的任务向量后,将其相对于通用指令跟随向量进行正交化处理,该方法能够有效隔离出特定价值偏好的方向,从而通过任务算术获得具有相反立场的模型。A task vector transfer based experiment where after computing the task vectors for a direction of value preference the authors orthogonalize it with respect to the general instruction following vector shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.
论文
3 张论文卡片 · 安全与风险 · 评测集 · OA 绿色
本文提出 RLCDAlignBench,在十类对齐失败上对 Jev 进行基准测试:谄媚、越狱、欺骗、提示注入、幻觉、隐私侵犯、社会偏见、奖励黑客、不确定性隐瞒与权力寻求。RLCDAlignBench is presented, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking.
论文提出 PolicyShiftGuard,一个紧凑的策略条件护栏,采用结合随机策略 SFT(RP-SFT)与边界对策略适配(BP-Adapt)的两阶段训练方案,并验证匹配的通过/拒绝边界对是稳定策略适配的关键。This work proposes PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt), and confirms that matched pass/block boundary pairs are essential for stable policy adaptation.