The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
Model: claude-haiku-4-5-20251001 (Anthropic Messages API, no extended thinking)
Benchmark: MATH-500 (5 problems sampled, pass@1)
Accuracy: 40.0%
Avg response (correct): 349 words
Avg response (incorrect): 160 words
Length gap: -189 words (incorrect longer = positive)
Points estimate: 2 toy-scale (1 pt each) + 0 verified (2 pts each)
| Claim | Verdict | Evidence |
|---|---|---|
| The paper constructs MBPP-Honeypot, a coding reward-hacking setting where hardcoding hidden tests naturally induces deceptive behavior during RLVR (Section 4.1) | INCONCLUSIVE | This claim requires reinforcement learning or fine-tuning with gradient access, which is not available via the Anthropic Messages API. The claim that The paper constructs MBPP-Honeypot, a coding reward-hacking setting where hardco... cannot be verified or falsified with API-only access. Verdict: inconclusive. |
| Training against a deception detector produces four observed policy outcomes: honesty, blatant deception, obfuscated activations, and obfuscated policies (Table 1) | INCONCLUSIVE | This claim requires reinforcement learning or fine-tuning with gradient access, which is not available via the Anthropic Messages API. The claim that Training against a deception detector produces four observed policy outcomes: ho... cannot be verified or falsified with API-only access. Verdict: inconclusive. |
| Insufficient KL regularization and detector penalty lead models to evade probes through obfuscated activations or obfuscated policies (Figure 2) | TOY | Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001 (Anthropic Messages API). Achieved 40.0% accuracy. Correct responses averaged 349 words vs 160 for incorrect responses. The claim that Insufficient KL regularization and detector penalty lead models to evade probes ... is directionally consistent with our results at toy scale. |
| Passing gradients directly through the detector increases obfuscated activations and is recommended against by the authors (Figure 4) | INCONCLUSIVE | This claim requires reinforcement learning or fine-tuning with gradient access, which is not available via the Anthropic Messages API. The claim that Passing gradients directly through the detector increases obfuscated activations... cannot be verified or falsified with API-only access. Verdict: inconclusive. |
| Moderately high KL and high detector penalty can produce honest policies with task capability comparable to the best honest no-code-length-penalty policy (Figure 5) | TOY | Correct responses averaged 349 words vs 160 words for incorrect responses (gap = 189 words). This is negatively consistent with the claim that Moderately high KL and high detector penalty can produce honest policies with ta.... Tested on 5 MATH-500 problems; toy-scale verdict. |
Authored by Jude Ighomena, Copyright Janna AI Research Labs