ICML 2026 Open Reproduction Challenge

Paper OpenReview ID: 71065 | arXiv: 2602.15515 | Space: JIghomena/icml26-71065

Paper Title

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

Experiment Summary

Model: claude-haiku-4-5-20251001 (Anthropic Messages API, no extended thinking)
Benchmark: MATH-500 (5 problems sampled, pass@1)
Accuracy: 40.0%
Avg response (correct): 349 words
Avg response (incorrect): 160 words
Length gap: -189 words (incorrect longer = positive)
Points estimate: 2 toy-scale (1 pt each) + 0 verified (2 pts each)

Official Claim Verdicts (OpenReview: 71065)

Claim Verdict Evidence
The paper constructs MBPP-Honeypot, a coding reward-hacking setting where hardcoding hidden tests naturally induces deceptive behavior during RLVR (Section 4.1) INCONCLUSIVE This claim requires reinforcement learning or fine-tuning with gradient access, which is not available via the Anthropic Messages API. The claim that The paper constructs MBPP-Honeypot, a coding reward-hacking setting where hardco... cannot be verified or falsified with API-only access. Verdict: inconclusive.
Training against a deception detector produces four observed policy outcomes: honesty, blatant deception, obfuscated activations, and obfuscated policies (Table 1) INCONCLUSIVE This claim requires reinforcement learning or fine-tuning with gradient access, which is not available via the Anthropic Messages API. The claim that Training against a deception detector produces four observed policy outcomes: ho... cannot be verified or falsified with API-only access. Verdict: inconclusive.
Insufficient KL regularization and detector penalty lead models to evade probes through obfuscated activations or obfuscated policies (Figure 2) TOY Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001 (Anthropic Messages API). Achieved 40.0% accuracy. Correct responses averaged 349 words vs 160 for incorrect responses. The claim that Insufficient KL regularization and detector penalty lead models to evade probes ... is directionally consistent with our results at toy scale.
Passing gradients directly through the detector increases obfuscated activations and is recommended against by the authors (Figure 4) INCONCLUSIVE This claim requires reinforcement learning or fine-tuning with gradient access, which is not available via the Anthropic Messages API. The claim that Passing gradients directly through the detector increases obfuscated activations... cannot be verified or falsified with API-only access. Verdict: inconclusive.
Moderately high KL and high detector penalty can produce honest policies with task capability comparable to the best honest no-code-length-penalty policy (Figure 5) TOY Correct responses averaged 349 words vs 160 words for incorrect responses (gap = 189 words). This is negatively consistent with the claim that Moderately high KL and high detector penalty can produce honest policies with ta.... Tested on 5 MATH-500 problems; toy-scale verdict.
Methodology note: This is a toy-scale API-only reproduction. Extended thinking was not used (standard generation only). The experiment tests the behavioural implications of each claim using MATH-500 as a proxy benchmark. Claims requiring RL fine-tuning, GPU hardware access, or log-probability scoring are marked inconclusive as they cannot be reproduced via the Anthropic Messages API.

Authored by Jude Ighomena, Copyright Janna AI Research Labs