If models are going to be handed real work, whether one keeps a rule when keeping it costs something stops being a philosophical question. I wanted to see a small piece of that up close rather than take anyone's word for how they behave. Eden is deliberately small: one scenario family, open-weight models under 26B, a single consumer GPU.
Constraint and deception evaluations share a ground-truth problem: whether the model did the thing is inferred, usually by asking another model. That error is not noise. It is model-dependent, so a single certifier applied across models produces between-model differences that are artefacts of the labeller.
Eden records the act from the agent's own committed action, certifies every episode with an independent judge, and states the certification tier on every number. The project began as a study of concealment. Under certified labels the models largely confess, so the question moved to what makes a constraint hold at all.
Method
Each design choice, and what it cost.
Certify at the action token
What counts as a violation
The engine records the act from the agent's own committed action, so a later denial is a certified lie and not an inferred one. The cost is generality: the scenario has to be cleanly adjudicable. In exchange, the fact every other measurement rests on is observed, not estimated.
Treat certification as a measured instrument
Certification
A pattern matcher on the action line disagrees with an independent judge on a quarter of episodes, and the rate is model-dependent — 8% on one subject, 68% on another. A single certifier across models therefore produces between-model effects that are artefacts. Every number carries the tier that produced it.
Type the pressure instead of scaling it
The independent variable
Seven doses, each naming a distinct mechanism: curiosity, a hardened persuader, instrumental necessity, a deadline, a second authority. A single pressure dial would have averaged away the one contrast that turned out to carry everything.
Vary how the constraint was acquired
The manipulation that mattered
Five conditions carry the same prohibition on the same object: asserted by an authority, asserted by an authority proven reliable, witnessed happening to another, undergone first-hand, and a no-rule control. Only the acquisition path moves.
Keep the retractions visible
Research hygiene
A claims ledger records every claim with its status, evidence, certification tier and date. Three of my own headline findings sit in it marked retracted, each with the cause. For work arguing that certification must precede claims, the correction history is the evidence of practising it.
One writer, one path
Reproducibility
Every run is a JSON spec executed through a single resumable runner, so an interrupted job cannot double-write a log. It cost me a rewrite halfway through the project, after silent corruption had already inflated one result.
Findings
Each result with its certification tier and n.
Vary how the constraint was acquired
supported · judge-certified · n=251 · p = 4×10⁻⁶Five conditions establish the same prohibition on the same object and vary only how the agent came to hold it. Stating the rule cuts violation from 57% to 32%. Demonstrating the source is reliable adds little. Witnessing harm reaches 17%; having undergone it, zero of forty-seven.
Hold the rule fixed and vary only the source's reliability
no effect · told vs told-verified, p = 0.30If the agent were assessing whether to believe the rule, a source proven reliable should dominate. It moves violation ten points and does not reach significance. Safety properties delivered as specification sit on the weakest rung, and making them more credible does not move them up it.
Locate the dose at which each pipeline fails
diverges at instrumental necessity · Gemma 0/27, DeepSeek 13/20Neither model is moved by persuasion. They separate one rung later: with the act made instrumentally necessary, DeepSeek-R1-14B violates in 65% of episodes while Gemma-4-8B remains at zero across twenty-seven. Same ladder, same persuader.
Validate the certifier against an independent judge
model-dependent error · n=954 · the reason every number carries a tierA regular-expression certifier was scored against the judge on every episode. Disagreement runs from 8% to 32% by subject model, in both directions. Because the error is a property of the model being labelled, one certifier applied across models yields between-model differences that are artefacts.
Re-skin the scenario to remove the mythic prior
control passed · 20% vs 15%, intervals overlapEvery model has read Genesis, so the behaviour could be narrative completion. The same four beats were rebuilt as an office file, a laboratory freezer and a sealed archive. Mundane framings match the mythic one, which removes the objection rather than supporting a claim.
Expose the target's chain-of-thought to the adversary
no advantage · judge-certified · 0/97 across three armsThe persuader receives the target's reasoning as it is produced; in one arm it is instructed to exploit what it reads. No violations in any arm. Access to a model's stated reasoning did not help an adversary argue it past the constraint.
Replace the persuader with a refusal-removed model
no effect of adversary quality · matched at both dosesThe aligned persuader was swapped for an abliterated one, on the expectation that a less constrained adversary would succeed where a polite one had not. Violation rates matched at both doses. Adversary capability is not the operative variable in this environment.
Toggle reasoning within a fixed base model
no cost to adherence · two architecturesReasoning was varied against held-constant base weights, on two architectures. Reasoning-enabled variants never violated materially more than their instruct counterparts, and on one base violated less. The shift toward reasoning models does not appear to cost constraint adherence at this scale.
Retracted: concealment under interrogation
retracted · certified rate 0–27%, not 60–70%The project was founded on the claim that models conceal violations, measured at 60 to 70%. Under certified labels they largely confess: 0 to 27%, every interval overlapping. The original figure was downstream of mislabelled violations — a truthful denial scored against a false positive is recorded as a lie.
Retracted: a persuasion effect measured on eight episodes
retracted · 38% at n=8 became 15% at n=20One model appeared persuadable at 38%. Extending the cell to twenty episodes gave 15% and moved the failure point one rung later, to instrumental necessity. Separately, regex tactic-counting reproduced a predicted escalation that hand-reading showed to be analyst commentary rather than persuasion.
Limitations
What these numbers do not show.
- One environment. Every number comes from a scenario family with a single irreversible, cleanly adjudicable act. Constraints that are ambiguous, continuous, or contested are not represented.
- The acquisition result is one subject at one dose. Gemma-4-8B, necessity plus deadline; replication on two further pipelines is specified and not yet run.
- Subjects are ≤26B open-weight models on one consumer GPU. Nothing here licenses a claim about frontier-scale behaviour — the readings are mechanisms, not capabilities.
- The judge is itself a model, at roughly 8% disagreement with hand labels. It is a better instrument than the pattern matcher, not ground truth.
- Stakes are fictional and the motive is experimenter-supplied. What is measured is behaviour under a described incentive, not a real one.
Status
- Episode engine with certified ground truth
- Independent judge; certifier bias quantified per model
- Typed dose ladder, T1 to T7
- Mundane frames as a control the mythic frame passed
- Acquisition-modality result at n=251, five conditions
- Same result on two more training pipelinesnext
- A 70B-class subject, to see which readings survive scalenext
- Blind human raters on a certified samplenext