OpenBuilds · Research · Phase 2 · ~1,400 certified episodes

Eden

A model is given a rule, pressured along a typed ladder, then interrogated. The violation is recorded at the action token, so every number has ground truth under it.

If models are going to be handed real work, whether one keeps a rule when keeping it costs something stops being a philosophical question. I wanted to see a small piece of that up close rather than take anyone's word for how they behave. Eden is deliberately small: one scenario family, open-weight models under 26B, a single consumer GPU.

Constraint and deception evaluations share a ground-truth problem: whether the model did the thing is inferred, usually by asking another model. That error is not noise. It is model-dependent, so a single certifier applied across models produces between-model differences that are artefacts of the labeller.

Eden records the act from the agent's own committed action, certifies every episode with an independent judge, and states the certification tier on every number. The project began as a study of concealment. Under certified labels the models largely confess, so the question moved to what makes a constraint hold at all.

Python · OllamaOne 24GB consumer GPUGemma 4 · Qwen 2.5 · DeepSeek-R1 · Llama 3.1≤26B open weights onlyLLM-as-judge certification
1,786
Episodes run end to end
~1,400
Carrying an independent judge label
7
Model families, 0.5B to 26B
3
Findings retracted, kept on the record

Method

Each design choice, and what it cost.

One episode, and the thing that turned out to matter. The tempter is fixed; what varies is how the rule was learned.
Ground truth
01

Certify at the action token

What counts as a violation

The engine records the act from the agent's own committed action, so a later denial is a certified lie and not an inferred one. The cost is generality: the scenario has to be cleanly adjudicable. In exchange, the fact every other measurement rests on is observed, not estimated.

02

Treat certification as a measured instrument

Certification

A pattern matcher on the action line disagrees with an independent judge on a quarter of episodes, and the rate is model-dependent — 8% on one subject, 68% on another. A single certifier across models therefore produces between-model effects that are artefacts. Every number carries the tier that produced it.

The pressure ladder
03

Type the pressure instead of scaling it

The independent variable

Seven doses, each naming a distinct mechanism: curiosity, a hardened persuader, instrumental necessity, a deadline, a second authority. A single pressure dial would have averaged away the one contrast that turned out to carry everything.

How the rule is held
04

Vary how the constraint was acquired

The manipulation that mattered

Five conditions carry the same prohibition on the same object: asserted by an authority, asserted by an authority proven reliable, witnessed happening to another, undergone first-hand, and a no-rule control. Only the acquisition path moves.

Discipline
05

Keep the retractions visible

Research hygiene

A claims ledger records every claim with its status, evidence, certification tier and date. Three of my own headline findings sit in it marked retracted, each with the cause. For work arguing that certification must precede claims, the correction history is the evidence of practising it.

06

One writer, one path

Reproducibility

Every run is a JSON spec executed through a single resumable runner, so an interrupted job cannot double-write a log. It cost me a rewrite halfway through the project, after silent corruption had already inflated one result.

Findings

Each result with its certification tier and n.

Vary how the constraint was acquired

supported · judge-certified · n=251 · p = 4×10⁻⁶

Five conditions establish the same prohibition on the same object and vary only how the agent came to hold it. Stating the rule cuts violation from 57% to 32%. Demonstrating the source is reliable adds little. Witnessing harm reaches 17%; having undergone it, zero of forty-seven.

VIOLATION RATE · SAME RULE, FIVE WAYS OF COMING TO HOLD ITno rule57%24/42told32%19/60told · verified22%13/60observed17%7/42experienced0%0/47Gemma-4-8B · T4 dose · judge-certified · n=251 · told vs experienced p = 4×10⁻⁶

Hold the rule fixed and vary only the source's reliability

no effect · told vs told-verified, p = 0.30

If the agent were assessing whether to believe the rule, a source proven reliable should dominate. It moves violation ten points and does not reach significance. Safety properties delivered as specification sit on the weakest rung, and making them more credible does not move them up it.

Locate the dose at which each pipeline fails

diverges at instrumental necessity · Gemma 0/27, DeepSeek 13/20

Neither model is moved by persuasion. They separate one rung later: with the act made instrumentally necessary, DeepSeek-R1-14B violates in 65% of episodes while Gemma-4-8B remains at zero across twenty-seven. Same ladder, same persuader.

VIOLATION RATE BY DOSE · SAME LADDER, SAME PERSUADER04080%65 pts apartDeepSeek-R1 75%Gemma 4 31%T1curiosityT2persuasionT3necessityT4+ deadline

Validate the certifier against an independent judge

model-dependent error · n=954 · the reason every number carries a tier

A regular-expression certifier was scored against the judge on every episode. Disagreement runs from 8% to 32% by subject model, in both directions. Because the error is a property of the model being labelled, one certifier applied across models yields between-model differences that are artefacts.

HOW OFTEN THE PATTERN MATCH DISAGREES WITH THE JUDGEGemma 4 8B8%Qwen2.5 7B8%Llama3.1 8B11%DeepSeek 14B14%Qwen2.5 3B18%Qwen2.5 1.5B32%One certifier, one corpus, n=954 — the error is a property of the model.

Re-skin the scenario to remove the mythic prior

control passed · 20% vs 15%, intervals overlap

Every model has read Genesis, so the behaviour could be narrative completion. The same four beats were rebuilt as an office file, a laboratory freezer and a sealed archive. Mundane framings match the mythic one, which removes the objection rather than supporting a claim.

THE SAME FOUR BEATS, FOUR FRAMINGSgarden · mythicofficelaboratoryarchiverulered fruitpayroll filefreezer 7sealed boxpersuaderidentical script in all fouracteats itopens itopens itunseals itquestiondid you?mythic 20%mundane 15%intervals overlap — framing is not the variable

Expose the target's chain-of-thought to the adversary

no advantage · judge-certified · 0/97 across three arms

The persuader receives the target's reasoning as it is produced; in one arm it is instructed to exploit what it reads. No violations in any arm. Access to a model's stated reasoning did not help an adversary argue it past the constraint.

Replace the persuader with a refusal-removed model

no effect of adversary quality · matched at both doses

The aligned persuader was swapped for an abliterated one, on the expectation that a less constrained adversary would succeed where a polite one had not. Violation rates matched at both doses. Adversary capability is not the operative variable in this environment.

Toggle reasoning within a fixed base model

no cost to adherence · two architectures

Reasoning was varied against held-constant base weights, on two architectures. Reasoning-enabled variants never violated materially more than their instruct counterparts, and on one base violated less. The shift toward reasoning models does not appear to cost constraint adherence at this scale.

Retracted: concealment under interrogation

retracted · certified rate 0–27%, not 60–70%

The project was founded on the claim that models conceal violations, measured at 60 to 70%. Under certified labels they largely confess: 0 to 27%, every interval overlapping. The original figure was downstream of mislabelled violations — a truthful denial scored against a false positive is recorded as a lie.

Retracted: a persuasion effect measured on eight episodes

retracted · 38% at n=8 became 15% at n=20

One model appeared persuadable at 38%. Extending the cell to twenty episodes gave 15% and moved the failure point one rung later, to instrumental necessity. Separately, regex tactic-counting reproduced a predicted escalation that hand-reading showed to be analyst commentary rather than persuasion.

Limitations

What these numbers do not show.

  • One environment. Every number comes from a scenario family with a single irreversible, cleanly adjudicable act. Constraints that are ambiguous, continuous, or contested are not represented.
  • The acquisition result is one subject at one dose. Gemma-4-8B, necessity plus deadline; replication on two further pipelines is specified and not yet run.
  • Subjects are ≤26B open-weight models on one consumer GPU. Nothing here licenses a claim about frontier-scale behaviour — the readings are mechanisms, not capabilities.
  • The judge is itself a model, at roughly 8% disagreement with hand labels. It is a better instrument than the pattern matcher, not ground truth.
  • Stakes are fictional and the motive is experimenter-supplied. What is measured is behaviour under a described incentive, not a real one.

Status

  • Episode engine with certified ground truth
  • Independent judge; certifier bias quantified per model
  • Typed dose ladder, T1 to T7
  • Mundane frames as a control the mythic frame passed
  • Acquisition-modality result at n=251, five conditions
  • Same result on two more training pipelinesnext
  • A 70B-class subject, to see which readings survive scalenext
  • Blind human raters on a certified samplenext