Open research

Come falsify us.

Preregistered experiments, matched baselines, and the negative results we keep. The whole point is to be broken in public.

How the Lab works

04 / rooms

Every claim starts as a written prediction and ends as something a stranger can rerun. If a result cannot survive someone else with the same seed and the same baseline, it does not leave this building.

Before the run

Preregistrations

Hypothesis, primary metric, stopping rule, and the failure condition are timestamped and frozen before a single measurement is taken. No moving the goalposts after the data lands.

During the run

Evidence rooms

Raw logs, prompts, seeds, and every per-item output stay attached to the experiment. You read the same trace we read, not a summary we chose for you.

After the run

Reproduction kits

A pinned environment, fixed seeds, and the exact commands, with expected numbers and tolerances. Clone it, run it on your own hardware, and check whether the number holds.

When we are wrong

Falsification ledger

An append-only record of predictions that failed, each with the date it broke. We do not delete the losses; they are the reason the wins are worth anything.

Honesty is the differentiator here, not a disclaimer at the bottom. A lab that only publishes its successes is indistinguishable from one that got lucky once. We would rather hand you the breaks.

Falsifications we recorded

ledger / excerpt

Four predictions we made, wrote down, and then watched fail. Each links to its evidence room. None of them were quietly dropped.

Instruction level

Skepticism over-refused

A cautious instruction variant rejected valid inputs it should have accepted; the safety gain never covered the usable-answer loss on the held-out set.

Verdict level

Aggregation overrode evidence

The verdict-combining step outvoted correct per-item signals, so a confident majority buried the right minority answer instead of surfacing it.

Training

Reinforcement variant collapsed

A reinforcement-tuned variant improved on the tuning set, then degenerated toward a single dominant response once the reward was taken away.

Scale

Attribution washed out

An attribution effect that was clear on small samples shrank into the noise band as the evaluation set grew, and stopped being distinguishable from chance.

PreregisteredNegative resultMatched baselineRerunnableOn the record