Systematically exploring what the lab could technically do. Method (Boyden / Wang / Khera / Marblestone): recursively split a solution set into MECE subsets by physical / measurable properties. Avoid retroactive splits (organizing around existing fields), catch-all "other" buckets, vague words, and category errors (mixed split dimensions). Redo trees as tech evolves.
Problem tiled: ways to reduce a frontier model's misalignment (make it verifiably safer) that a new lab could own. Evals/measurement are a different tree — mixing them is a category error.
Split by field-names: pretraining · RLHF/post-training · interpretability · evals · oversight · governance. Pitfalls: retroactive (only re-surfaces known ideas), category error (evals = measurement, governance = distribution — different dimensions), catch-all ("oversight/other" dead-ends). Discard.
MECE + exhaustive: an intervention edits the weights, or edits runtime activations, or leaves the model untouched.
A2 and B2 share one dependency: both need a validated internal readout (a probe). So re-tile on the orthogonal physical axes — must you read internals? × do you change the model or the environment? (a Zwicky box):
| Change the model | Change the environment | |
|---|---|---|
| Black-box behaviour only |
RLHF · DPO · SFT A1 crowded — everyone |
prompts · filters · scaffolds · access C1–C3 crowded — the startup swarm |
| White-box read internals |
probe → train-away A2 · structural prune A3 ← the empty tile · Neolab's home |
probe-gated runtime block B2 · interp audit |
White-box × change-the-model, at scale, on open weights. Use a validated probe to drive weight edits (A2 train-away + A3 structural prune). Almost nobody does this on the newest open-weight models — it needs a datacenter + the will. Measurement already exists (FAR's probes); the optimization half is open. That gap is the lab.
Press the empty technical tile (A2+A3) and the governance tile (C3, per-country access) together. Each alone has players. Nobody presses both. The combination is the underexplored idea — a lab that produces realigned open weights and coordinates their access via a per-country institution.
Pitfalls avoided: no catch-all bucket; split by physical loci (weights/activations/environment) not field-names; measurement kept in a separate tree; one consistent dimension per split.
Tile sources of legitimate coordinating power by the property who is bound, and by what:
Read-off: Neolab's legitimacy = participation × technical-indispensability (cheap + compounding), optionally backed by ownership of the pipeline/compute. Avoid force as the primary. This is the co-authorable moat FAR won't build.