field note 00 · 0xgrizz
Inference is the qualifier.
Training is the main event.
01 · the middle
The part between prompt and product.
AI marketing has trained us to look at the edges. On one edge, a prompt goes in. On the other, an answer appears. Latency is easy to time, tokens are easy to count, and a fast demo is easy to understand. The harder object sits in the middle: deciding which experience should change the model, proving that experience is real, applying the gradient, and showing that the next checkpoint is better outside the training batch.
That middle is Reliquary. Miners do not merely render text. They compete to locate prompts at the current policy’s learning frontier, where a group contains enough success and failure contrast to carry a useful training signal. The validator recomputes rewards, verifies proof-carrying inference, seals the selected batch, and only then permits the optimizer to move.
Speed can win a slot. It cannot win the argument. The public checkpoint has to move.
This distinction changes the category. If inference latency is the whole contest, the event ends before learning begins. If the outcome is held-out model improvement under a fixed evidence contract, inference becomes what it should be: qualifying.
02 · the starting image
128 trajectories, deliberately held.
Season Zero begins with a real production receipt, not a render. At sealed window 25,259, the public archive recorded a balanced batch of 16 selected groups—8 Math and 8 Code—for 128 total trajectories. The batch was ready and not quarantined.
The same receipt says training_attempted=false, trained=false, and blocked_reason=training_checkpoint_ceiling. The public validator state reported checkpoint 53 at revision 320a1896…. That is not a failure to hide. It is the proof that a training gate can remain closed when the operating contract says stop.
frozen receipt · fetched 2026-07-24T03:18:49Z · descriptive production state
03 · the circuit
Seven gates, all inspectable.
A useful competition cannot have a hidden shortcut between fast output and declared victory. FLEE7 maps the product people can watch onto the protocol operators already run. Each gate names the artifact required to pass.
- 01findSearch the frontierprompt + checkpoint
- 02runGenerate the fieldeight rollouts
- 03proveVerify the workGRAIL receipt
- 04separateRank the signalreward geometry
- 05qualifySelect the batchsealed window
- 06moveTrain the modelGRPO cycle
- 07publishProve the gaincheckpoint + eval
Search asks whether a miner can find the narrow part of the prompt distribution that teaches. Generate fixes the group geometry. Verify binds the work to the declared model. Rank separates usable contrast from unanimous noise. Select freezes the batch. Train moves the weights. Prove publishes the checkpoint, evaluation, and chain of custody. A fleet that misses any gate does not carry a result forward.
04 · what a championship must control
Fairness is a system, not a vibe.
The visual language can be spectacular; the comparison contract should be boring. Mature training benchmarks converge on the same principle: decide what is fixed, decide what can change, declare the quality target, repeat the run, and publish enough evidence to replay the result.
MLCommons’ AlgoPerf competition is a useful precedent. It fixes workloads and hardware so teams compete on training algorithms, evaluates time to target across multiple workloads, separates tuning rule sets, repeats runs, and publishes logs and submissions. Reliquary is not claiming compatibility with or endorsement from MLCommons. We are borrowing the discipline: make the source of advantage legible.
| control | season zero draft | why |
|---|---|---|
| start | same checkpoint + revision | No fleet begins closer to the target. |
| field | fixed tasks + hidden holdout | Training data cannot become the answer key. |
| budget | normalized compute envelope | More hardware is measured, not disguised as method. |
| evaluation | precommitted metric + seeds | The finish line cannot move after results appear. |
| repeats | matched repetitions | One lucky run cannot become a title. |
| result | provisional until replay | The public record outranks the live show. |
The primary outcome should be held-out model improvement per normalized compute budget, conditional on proof validity. Useful-group rate and pace can explain how the fleet got there. They should not overrule a checkpoint that failed to improve.
05 · current evidence
One result is a starting line.
In a historical controlled 300-step comparison using Qwen3-4B-Instruct, the Reliquary arm measured pass@1 0.470 → 0.610 versus vanilla GRPO: +14 percentage points at equal step count. It is a strong reason to run the next experiment. It is not a production guarantee, a universal efficiency multiplier, or proof that every future checkpoint improves.
The current production system is also not the historical experiment. Today’s public contract uses Qwen3.5-2B at ReliquaryForge/qwen3.5-2b-reliquary-v3, groups of 8, and 100-second windows. It publishes every four successful balanced optimizer steps, with earlier publication on behavior-policy drift. Season Zero has to preserve that separation between what happened once, what is live now, and what the next experiment proposes.
06 · novelty
Competition can be the research instrument.
Difficulty answers one question: did a rollout group contain learning contrast? It does not answer another: did two successful trajectories reveal distinct strategies, or the same behavior twice? Reliquary’s novelty work starts by refusing to collapse those questions.
Today, behavioral novelty has zero influence on ranking, rewards, selection, training, and emissions. The proposed path is shadow-first: define observable descriptors, preserve a parallel archive, audit identity leakage and gaming pressure, replay against frozen windows, and only then preregister an experiment. A competition gives that work pressure, characters, and repetition. The protocol gives it receipts.
A yearly season is valuable only if every edition leaves better rules, better evidence, and a checkpoint future researchers can challenge.inspect the novelty ledger →
07 · Montréal
A deadline, not an affiliation.
Exploit Summit 2026 is scheduled for 28–29 September at New City Gas in Montréal. Its public program is built around debates, workshops, demos, game-theoretic competition, and research. That makes Montréal a meaningful deadline for a Bittensor training protocol to show its work.
Reliquary does not currently have a publicly announced speaker slot, sponsorship, or official activation. Until an agreement exists, Season Zero is an independent Reliquary program built on the road to Montréal. The event owns its name and program; Reliquary owns the competition system, evidence, and result.
the compact
No claim without a receipt.
No title before replay.
Every season leaves a checkpoint.
— 0xgrizz · founding maintainer
enter season zero →