Research
We study how to post-train models on a network nobody owns, and publish what we learn: the mechanism, the runs, the failures.
Reliquary is also a research lab. We run our own training to push decentralized post-training forward, starting with reinforcement learning, and to build specialized models.
Reliquary-4B, from Qwen3-4B-Base. Sampled pass@1, single run, not replicated. Technical report §6. These figures describe the published run, not current live telemetry.
We study how to post-train models on a network nobody owns, and publish what we learn: the mechanism, the runs, the failures.
We publish training records and checkpoints. Public evidence makes the network’s work inspectable.
We train models for specific tasks and release the weights, starting with reasoning in math and code.
The reported GRPO run is on-policy: the model learns from attempts generated by the checkpoint being trained. Miners supply them, validators check them, and selected work feeds the trainer. The next checkpoint returns to the network.
Generate answers on their own GPUs, from the current checkpoint.
The reported design checks submitted token probabilities against the published checkpoint, allowing for numerical differences across supported hardware.
Trains the model, publishes the checkpoint, and the loop repeats.
New checkpoint, back to the miners
Miners supply the rollout generation. Verification and training remain separate.
In the published run, GRPO compares 16 attempts at the same prompt. If all outcomes receive the same reward, their relative rewards carry no learning signal. Mixed outcomes can carry a signal, but are not a payment guarantee: verification and selection still apply.
Share of code prompts that give a mixed group under uniform sampling, measured on a fixed, index-held-out suite of 500 prompts with 16 rollouts each. The suite is not content-disjoint. This result describes one run, not every training trajectory.
| Checkpoint | Useful (mixed) groups |
|---|---|
| Base model | 88.2% |
| Update 8,800 | 42.8% |
| Update 12,842 | 41.4% |
This chart records the measured mixed-group share for a fixed, index-held-out suite. It does not measure production savings.Technical report §§3.2–3.3 ↗
Reliquary-4B is our reported release. Teutonic-I is an ecosystem research programme. Specialized models are a planned direction. Research plans are separate from live network activity and available workspace execution.
Only the RL ran on the subnet.
A research programme for the full pipeline. Stages are planned, not a live execution status.
A planned research direction, not a completed training run or available service.
Qwen3-4B-Base to Reliquary-4B, with RL only and no SFT. 6.6M rollouts reported in the run, with no trainer-side rollout generation.
37.2% → 72.6%
52.1% → 72.9%
15.6% → 47.5%
Sampled pass@1, 8 samples per problem, T = 1.0, same prompts for both models. MATH-500 is near-domain; AMC23 is not. Single run, not replicated. Deltas are computed before rounding. No matched centralized control. Technical report §6.2.
Independent miners supply generation; verification and training are separate. The reported run demonstrates this architecture end to end. Capacity, comparative costs and future workload compatibility still require measurement and qualification.
Independent miners supply rollouts on their own GPUs. The reported run required no trainer-side rollout generation.
The reported design verifies token probabilities with a tolerance rather than bit-exact GPU replay. Supported hardware and runtime requirements still apply.
Miners choose prompts. Selected groups must meet verification and eligibility rules before they can receive rewards.
In the reported design, the validator scores submitted tokens against the checkpoint instead of regenerating completions.
The reported trainer consumes selected, verified rollouts and publishes checkpoints. Generation and verification remain separate stages.
SFT data, RL and evaluation are distinct tasks. New models still require compatible environments, execution profiles and qualification.
Read the commentary ↗“one of the most compelling structures i have seen for a scalable post training mechanism on Bittensor”