skip to content
DISTRIBUTED POST-TRAININGBITTENSOR SUBNET 81
RELIQUARY
Menu
RELIQUARY LAB / BITTENSOR SUBNET 81

We train models
on a decentralized network.

Reliquary is also a research lab. We run our own training to push decentralized post-training forward, starting with reinforcement learning, and to build specialized models.

training rollouts from independent miners
6.6M
trainer-generated training rollouts
0
points on MATH-500 (near-domain)
+35.4
points on AMC23
+31.9

Reliquary-4B, from Qwen3-4B-Base. Sampled pass@1, single run, not replicated. Technical report §6. These figures describe the published run, not current live telemetry.

01 / WHAT WE DO

A lab that trains in the open.

01

Research

We study how to post-train models on a network nobody owns, and publish what we learn: the mechanism, the runs, the failures.

02

Training

We publish training records and checkpoints. Public evidence makes the network’s work inspectable.

03

Specialized models

We train models for specific tasks and release the weights, starting with reasoning in math and code.

02 / THE SYSTEM

Reinforcement learning, run by a network of miners.

The reported GRPO run is on-policy: the model learns from attempts generated by the checkpoint being trained. Miners supply them, validators check them, and selected work feeds the trainer. The next checkpoint returns to the network.

  1. 01

    Miners

    Generate answers on their own GPUs, from the current checkpoint.

  2. 02

    Verification

    The reported design checks submitted token probabilities against the published checkpoint, allowing for numerical differences across supported hardware.

  3. 03

    Reliquary

    Trains the model, publishes the checkpoint, and the loop repeats.

New checkpoint, back to the miners

Miners supply the rollout generation. Verification and training remain separate.

Find the prompts that teach the model.

In the published run, GRPO compares 16 attempts at the same prompt. If all outcomes receive the same reward, their relative rewards carry no learning signal. Mixed outcomes can carry a signal, but are not a payment guarantee: verification and selection still apply.

All fail0/16 pass · No relative-reward signal
All pass16/16 pass · No relative-reward signal
Mixed7/16 pass · Learning signal
  • Miners choose the prompts.
  • Selected, verified mixed-result groups are eligible for rewards.
  • Generation has a cost. Supplying a mixed group does not guarantee selection or payment.

Useful code groups became less common in this run.

Share of code prompts that give a mixed group under uniform sampling, measured on a fixed, index-held-out suite of 500 prompts with 16 rollouts each. The suite is not content-disjoint. This result describes one run, not every training trajectory.

Share of code prompts that give a useful groupIn the reported run, uniform sampling gives mixed groups for 88.2% of a fixed index-held-out code suite at the base model, 42.8% at update 8,800 and 41.4% at update 12,842. This is not a live chart or a content-disjoint evaluation.0%25%50%75%100%04,0008,00012,000Optimizer updates88.2% · base model41.4% · end of run
Scroll to inspect the chart ↔Three reported checkpoints · inspect a point or open the data table.
Show reported data
Code mixed-group share under uniform sampling; technical report §3.3
CheckpointUseful (mixed) groups
Base model88.2%
Update 8,80042.8%
Update 12,84241.4%

This chart records the measured mixed-group share for a fixed, index-held-out suite. It does not measure production savings.Technical report §§3.2–3.3 ↗

03 / OUR MODELS

From our own model to Bittensor's.

Reliquary-4B is our reported release. Teutonic-I is an ecosystem research programme. Specialized models are a planned direction. Research plans are separate from live network activity and available workspace execution.

  1. ReleasedOur model

    Reliquary-4B

    Reported on the network
    • SFT data (not run on the network)
    • RL (run on the network)
    • Evaluation (not run on the network)

    Only the RL ran on the subnet.

    Base
    From Qwen3-4B-Base, RL only
    Focus
    Reasoning in math and code
    Result
    +35.4 pts on MATH-500 (near-domain), +31.9 on AMC23
    Weights ↗
  2. Research programmeA Bittensor model

    Teutonic-I

    Planned pipeline
    • SFT data (planned, not a verified execution)
    • RL (planned, not a verified execution)
    • Evaluation (planned, not a verified execution)

    A research programme for the full pipeline. Stages are planned, not a live execution status.

    Base
    dendriteholdings/Teutonic-I
    Focus
    Math, code, instruction following and agentic tool use
    Evaluation
    Proposed held-out evaluation: IFBench, BFCL v3, AIME, LiveCodeBench, GPQA and MMLU-Pro.
    Inspect network activity ↗
  3. NextDomain by domain

    Specialized models

    Planned pipeline
    • SFT data (planned, not a verified execution)
    • RL (planned, not a verified execution)
    • Evaluation (planned, not a verified execution)

    A planned research direction, not a completed training run or available service.

    Base
    Models chosen domain by domain
    Focus
    Domain-specific post-training
    Evaluation
    Planned releases with weights and evaluations.
    Our research ↗

Reliquary-4B, in detail.

Qwen3-4B-Base to Reliquary-4B, with RL only and no SFT. 6.6M rollouts reported in the run, with no trainer-side rollout generation.

MATH-500 · near-domain+35.4

37.2% → 72.6%

HumanEval+ · code+20.9

52.1% → 72.9%

AMC23 · not near-domain+31.9

15.6% → 47.5%

Sampled pass@1, 8 samples per problem, T = 1.0, same prompts for both models. MATH-500 is near-domain; AMC23 is not. Single run, not replicated. Deltas are computed before rounding. No matched centralized control. Technical report §6.2.

More detail
  • Held-out, exact-content-disjoint subsets: +31.1 math (95% CI 28.5 to 33.8, 282 problems), +20.1 code (18.6 to 21.6, 450 problems). §6.2
  • pass@8: AMC23 55.0 to 80.0, AIME 2024 10.0 to 23.3. §6.2
  • Greedy pass@1: MATH-500 52.8 to 77.8, HumanEval+ 78.7 to 76.8. §6.4
  • 94% of the math gain came in the first ~800 updates; math plateaued from ~update 4,600, a reference-quality issue. §3.4, §6.3
  • About 16,000 submissions were rejected for token mismatch during the reported run. §4.2
04 / SCALABLE DESIGN

Scalable design for RL.

Independent miners supply generation; verification and training are separate. The reported run demonstrates this architecture end to end. Capacity, comparative costs and future workload compatibility still require measurement and qualification.

Grows with the network
  1. 01

    Generation scales out.

    Independent miners supply rollouts on their own GPUs. The reported run required no trainer-side rollout generation.

  2. 02

    Heterogeneous hardware.

    The reported design verifies token probabilities with a tolerance rather than bit-exact GPU replay. Supported hardware and runtime requirements still apply.

  3. 03

    Selection happens at the edge.

    Miners choose prompts. Selected groups must meet verification and eligibility rules before they can receive rewards.

Stays light on our side
  1. 04

    Verification, not regeneration.

    In the reported design, the validator scores submitted tokens against the checkpoint instead of regenerating completions.

  2. 05

    The trainer only trains.

    The reported trainer consumes selected, verified rollouts and publishes checkpoints. Generation and verification remain separate stages.

  3. 06

    A reusable pipeline.

    SFT data, RL and evaluation are distinct tasks. New models still require compatible environments, execution profiles and qualification.

“one of the most compelling structures i have seen for a scalable post training mechanism on Bittensor”

const · @const_reborn · Public commentary
Read the commentary ↗
05 / RESEARCH

What we publish.

  1. technical report · september 2026

    A Market Mechanism for Rollout Selection in Group-Relative RL

  2. field note · the 2B run

    Reliquary 2B: training done

  3. field note · behavioral diversity

    Eight rollouts are not eight votes

All research ↗