skip to content
DECENTRALISED POST-TRAININGBITTENSOR SUBNET 81
RELIQUARY
PLATFORM / MODEL EVALUATIONCOMING SOON

Know what changed in your model.

Test models on the tasks you care about. The planned reports keep scores, example answers and the test setup together, so you can trace a result back to the evidence.

ILLUSTRATIVE EXAMPLEPlanned evaluation package
Input
Pinned checkpoint
Suite
Task set + grading contract
Method
Recorded sampling configuration
Output
Task results + comparison report

An illustrative workflow, not an evaluation in progress or a measured score.

01 / WHAT YOU GET

See more than a score.

Review individual answers alongside the aggregate result, with the checkpoint, suite revision and sampling setup recorded. Published Lab benchmarks remain historical research evidence, not results from this upcoming service.

  • Task-level results
  • Recorded test setup
  • Comparison report
02 / HOW IT WORKS

From a checkpoint to an inspectable report.

  1. 01

    Pin the checkpoint

    Choose the model revision to assess and a separate baseline when you want a comparison.

  2. 02

    Define the evaluation

    Choose the tasks, grading method and sampling. Use the same conditions for the models you compare.

  3. 03

    Review the results

    Inspect example answers and grader evidence beside the overall results. Keep the test setup and limitations visible.

From checkpoint to reportWorkflow concept
  1. 01
    Select a checkpointyour-model / checkpoint
  2. 02
    Choose a task suiteCode · Math · Logic
  3. 03
    Inspect the reportTask results, grader details and examples
report / exampleNot measured
Task
Deduplicate a list
Grader
Python task tests
Evidence
Model response + test output
Inspect example task
input:  [3, 1, 3]
expect: [3, 1]
Illustrative workflow · custom evaluation is coming soon. No evaluation was executed.
03 / SCOPE & PRICING

Built around your workload.

No public evaluation rate is announced. Tell us about your model, task suite and comparison requirements to help define the scope.

Read the documentation
Model identity
Pinned checkpoint revisions, with a distinct baseline when needed.
Task suite
A versioned task set and a documented grading contract.
Evaluation method
Sampling, sample count, metrics and comparison conditions.
04 / AVAILABILITY

Custom evaluation is coming soon.

Early-access requests do not start an evaluation or grant execution permission. Published model results can already be inspected in the Lab.

Request early access

Tell us about your workload. Requesting access does not start a job, reserve compute or grant execution permission.

Inspect the published Lab results