skip to content
syncing conviction
conviction watch · syncing finalized state
syncing
→
Conviction state: syncing conviction
Reliquary
dashboard
conviction
research
roadmap
docs
source
x
menu
Forge · live training · team only · Reliquary
forge · live training
5H47sFL6-recovery-checkpoint-182
·
step 28,477
live
▸ open in wandb
status
running
runtime
4d 05h
last seen
—
loss
-1.12e-5
kl
7.89e-3
grad_norm
2.938
reward μ
0.4922
steps / h
280
target met
lr
1.00e-6
gpu util
0.0%
gpu mem
0.0%
ai advisor
reading the last 160 points…
model quality
computing quality signals…
validator rejections
tailing validator logs over ssh…
PPO loss
primary objective
KL divergence
budget kl_beta = 0.01
grad_norm
clip @ 1
learning rate
cosine schedule
rewards
mean ± std
degenerate-group ratio
zero-variance reward groups
valid rollout ratio
GRAIL accepted / submitted
model improvement · checkpoint evals
held-out pass@1 · math + code
no data
No checkpoint evals yet
Held-out benchmarks appear here as soon as the eval pipeline publishes its first checkpoint for the current base. Expected shortly after a base reset.
accepted rollouts · rolling window
verified into training · last 3 windows
accepted rollouts
verified miner rollouts stream once the validator seals a window
gpu util
0.0%
gpu mem
0.0%
sm occupancy
0.0%
gpu temp
53°C
power
90 W
run config
22 keys · hide
b_batch
8
grad_clip_norm
1
grad_norm_skip_threshold
50
kl_base_model
Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
kl_beta
0.01
kl_reference_mode
fixed
kl_reference_repo_id
Qwen/Qwen3.5-4B
kl_reference_revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
kl_reference_storage_bytes
9078531392
learning_rate
0.000001
lr_cosine_max_windows
10000
lr_warmup_windows
10
m_rollouts_per_prompt
8
pi_old_source
verify_model
ppo_clip_epsilon
0.2
ppo_ratio_outside_clip_skip_threshold
0.1
reliquary_version
0.1.0
shape_len_frac
0.5
shape_penalty
0
train_until_checkpoint_n
0
wandb_training_version
recovery-checkpoint-182
window_length
5