Command Info

Name: linksite:jp-bench-watch

Description: Periodically evaluate JP bench orchestration progress until landing is complete

Status: completed

Start Time: 2026-07-10 09:17:40

End Time: 2026-07-10 09:17:41

Total Time: 1 second

Dispatched Jobs Count: 0

Successful Jobs Count: 0

Failed Jobs Count: 0

Output

09:17:41 [2026-07-10T09:17:40+00:00] JP bench 8/51 scored (15.7%) · missing 43
orchestrator: not detected
next cell: jp-1 × moonshotai/kimi-k2.6
log: ❌ Cell failed: jp-2 × moonshotai/kimi-k2.6 — Generation timed out after 30 minutes
gaps (57):
· [medium] jp-1: missing 2 challenger(s): moonshotai/kimi-k2.6, fugu-ultra
· [info] jp-1: A/B comparison ready (4 challengers + baseline)
· [medium] jp-2: missing 2 challenger(s): moonshotai/kimi-k2.6, fugu-ultra
· [info] jp-2: A/B comparison ready (3 challengers + baseline)
· [medium] jp-3: missing 4 challenger(s): moonshotai/kimi-k2.6, fugu-ultra, anthropic/claude-sonnet-5…
· [high] jp-4: missing 5 challenger(s): deepseek/deepseek-v3.2, moonshotai/kimi-k2.6, fugu-ultra…
· … +51 more (see gaps report file)
improvements (5):
· [high] 113 linksite eval/finalize failed_jobs in 24h — inspect failed_jobs table
· [medium] evaluator.all_axes_failed ×1 in recent logs — All eval axes failed — check model routing / API keys
· [info] finalize_generation.skip_auto_eval_bench ×1 in recent logs — Bench double-eval prevented (expected on is_bench_run)
· [info] clobbering_inflight ×7 in recent logs — Eval dispatch clobbered an in-flight run — run docs should isolate, verify aggregator
· [info] Opus limited to jp-1 per HOWTO (~$8/post) — verify cost on landing card