STG Reviewer Validation Benchmark v1.0: Matched MuJoCo Baselines, Memory/Torsion Ablations, and Negative Specificity Results
This record provides the complete code, frozen configurations, calibration and test seeds, raw episode- and step-level logs, processed tables, statistical outputs, figures, environment manifests, checksums, and a self-contained Google Colab workflow for a reviewer-requested comparative validation of the Spiral-Time Governor (STG) as an execution-level safety filter for LLM-controlled legged robots. The benchmark compares four policies under matched conditions: (i) an always-execute condition, (ii) a simple instantaneous deterministic threshold filter without historical memory or torsion, (iii) a no-memory/no-torsion STG ablation, and (iv) the frozen full STG v1 implementation. The methods were evaluated in four MuJoCo scenarios: the quadruped escape-bowl task, a rough-heightfield task, a repeated stair/step-up task, and a physically defined foothold-ascent proxy. Separate calibration seeds were used for threshold selection, followed by 50 frozen and matched test seeds per policy and scenario, yielding 800 final test episodes. No parameters were retuned after inspection of the final test results. All three governed methods prevented every action classified as unsafe by the shared certified norm monitor and produced an intervention rate of approximately 10.13%, whereas the always-execute condition admitted a mean of 30.4 unsafe proposals per episode. However, the simple threshold filter, the no-memory/no-torsion ablation, and full STG v1 produced identical execution decisions and identical physical outcome metrics under the implemented benchmark. Full STG v1 therefore showed no measurable safety or task-performance advantage over the simpler deterministic filter and incurred greater decision latency. No evaluated policy achieved task success in the four scenarios because the benchmark used an open-loop mock proposal generator rather than a task-capable locomotion controller. Consequently, this record does not establish climbing competence, autonomous locomotion capability, hardware validation, or superiority of STG-specific memory and torsion terms. The raw hallucination rate remained unchanged across policies by design because STG v1 filters execution rather than LLM claim generation. No external stochastic LLM experiment is included in this record. The negative and null findings are preserved without post hoc adjustment. This package is released as a reproducibility dataset and comparative safety-filter benchmark, not as evidence that full STG v1 outperforms a matched instantaneous threshold baseline.