CORTEXA
← Browse
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26Cited by 0

STG Reviewer Validation Benchmark v1.0: Matched MuJoCo Baselines, Memory/Torsion Ablations, and Negative Specificity Results

Marcel Krüger

This record provides the complete code, frozen configurations, calibration and test seeds, raw episode- and step-level logs, processed tables, statistical outputs, figures, environment manifests, checksums, and a self-contained Google Colab workflow for a reviewer-requested comparative validation of the Spiral-Time Governor (STG) as an execution-level safety filter for LLM-controlled legged robots. The benchmark compares four policies under matched conditions: (i) an always-execute condition, (ii) a simple instantaneous deterministic threshold filter without historical memory or torsion, (iii) a no-memory/no-torsion STG ablation, and (iv) the frozen full STG v1 implementation. The methods were evaluated in four MuJoCo scenarios: the quadruped escape-bowl task, a rough-heightfield task, a repeated stair/step-up task, and a physically defined foothold-ascent proxy. Separate calibration seeds were used for threshold selection, followed by 50 frozen and matched test seeds per policy and scenario, yielding 800 final test episodes. No parameters were retuned after inspection of the final test results. All three governed methods prevented every action classified as unsafe by the shared certified norm monitor and produced an intervention rate of approximately 10.13%, whereas the always-execute condition admitted a mean of 30.4 unsafe proposals per episode. However, the simple threshold filter, the no-memory/no-torsion ablation, and full STG v1 produced identical execution decisions and identical physical outcome metrics under the implemented benchmark. Full STG v1 therefore showed no measurable safety or task-performance advantage over the simpler deterministic filter and incurred greater decision latency. No evaluated policy achieved task success in the four scenarios because the benchmark used an open-loop mock proposal generator rather than a task-capable locomotion controller. Consequently, this record does not establish climbing competence, autonomous locomotion capability, hardware validation, or superiority of STG-specific memory and torsion terms. The raw hallucination rate remained unchanged across policies by design because STG v1 filters execution rather than LLM claim generation. No external stochastic LLM experiment is included in this record. The negative and null findings are preserved without post hoc adjustment. This package is released as a reproducibility dataset and comparative safety-filter benchmark, not as evidence that full STG v1 outperforms a matched instantaneous threshold baseline.

View free PDFSource page

Related papers

openalexZenodo (CERN European Organization for Nuclear Research)2026-07-24

Artificial Intelligence In Pharmaceutical Process Validation: A Review

Taufik Mulla*, Siddheshwar Sonavane, Ayush Tambe, Megha Hange, P. N. Sable

For decades, pharmaceutical process validation has rested on a relatively narrow set of habits: a fixed qualification protocol, a small handful of conformance batches, and a periodic review of trends to argue that manufacturing is operating consistently. That posture is being cha…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

Horizon-Dependent Value of Geodetic Context in Operational Earthquake Forecasting: A Leakage-Free, Multi-Region Study with Pre-Registered Negative Results

Felipe Santibañez-Leal

Operational earthquake forecasting (OEF) issues calibrated conditional probabilities of future seismicity; the Epidemic-Type Aftershock Sequence (ETAS) model is its de-facto benchmark, and no machine-learning temporal point process has been shown to beat a well-fit ETAS prospecti…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-26

Flexbrick photovoltaic ceramic textile façade: a coupled EnergyPlus–Radiance–PV dataset and surrogate-optimisation pipeline across European climates

Mirco Riganti

This deposit accompanies the study "Machine-Learning Surrogates and Simulation-Validated Multi-Objective Optimisation of a BIPV Ceramic Textile Façade: Energy, Daylight and Thermal Comfort Trade-Offs Across European Climates". It quantifies how the binary weave pattern of a 28 x…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-25

D013 Prime Elementology — Prime Spectral Descriptor Atlas for Synthetic RF

Thành Trung Phan

D013 Prime Elementology — Prime Spectral Descriptor Atlas for Synthetic RF Dataset ID: D013Version: 2.0Dataset Type: Synthetic Research DatasetAuthor: Phan Thành TrungORCID: 0009-0000-7520-6781DOI: 10.5281/zenodo.21569013 1. Overview D013 Prime Elementology — Prime Spectral Descr…

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)

Data and Code to reproduce results in paper "A Systematic Literature Review on Graph-Based Models in Credit Risk Assessment"

Lennart John Baals, Yiting Liu, Joerg Osterrieder, Branka Hadji Misheva

Data and Code to reproduce results in paper "A Systematic Literature Review on Graph-Based Models in Credit Risk Assessment" This repository contains the necessary codes to reproduce results in the paper: Baals, L. J., Liu, Y., Osterrieder, J., & Hadji-Misheva, B. (2025). A Syste…

Also available via: European Organization for Nuclear Research

View free PDFSource page
openalexZenodo (CERN European Organization for Nuclear Research)2026-07-23

PREreview of "Demanding peer review is associated with higher impact in published science"

Alan Colín‐Arce, Kylie Yui Dan

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/21515890. This work seeks to explain if comprehensive peer review is associated with higher citation impact. The study utilizes LLMs to parse…

View free PDFSource page