CORTEXA
← Browse
arxivcs.AI2026-07-22

Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems

Chengxiao Dai, Zhaokun Yan, Chenjun Lei, Qiao Li, Luyan Zhang

In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all. Yet existing automated remediation systems are designed to generate actions rather than to decide whether intervention is warranted, leaving safety as an afterthought enforced by manual approval. This paper makes three contributions to close this gap: (i) we reformulate safe remediation as a risk-constrained intervention decision problem and cast it as a Constrained Markov Decision Process (CMDP), in which the agent maximizes repair success subject to a bounded false remediation rate (FRR); (ii) we introduce a three-dimensional risk decomposition comprising blast radius, reversibility, and epistemic uncertainty, providing operators with an interpretable per-action safety interface; and (iii) we design a context-adaptive human-in-the-loop (HITL) gate that turns escalation from a binary failsafe into a bandwidth-aware control layer responsive to on-call load and business criticality. The full policy is learned offline from historical incident logs, enabling explicit control of the expected FRR. Experiments on the Train Ticket microservice benchmark with Chaos Mesh fault injection and an RCAEval-aligned fault taxonomy show that our framework reduces FRR by 39% while improving repair success by 2.5 points over a strong runbook baseline, and reduces on-call escalation load by 17% relative to a fixed-threshold variant.

View free PDFSource page

Related papers

arxivcs.CLcs.AI2026-07-05

Risk-Constrained Freshness-Aware Semantic Caching for Open-Web Retrieval-Augmented LLMs

Muhammad Mansoor, Tahir Ahmad, Yeo-Chan Yoon

Semantic caching reduces the latency and cost of retrieval-augmented generation (RAG) by serving cached answers to semantically similar queries, but most existing methods do not model the time-varying freshness of open-web evidence. We present FreshCache, a three-tier semantic ca…

View free PDFSource page
arxivcs.AI2026-07-15

CausalGraphX: A Counterfactual Graph Neural Network Framework for Explainable Systemic Risk Assessment

Rabimba Karanjai, Hemanth Madhavarao, Lei Xu, Weidong Shi

The interconnected nature of global financial systems makes them vulnerable to systemic risks, where the failure of a few institutions can trigger catastrophic cascading defaults. Traditional risk models often fail to capture the complex, non-linear dynamics of these networks. Wh…

View free PDFSource page
arxivcs.ROcs.AIcs.CVcs.HC2026-06-27

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

Yash Tandon, Giovanni Tapia Lopez, Marcus Blennemann, Mohan Trivedi, Ross Greer

Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around collision avoidance and minimal risk condition (MRC) behaviors such as slowing or stopping when uncertainty arises. Although effective in reduci…

View free PDFSource page
arxivcs.AI2026-07-11

Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models

Xuankun Rong, Wenke Huang, Bo Du, Dacheng Tao, Mang Ye

As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities. Human decision-making combines relatively persistent risk preferences with c…

View free PDFSource page
arxivcs.AI2026-07-14

Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring

Anna Gatzioura, Vrettos Moulos, Nina Baranowska

According to the recent European legislation, high-risk AI systems will have to adapt in order to comply with requirements related to specific areas, like risk management, data quality and governance, logging and traceability, technical documentation, transparency, human oversigh…

View free PDFSource page
arxivcs.AIcs.CY2026-07-08

Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems

M. L. Kaluzhsky, V. A. Efirov

The article investigates the fundamental problem of ensuring the stability of operator control and preserving goal-targeting in hybrid human-machine decision support systems (DSS) of a new generation. Based on a two-month continuous longitudinal experiment on the joint design of…

View free PDFSource page