CORTEXA
← Browse
arxivcs.LG2026-07-15

Counterfactual Optimal Action Trees (COAT): Interpretable Prescriptive Policies from Observational Data

Youssef Drissi, Markus Ettl, Shivaram Subramanian, Wei Sun, Zack Xue

We introduce COAT (Counterfactual Optimal Action Tree), a framework for learning interpretable prescriptive policies from observational data. COAT combines counterfactual outcome estimation with large-scale mixed-integer optimization, using column generation to translate causal predictions into feasible, transparent decisions under business and regulatory constraints. We apply COAT to airline ancillary pricing, a setting characterized by complex business rules and limited experimental flexibility. In a 17-week field pilot with a major global airline, COAT increased upsell revenue per booking by 6.9%, with the airline projecting \$50-\$150 million in incremental annual premium seat revenue across eligible domestic markets. The success of the pilot led to scaled adoption and informed broader AI-driven decision initiatives within the organization.

View free PDFSource page

Related papers

arxivcs.LGstat.ML2026-06-30

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

Mingyi Li, Taira Tsuchiya, Kenji Yamanishi

We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dependent regret bounds. Recent work (Dann et al., 2023; Li et al., 2026) has shown that policy optimiz…

View free PDFSource page
arxivcs.LGcs.AIcs.SE2026-07-08

ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies

Ignacio D. Lopez-Miguel, Ezio Bartocci, Thomas Eiter, Martin Tappler

Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORCAID, a novel method for extracting interpretable r…

View free PDFSource page
arxivcs.ROcs.AIcs.LGeess.SYmath.OC2026-07-16

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan, Glen Chou

World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activation…

View free PDFSource page
arxivcs.LG2026-06-30

Offline Reinforcement Learning for Fluid Controls: Data-based Multi-observational Policy Extraction

Deepak Akhare, Luning Sun, Xin-Yang Liu, Xiantao Fan, Timo Bremer, Ben Zhu, et al.

Active flow control is a fundamental application in engineering. Recent advances in deep reinforcement learning have made progress in this field. However, the classical online RL approaches require extensive real-time interactions with the high fidelity environment, while each se…

View free PDFSource page
arxivcs.LG2026-07-17

Data-Native Global Optimization for Big Data K-means Clustering

Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan

Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids. We target arbitrarily tall data: a fixed feature space ma…

View free PDFSource page