CORTEXA
← Browse

Evgenii Opryshko

2 papers indexed

arxivcs.AI2026-07-22

ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management

Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko, Mark C. Jeffrey

We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows. ARBIGRAPH represents each task as a natural-language problem with an executable P…

View free PDFSource page
arxivcs.LGcs.AI2026-06-27

Modification-Considering Value Learning for Reward Hacking Mitigation in RL

Evgenii Opryshko, Umangi Jain, Igor Gilitschenski

Reinforcement learning agents can exploit misspecified reward signals to achieve high apparent returns while failing on the intended objective, a failure mode known as reward hacking. Existing practical defenses typically constrain policy updates to stay near a known safe referen…

View free PDFSource page