CORTEXA
← Browse

Noah Shi

5 papers indexed

arxivcs.LGcs.AI2026-07-08

RLVP: Penalize the Path, Reward the Outcome

Bojie Li, Noah Shi

Agents acting on our behalf in the real world (e.g. placing phone calls) must learn online from costly, often irreversible interactions rather than cheap simulator steps. Two things follow. First, deployability depends on the path, not only the outcome. An agent must respect outc…

View free PDFSource page