CORTEXA
← Browse

Sergey Nikolenko

3 papers indexed

arxivcs.LGcs.AI2026-07-18

Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration

Maksim Sheverev, David Finkelstein, Sergey Nikolenko

Long-term memory is becoming a core component of LLM agents, but most memory benchmarks evaluate conversations or compact summaries, while research agents need to restore evidence from full scientific papers. We introduce two full-text scientific-memory benchmarks, Public AI Memo…

View free PDFSource page
arxivcs.LGcs.AI2026-07-18

First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers

Irina Piontkovskaia, Sergey Nikolenko

Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order pre…

View free PDFSource page
arxivcs.AIcs.LGcs.SE2026-07-07

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin, et al.

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, u…

View free PDFSource page