CORTEXA
← Browse

Vinay Kumar Sankarapu

2 papers indexed

arxivcs.LGcs.CLcs.ET2026-07-21

CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability

Pratinav Seth, Hem Gosalia, Aditya Kasliwal, Vinay Kumar Sankarapu

Circuit analysis can support not only model explanation but also downstream interventions such as pruning, editing, steering, and selective fine-tuning. However, conducting such analyses currently requires stitching together separate implementations for discovery, evaluation, and…

View free PDFSource page
arxivcs.CLcs.ETcs.LG2026-07-06

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

Ananth Eswar, Pratinav Seth, Utsav Avaiya, Vinay Kumar Sankarapu

Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing for safety, yet whether they identify causally important rows is rarely tested directly. We address this with two paired audits bu…

View free PDFSource page