arxivcs.AI2026-07-31
ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models
Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently ch…