Workshop paper
SemCLIP: A Semantic Memory-Aligned Vision Language Model
Tanveer Syeda-Mahmood, Niharika DSouza, et al.
NeurIPS 2025
This talk will focus on designing and evaluating agentic benchmarks with a strong emphasis on in-domain evaluation and real-world task reliability. Drawing from the development of AssetOpsBench, we’ll discuss practical considerations for measuring agent behavior, task completion quality, and decision robustness. The session will highlight what works, what doesn’t, and what matters most when building benchmarks for agent-based systems.
Tanveer Syeda-Mahmood, Niharika DSouza, et al.
NeurIPS 2025
Giovanni De Felice, Arianna Casanova Flores, et al.
NeurIPS 2025
Ramon Nartallo-kaluarachchi, Robert Manson Sawko, et al.
NeurIPS 2025
Max Esposito, Besart Shyti
NeurIPS 2025