Bench2Run

最近发表的 Benchmark

按论文发表日整理公开资源证据,供研究者快速浏览、核对和继续阅读。

2,851公开目录 2,557有已核验代码 521有已核验数据集

资源归属核验不等于已安装、可运行或端到端复现。

浏览完整目录

本期收录

论文发表日 2026-09-22 · 2

BackTrend 详情

EducationScience & EngineeringWeb & Software

BackTrend is a retrospective benchmark for evaluating scientific weak-signal prediction, containing 25 mature AI/ML topics and 66 human-validated weak signals, with tasks for LLMs, RAG systems, and agentic research systems.

显示该日全部 2 条;完整目录支持搜索和按代码/数据集核验状态筛选。

展开前一天的数据(2026-09-21)

2026-09-21

EvoPathBench 详情

EducationFinance

EvoPathBench is a benchmark for process-level evaluation of self-evolving agents, tracking individual capabilities across artifact checkpoints using public trading data and calibrated trajectories.

DUMA-Bench 详情

发布团队 · ai-security-lab-itmo ↗CybersecurityRobotics & Embodied

DUMA-Bench is a benchmark and evaluation protocol for measuring LLM agent security under dual-control interaction, where both agent and user influence the environment. It extends tau^2-bench with adversarial environments covering eight vulnerability classes and evaluates 14 model

UK-PRBench 详情

Law & Policy

UK-PRBench is a new benchmark for paragraph-level precedent retrieval in UK case law, constructed from UK National Archives judgments. It evaluates state-of-the-art retrieval models and shows that this task remains challenging.

MemCalib 详情

MemCalib is a benchmark for evaluating how LLM agents use memory in realistic memory-system scenarios, revealing over- and under-use issues. It also introduces an optimization algorithm, MemCalib-RL, but the primary contribution is the benchmark.

MCP-GRANITE 详情

MCP-GRANITE is an open-source benchmark framework for evaluating MCP-based LLM agents under edge and IoT scenarios, with 81 multi-step scenarios across 9 domains at 4 granularity levels, measuring task completion, tool selection, argument accuracy, latency, and resource usage.

MCPGen 详情

MCPGen is an executable benchmark for Model Context Protocol workflow development, containing 100 MCP projects across 16 domains, evaluating LLMs on three diagnostic tasks with static analysis, unit/integration tests, and end-to-end execution.