Bench2Run

Benchmark 目录

共 2,853 条公开记录,全部通过同一发布门槛(审核通过 · 资源归属已核验 · 已发布)。用下面的筛选缩小范围;代码与数据集状态在各条目内单独标注。

代码
数据集
领域
模态

公开记录

显示前 24 条 / 共 2,853 条

BackTrend 详情

EducationScience & EngineeringWeb & Software

BackTrend is a retrospective benchmark for evaluating scientific weak-signal prediction, containing 25 mature AI/ML topics and 66 human-validated weak signals, with tasks for LLMs, RAG systems, and agentic research systems.

MSI-Bench 详情

Social & Games

MSI-Bench is a new benchmark for evaluating multi-speaker voice interaction in AI agents, with 1,152 test cases in English and Mandarin, targeting memory, instruction following, and reasoning.

EvoPathBench 详情

EducationFinance

EvoPathBench is a benchmark for process-level evaluation of self-evolving agents, tracking individual capabilities across artifact checkpoints using public trading data and calibrated trajectories.

DUMA-Bench 详情

发布团队 · ai-security-lab-itmo ↗CybersecurityRobotics & Embodied

DUMA-Bench is a benchmark and evaluation protocol for measuring LLM agent security under dual-control interaction, where both agent and user influence the environment. It extends tau^2-bench with adversarial environments covering eight vulnerability classes and evaluates 14 model

UK-PRBench 详情

Law & Policy

UK-PRBench is a new benchmark for paragraph-level precedent retrieval in UK case law, constructed from UK National Archives judgments. It evaluates state-of-the-art retrieval models and shows that this task remains challenging.

MemCalib 详情

MemCalib is a benchmark for evaluating how LLM agents use memory in realistic memory-system scenarios, revealing over- and under-use issues. It also introduces an optimization algorithm, MemCalib-RL, but the primary contribution is the benchmark.

IMPLICIT-Bench 详情

IMPLICIT-Bench is a benchmark for measuring implicit bias in text-to-image models using 5,493 controlled prompt triplets across 11 bias categories, with validation via multi-model agreement, CLIP-based verification, and human evaluation.

MCP-GRANITE 详情

MCP-GRANITE is an open-source benchmark framework for evaluating MCP-based LLM agents under edge and IoT scenarios, with 81 multi-step scenarios across 9 domains at 4 granularity levels, measuring task completion, tool selection, argument accuracy, latency, and resource usage.

MCPGen 详情

MCPGen is an executable benchmark for Model Context Protocol workflow development, containing 100 MCP projects across 16 domains, evaluating LLMs on three diagnostic tasks with static analysis, unit/integration tests, and end-to-end execution.

OmicsBench 详情

EducationLaw & PolicyScience & Engineering

OmicsBench is a new reasoning benchmark for multi-omics sequences with 1,160 expert-validated questions across six tasks, requiring traceable evidence chains and evaluated with instance-specific rubrics.

RiverVLN 详情

Robotics & Embodied

RiverVLN is introduced as the first benchmark for long-horizon USV VLN under continuous riverine motion, with a Unity-ROS closed-loop evaluation protocol and real-world deployment tests.

MuLA-Bench 详情

MuLA-Bench is a new benchmark with 5,038 open-ended questions over 1,769 in-the-wild recordings across 16 languages, designed to evaluate audio-language models on long-form audio understanding with controlled language-domain and acoustic tracks.

TicTacBench 详情

Science & Engineering

TicTacBench is a benchmark with 30 tasks for evaluating coding agents' ability to perform RTL-level timing closure under post-PnR evaluation, using suboptimal RTL designs, timing constraints, and verification.

CraftBench-UE 详情

Social & Games

CraftBench-UE is an evaluation harness and benchmark with 70 tasks for coding agents in Unreal Engine, using deterministic build, asset, and runtime checks without an LLM judge.

ISA-Bench 详情

Social & Games

ISA-Bench is a benchmark of programming games with constrained instruction sets, providing full execution stacks for automated evaluation and iterative refinement. It evaluates LLMs' reasoning across unfamiliar instruction set architectures.

MIS-Bench 详情

MIS-Bench is a new benchmark for evaluating multimodal LLMs on assessing psychotherapeutic interpersonal skills, consisting of 996 annotated videos across 8 dimensions. It reveals gaps in current MLLMs' performance and introduces a fine-tuning method to improve scoring.

SEABED 详情

SEABED is a new benchmark dataset for evaluating audio reasoning in Southeast Asian speech, comprising 5,404 question-answer pairs across six tasks, and includes evaluation of audio LLMs.

BrainWideBench 详情

Education

BrainWideBench is a benchmark for evaluating across-animal transfer in multi-region neural recordings, with three task suites for behavior decoding, neural activity prediction, and anatomical organization recovery.

QuranicMMLU 详情

QuranicMMLU is a benchmark for evaluating generative AI on Quranic Arabic linguistic knowledge, with 980 human-reviewed questions across five linguistic pillars, stratified by cognitive level and verse difficulty, and tested on 12 systems.

persona-driven IE benchmark 详情

The paper introduces a new benchmark for personalized information extraction with 292 simulated enterprise users, paired with a persona-generation pipeline, to evaluate LLM-based prompt adaptation methods.