GameHorizon-Bench 详情→
GameHorizon Suite introduces GameHorizon-Bench, a benchmark for evaluating gameplay capabilities across horizons with offline and online tracks, covering 21 games and 47 models.
共 2,853 条公开记录,全部通过同一发布门槛(审核通过 · 资源归属已核验 · 已发布)。用下面的筛选缩小范围;代码与数据集状态在各条目内单独标注。
GameHorizon Suite introduces GameHorizon-Bench, a benchmark for evaluating gameplay capabilities across horizons with offline and online tracks, covering 21 games and 47 models.
BackTrend is a retrospective benchmark for evaluating scientific weak-signal prediction, containing 25 mature AI/ML topics and 66 human-validated weak signals, with tasks for LLMs, RAG systems, and agentic research systems.
MSI-Bench is a new benchmark for evaluating multi-speaker voice interaction in AI agents, with 1,152 test cases in English and Mandarin, targeting memory, instruction following, and reasoning.
EvoPathBench is a benchmark for process-level evaluation of self-evolving agents, tracking individual capabilities across artifact checkpoints using public trading data and calibrated trajectories.
DUMA-Bench is a benchmark and evaluation protocol for measuring LLM agent security under dual-control interaction, where both agent and user influence the environment. It extends tau^2-bench with adversarial environments covering eight vulnerability classes and evaluates 14 model
UK-PRBench is a new benchmark for paragraph-level precedent retrieval in UK case law, constructed from UK National Archives judgments. It evaluates state-of-the-art retrieval models and shows that this task remains challenging.
MemCalib is a benchmark for evaluating how LLM agents use memory in realistic memory-system scenarios, revealing over- and under-use issues. It also introduces an optimization algorithm, MemCalib-RL, but the primary contribution is the benchmark.
IMPLICIT-Bench is a benchmark for measuring implicit bias in text-to-image models using 5,493 controlled prompt triplets across 11 bias categories, with validation via multi-model agreement, CLIP-based verification, and human evaluation.
MCP-GRANITE is an open-source benchmark framework for evaluating MCP-based LLM agents under edge and IoT scenarios, with 81 multi-step scenarios across 9 domains at 4 granularity levels, measuring task completion, tool selection, argument accuracy, latency, and resource usage.
MCPGen is an executable benchmark for Model Context Protocol workflow development, containing 100 MCP projects across 16 domains, evaluating LLMs on three diagnostic tasks with static analysis, unit/integration tests, and end-to-end execution.
VibeMemBench is a benchmark for evaluating memory systems in coding agents, using 111 coding targets from 90 SWE-rebench V2 repositories with 3,634 history trajectories. It measures whether memory systems improve executable task outcomes compared to baselines.
OmicsBench is a new reasoning benchmark for multi-omics sequences with 1,160 expert-validated questions across six tasks, requiring traceable evidence chains and evaluated with instance-specific rubrics.
RiverVLN is introduced as the first benchmark for long-horizon USV VLN under continuous riverine motion, with a Unity-ROS closed-loop evaluation protocol and real-world deployment tests.
MuLA-Bench is a new benchmark with 5,038 open-ended questions over 1,769 in-the-wild recordings across 16 languages, designed to evaluate audio-language models on long-form audio understanding with controlled language-domain and acoustic tracks.
TicTacBench is a benchmark with 30 tasks for evaluating coding agents' ability to perform RTL-level timing closure under post-PnR evaluation, using suboptimal RTL designs, timing constraints, and verification.
CraftBench-UE is an evaluation harness and benchmark with 70 tasks for coding agents in Unreal Engine, using deterministic build, asset, and runtime checks without an LLM judge.
LD-RSVIS is a new benchmark for robust and general referring surgical video instrument segmentation.
ISA-Bench is a benchmark of programming games with constrained instruction sets, providing full execution stacks for automated evaluation and iterative refinement. It evaluates LLMs' reasoning across unfamiliar instruction set architectures.
MIS-Bench is a new benchmark for evaluating multimodal LLMs on assessing psychotherapeutic interpersonal skills, consisting of 996 annotated videos across 8 dimensions. It reveals gaps in current MLLMs' performance and introduces a fine-tuning method to improve scoring.
SEABED is a new benchmark dataset for evaluating audio reasoning in Southeast Asian speech, comprising 5,404 question-answer pairs across six tasks, and includes evaluation of audio LLMs.
BrainWideBench is a benchmark for evaluating across-animal transfer in multi-region neural recordings, with three task suites for behavior decoding, neural activity prediction, and anatomical organization recovery.
QuranicMMLU is a benchmark for evaluating generative AI on Quranic Arabic linguistic knowledge, with 980 human-reviewed questions across five linguistic pillars, stratified by cognitive level and verse difficulty, and tested on 12 systems.
RecreationWorld introduces RecreationBench, a benchmark of 250 tasks across platforms for hybrid computer-use agents, with reference-grounded programmatic and visual assertions for automatic scoring.
The paper introduces a new benchmark for personalized information extraction with 292 simulated enterprise users, paired with a persona-generation pipeline, to evaluate LLM-based prompt adaptation methods.