GameHorizon-Bench 详情→
GameHorizon Suite introduces GameHorizon-Bench, a benchmark for evaluating gameplay capabilities across horizons with offline and online tracks, covering 21 games and 47 models.
按论文发表日整理公开资源证据,供研究者快速浏览、核对和继续阅读。
资源归属核验不等于已安装、可运行或端到端复现。
浏览完整目录GameHorizon Suite introduces GameHorizon-Bench, a benchmark for evaluating gameplay capabilities across horizons with offline and online tracks, covering 21 games and 47 models.
BackTrend is a retrospective benchmark for evaluating scientific weak-signal prediction, containing 25 mature AI/ML topics and 66 human-validated weak signals, with tasks for LLMs, RAG systems, and agentic research systems.
显示该日全部 2 条;完整目录支持搜索和按代码/数据集核验状态筛选。
EvoPathBench is a benchmark for process-level evaluation of self-evolving agents, tracking individual capabilities across artifact checkpoints using public trading data and calibrated trajectories.
DUMA-Bench is a benchmark and evaluation protocol for measuring LLM agent security under dual-control interaction, where both agent and user influence the environment. It extends tau^2-bench with adversarial environments covering eight vulnerability classes and evaluates 14 model
UK-PRBench is a new benchmark for paragraph-level precedent retrieval in UK case law, constructed from UK National Archives judgments. It evaluates state-of-the-art retrieval models and shows that this task remains challenging.
MemCalib is a benchmark for evaluating how LLM agents use memory in realistic memory-system scenarios, revealing over- and under-use issues. It also introduces an optimization algorithm, MemCalib-RL, but the primary contribution is the benchmark.
MCP-GRANITE is an open-source benchmark framework for evaluating MCP-based LLM agents under edge and IoT scenarios, with 81 multi-step scenarios across 9 domains at 4 granularity levels, measuring task completion, tool selection, argument accuracy, latency, and resource usage.
MCPGen is an executable benchmark for Model Context Protocol workflow development, containing 100 MCP projects across 16 domains, evaluating LLMs on three diagnostic tasks with static analysis, unit/integration tests, and end-to-end execution.