研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

112 个 · 评测基准 · 学术写作

排序 Stars 周增
modelscope/evalscope
Python · 2026-08-10 评测基准 评测集 研究原型 Stars 3218 周增 -7

一个精简且可定制的高效大模型(LLM、VLM、AIGC)评估与性能基准测试框架。A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

ragevaluationllm-infra
yb2460/harness-anything
Python · 2026-07-28 评测基准 工具 研究原型 Stars 1419 周增 +14

Harness Anything - AI agent 控制中枢:支持 WPS、MS Office、Zotero、Photoshop、47 个 CLI 命令、27 项学术技能、SVG 转 PPTXHarness Anything - AI agent control hub: WPS, MS Office, Zotero, Photoshop, 47 CLI commands, 27 academic skills, SVG-to-PPTX

agentengineering
Sushegaad/Claude-Skills-Governance-Risk-and-Compliance
HTML · 2026-07-20 评测基准 评测集 研究原型 Stars 814 周增 +7

面向治理、风险与合规(GRC)的 Claude Skills:针对 ISO 27001、SOC 2、FedRAMP、GDPR、HIPAA、NIST CSF、PCI DSS、EU AI Act、ISO 42001、ISO 27701、DORA、CSRD、印度 DPDPA、CMMC 2.0、NIST AI Risk、SWIFT、澳大利亚 ISM、EU NIS2、CCPA/CPRA 等的专家级合规指导。使用 skills 基准 97%,不使用 81%。Claude Skills for Governance, Risk, & Compliance (GRC): Expert-level compliance guidance for ISO 27001, SOC 2, FedRAMP, GDPR, HIPAA, NIST CSF, PCI DSS, EU AI Act, ISO 42001, ISO 27701, DORA, CSRD, India's DPDPA, CMMC 2.0, NIST AI Risk, SWIFT, Australia's ISM, EU NIS2, CCPA/CPRA, and others. Benchmark 97% (with skills) vs 81% (without skills).

evaluationrisk
Merck/BioPhi
Python · 2025-05-13 评测基准 框架 实验 Stars 259 周增 +0

BioPhi 是开源抗体设计平台,提供自动化抗体人源化方法(Sapiens)、人源性评估(OASis)及计算机辅助抗体序列设计界面BioPhi is an open-source antibody design platform. It features methods for automated antibody humanization (Sapiens), humanness evaluation (OASis) and an interface for computer-assisted antibody sequence design.

evaluation
Epsilon617/Codex-Academic-Skills
Python · 2026-06-16 评测基准 收藏榜 实验 Stars 169 周增 +0

适用于 OpenAI Codex 的研究类 skills 精选清单,覆盖写作、文献综述、评估与研究工作流A curated list of research-oriented skills usable in OpenAI Codex, covering writing, literature review, evaluation, and research workflows.

agentevaluation
clover-Saber/Review-Helper
JavaScript · 2026-01-25 评测基准 工具 实验 Stars 127 周增 +0

Review Helper 是一款 AI 工具,帮助研究者快速摘要、整理和评估学术论文。它通过自动化关键任务来简化文献综述流程,使研究更高效、更准确。Review Helper is an AI tool that helps researchers quickly summarize, organize, and evaluate academic papers. It streamlines literature reviews by automating key tasks, making research more efficient and accurate.

asreview/synergy-dataset
Python · 2026-05-21 评测基准 数据集 研究原型 Stars 111 周增 +0

SYNERGY——系统综述研究筛选的开放机器学习数据集SYNERGY - Open machine learning dataset on study selection in systematic reviews

LitLLM/litllms-for-literature-review-tmlr
Python · 2025-04-20 评测基准 评测集 研究原型 Stars 61 周增 -7

论文 LitLLMs, LLMs for Literature Review: Are we there yet?(TMLR 2025)的代码仓库。Code for LitLLMs, LLMs for Literature Review: Are we there yet? (TMLR 2025)

llm-infra
evidencesynthesis-tools/awesome-evidence-synthesis
未知语言 · 2026-08-26 评测基准 收藏榜 生产可用 Stars 22 周增 +0

Awesome 列表:源自 Evidence Synthesis Tools Directory 的系统综述、Meta 分析与证据综合开源工具精选:octocat: Awesome list of open-source tools for systematic reviews, meta-analysis, and evidence synthesis derived from Evidence Synthesis Tools Directory.

literaf/dsh-ai4scholar
TypeScript · 2026-09-11 评测基准 应用 实验 Stars 21 周增 +0

AI4Scholar for DeepSeek Harness (dsh):38 个原生学术工具——Semantic Scholar、PubMed、Google Scholar、arXiv、bioRxiv/medRxiv、DOI、full text、auto-cite、figures、unified search。由 ai4scholar.net 提供支持AI4Scholar for DeepSeek Harness (dsh): 38 native academic tools — Semantic Scholar, PubMed, Google Scholar, arXiv, bioRxiv/medRxiv, DOI, full text, auto-cite, figures, unified search. Powered by ai4scholar.net

Rogo-Technologies/big-finance-benchmark
Python · 2026-10-07 评测基准 评测集 实验 Stars 20 周增 +0

针对工作流落地型金融研究问题的 Big Finance 基准测试的参考实现 harness。Reference harness for the Big Finance benchmark of workflow-grounded financial-research questions

agentevaluationllm-infra
choxos/jev-reviewer
JavaScript · 2026-09-19 评测基准 工具 实验 Stars 19 周增 +35

系统综述的数据提取,直接引用自论文。向试验报告及其补充材料提交你的提取表单或 RoB 2、ROBINS-I、QUADAS-2、TIDieR 模板;Jev 指向具体行,每条答案均为带页码的逐字引用,由你核对后导出表格。文件保留在浏览器本地。Data extraction for systematic reviews, quoted from the papers. Ask a trial report and its supplements your extraction form or a RoB 2, ROBINS-I, QUADAS-2 or TIDieR template; Jev points at the lines, every answer is a verbatim quote with its page, you check it and export the table. Files stay in your browser.

riskengineering
bgcarlisle/Numbat
PHP · 2026-08-20 评测基准 工具 生产可用 Stars 8 周增 +0

Numbat Systematic Review ManagerNumbat Systematic Review Manager

anas1412/orb-mt5
Python · 2026-09-13 评测基准 应用 实验 Stars 5 周增 +2

面向 MetaTrader 5 的可配置开盘区间突破 EA,附带可复现的研究 harness;所有参数均为输入项,兼容 Windows 与 Linux/WineConfigurable opening-range breakout EA for MetaTrader 5, with a reproducible research harness. Every parameter is an input; runs on Windows and Linux/Wine.

SYSUSELab/From-Data-to-Code
SCSS · 2026-08-13 评测基准 收藏榜 研究原型 Stars 3 周增 +0

系统综述与论文清单:探索面向代码的 LLM 中数据与代码质量问题的映射、检测与治理。Systematic review and paper list exploring the mapping, detection, and governance of data and code quality issues in Large Language Models for Code.

llm-infra
Bardakor/The-De-margining-Artifact
Python · 2026-08-18 评测基准 工具 研究原型 Stars 3 周增 +0

可复现研究代码:检验足球预测优势是否依赖于博彩公司去水方法。Reproducible research code testing whether football forecasting edges depend on bookmaker de-margining methods.

SamInMotion/Medical-intervention-text-classification
Python · 2026-08-13 评测基准 评测集 实验 Stars 2 周增 +0

硕士论文(UiB,计算语言学,2023)及扩展工作。针对系统综述筛选的文本分类。原始 11 工作流分析结合 NEO ontology 集成,以及 Cohen 等人(2006)药物类别基准的扩展,比较 BoW 与 BiomedBERT。Master's thesis (UiB, Computational Linguistics, 2023) and extension work. Text classification for systematic review screening. Original 11-workflow analysis with NEO ontology integration plus Cohen et al. (2006) drug-class benchmark extension comparing BoW and BiomedBERT.

evaluation
literaf/ai4scholar-plugin-dsh
TypeScript · 2026-08-16 评测基准 应用 实验 Stars 2 周增 +0

AI4Scholar for DeepSeek Harness (dsh):38 个原生学术工具——Semantic Scholar、PubMed、Google Scholar、arXiv、bioRxiv/medRxiv、DOI、全文、自动引用、图表、统一搜索。由 ai4scholar.net 提供支持AI4Scholar for DeepSeek Harness (dsh): 38 native academic tools — Semantic Scholar, PubMed, Google Scholar, arXiv, bioRxiv/medRxiv, DOI, full text, auto-cite, figures, unified search. Powered by ai4scholar.net

akaieuan/akaOSS
TypeScript · 2026-09-22 评测基准 收藏榜 实验 Stars 2 周增 +0

akaOSS studio——五个面向人在回路 AI 测量与开发者工具的开源项目,Assist-Not-Complete 论文,可复现研究信息流,以及 HITL Kit 的 shadcn registry。站点 akaoss.dev。The akaOSS studio — five open-source projects for human-in-the-loop AI measurement and developer tooling, the Assist-Not-Complete paper, a reproducible research feed, and the HITL Kit shadcn registry. Live at akaoss.dev.

evaluation
zhoupengyun572-cell/dsh-hana-research
JavaScript · 2026-09-12 评测基准 应用 实验 Stars 1 周增 +0

面向 DeepSeek Harness 的本地文献综述、PDF 标注、证据综合与研究笔记工作台。A local literature review, PDF annotation, evidence synthesis, and research notes workbench for DeepSeek Harness.

suanlab/pinns-neural-operators-review
Python · 2026-09-07 评测基准 评测集 实验 Stars 1 周增 +0

论文《Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review》(Neural Networks)的补充材料:PRISMA 数据集、PDE 复杂度评分标准以及概念验证的统一基准。Supplementary materials for 'Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review' (Neural Networks): PRISMA datasets, PDE complexity rubric, and proof-of-concept unified benchmark

evaluation
rafaelgross/baseref
PHP · 2026-09-12 评测基准 工具 研究原型 Stars 1 周增 +0

系统性综述systematic-review

mikhaeelatefrizk/affect-labeling-review
Python · 2026-09-05 评测基准 数据集 研究原型 Stars 1 周增 +0

对情感标注(Lieberman 等 2007 范式)的系统综述与 Meta 分析。随机效应 Meta 分析(k=8),遵循 PRISMA 2020、RoB 2 / ROBINS-I,约 10,500 词手稿。开放数据与代码。Systematic review and meta-analysis of affect labeling (Lieberman et al. 2007 paradigm). Random-effects meta-analysis (k=8), PRISMA 2020, RoB 2 / ROBINS-I, ~10,500-word manuscript. Open data and code.

LoukiaSpin/Statistical-quality-hyoscine-labour-duration-systematic-reviews
R · 2026-09-01 评测基准 评测集 实验 Stars 1 周增 +0
HmZ9874/umd-planetary-memory
Python · 2026-08-12 评测基准 评测集 实验 Stars 1 周增 +0

UMD 行星长期记忆算法、公式、基准、SDK 与可复现研究。UMD planetary long-term memory algorithms, formulas, benchmarks, SDKs, and reproducible research

ragevaluationllm-infra
herbertkokholm/attest
Python · 2026-10-05 评测基准 工具 研究原型 Stars 1 周增 +0

针对 LLM 集成证据筛选(系统综述标题/摘要筛选)的筛选-自验证内核。Screening-and-self-validation kernel for LLM-ensemble evidence screening (systematic review title/abstract screening).

llm-infra
florianjehn/Societal_Collapse
HTML · 2026-09-01 评测基准 数据集 研究原型 Stars 1 周增 +0

"社会崩溃"主题活体文献综述的永久存储仓库。Repository for permanent storage of a living literature review for societal collapse

rag
delemmaao/BSc_Astrophysics_Projects
Jupyter Notebook · 2026-10-03 评测基准 应用 实验 Stars 1 周增 +0

本人毕业学年所有天体物理课题汇总。毕业设计聚焦轨道转移优化,采用梯度下降法及 Markov Chain Monte Carlo 方法进行不确定性评估These are all the astrophysics projects in my final year. My final-year project focuses on orbital transfer optimisation using gradient descent and the Markov Chain Monte Carlo method for uncertainty evaluation.

llm-infraevaluationengineering
cabbi-bio/miscanthus-yield-maps-review-viz
JavaScript · 2026-08-31 评测基准 应用 实验 Stars 1 周增 +0

GitHub Page 演示:服务于 2026 年文献综述的 Miscanthus 产量地图。Demonstration of GitHub Page with Miscanthus Yield mapping for Literature Review 2026

bristlepine/ilri-climate-adaptation-effectiveness
HTML · 2026-09-23 评测基准 工具 研究原型 Stars 1 周增 +0

ILRI 农业-食品系统气候适应证据综合与系统综述咨询仓库,聚焦"衡量关键:追踪小农户气候适应的有效性"。A repository for the ILRI consultancy on evidence synthesis and systematic reviews of climate adaptation in agri-food systems, focused on “Measuring what matters: tracking the effectiveness of climate adaptation for smallholder producers.”

WENHA0ZHANG/text_mining-based_literature_reviewer
Jupyter Notebook · 2026-08-20 评测基准 工具 研究原型 Stars 0 周增 +0

由文本挖掘驱动的科学文献综述。A text mining-driven review of scientific literature.

wbendinelli/citation-audit
Python · 2026-09-04 评测基准 工具 研究原型 Stars 0 周增 +0

对两篇文章所获引用的质量审计:可复现的方法、数据、分析与报告Auditoria da qualidade das citações recebidas por dois artigos: método, dados, análises e relatório reprodutíveis

vaidehimhamane15-ui/Task---1
未知语言 · 2026-08-18 评测基准 教程 实验 Stars 0 周增 +0

通过系统综述和患者病例场景分析识别潜在药物不良反应(ADR),包括对症状、用药史、剂量及时间线的评估,以确定可疑药物并判断所报告反应是否可能与药物相关。项目展示了基础的 Pharmacovigilance 能力identifying potential Adverse Drug Reactions through systematic review and analysis of patient case scenarios. It includes evaluation of symptoms, medication history, dosage, and timelines to determine the suspected drug and assess whether the reported reaction is potentially drug-related. The project demonstrates basic Pharmacovigilance

evaluation
Triple3A/AV-CAV-Congestion-Review-Data
Python · 2026-10-05 评测基准 数据集 研究原型 Stars 0 周增 +0

支撑 AV/CAV 交通拥堵缓解半系统综述的数据与代码。Data and coding supporting a semi-systematic review of AV/CAV congestion mitigation.

tav0-m/mm-ipsa-research
Python · 2026-09-12 评测基准 应用 实验 Stars 0 周增 +0

围绕矩匹配情景生成与时序投资组合评估的可复现研究。Reproducible research on moment-matching scenario generation and temporal portfolio evaluation

evaluation
ssemerikov/gamification-of-history-education---a-systematic-review
Python · 2026-08-30 评测基准 收藏榜 研究原型 Stars 0 周增 +0