Repositories · organized/repo_cards

仓库/Skill 库

23 个 · 评测集 · 学术写作

排序 Stars 周增
modelscope/evalscope
Python · 2026-08-10 评测基准 评测集 研究原型 Stars 3218 周增 -7

一个精简且可定制的高效大模型(LLM、VLM、AIGC)评估与性能基准测试框架。A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

ragevaluationllm-infra
Sushegaad/Claude-Skills-Governance-Risk-and-Compliance
HTML · 2026-07-20 评测基准 评测集 研究原型 Stars 814 周增 +7

面向治理、风险与合规(GRC)的 Claude Skills:针对 ISO 27001、SOC 2、FedRAMP、GDPR、HIPAA、NIST CSF、PCI DSS、EU AI Act、ISO 42001、ISO 27701、DORA、CSRD、印度 DPDPA、CMMC 2.0、NIST AI Risk、SWIFT、澳大利亚 ISM、EU NIS2、CCPA/CPRA 等的专家级合规指导。使用 skills 基准 97%,不使用 81%。Claude Skills for Governance, Risk, & Compliance (GRC): Expert-level compliance guidance for ISO 27001, SOC 2, FedRAMP, GDPR, HIPAA, NIST CSF, PCI DSS, EU AI Act, ISO 42001, ISO 27701, DORA, CSRD, India's DPDPA, CMMC 2.0, NIST AI Risk, SWIFT, Australia's ISM, EU NIS2, CCPA/CPRA, and others. Benchmark 97% (with skills) vs 81% (without skills).

evaluationrisk
Ar9av/PaperOrchestra
Python · 2026-08-09 Agent 智能体 评测集 研究原型 Stars 636 周增 +7

基于 Google PaperOrchestra 论文实现的全自动 AI 研究论文写作器,通过技能-基准测试 + 自动评分器,配合任意编码 Agent(Claude Code、Cursor、Antigravity、Cline、Aider)。无需 API Key,无需 LLM SDK。An automated AI research-paper writer based off Google's PaperOrchestra paper's implementation through a skills - benchmark + autoraters using any coding agent (Claude Code, Cursor, Antigravity, Cline, Aider). No API keys, no LLM SDKs.

agentevaluationllm-infraengineering
LitLLM/litllms-for-literature-review-tmlr
Python · 2025-04-20 评测基准 评测集 研究原型 Stars 61 周增 -7

论文 LitLLMs, LLMs for Literature Review: Are we there yet?(TMLR 2025)的代码仓库。Code for LitLLMs, LLMs for Literature Review: Are we there yet? (TMLR 2025)

llm-infra
YZCU/OOTB
C · 2025-07-11 多模态 评测集 实验 Stars 53 周增 +0

[ISPRS 2024] 卫星视频单目标跟踪:系统综述与定向目标跟踪基准[ISPRS 2024] Satellite Video Single Object Tracking: A Systematic Review and An Oriented Object Tracking Benchmark

multimodalevaluation
SamInMotion/Medical-intervention-text-classification
Python · 2026-08-13 评测基准 评测集 实验 Stars 2 周增 +0

硕士论文(UiB,计算语言学,2023)及扩展工作。针对系统综述筛选的文本分类。原始 11 工作流分析结合 NEO ontology 集成,以及 Cohen 等人(2006)药物类别基准的扩展,比较 BoW 与 BiomedBERT。Master's thesis (UiB, Computational Linguistics, 2023) and extension work. Text classification for systematic review screening. Original 11-workflow analysis with NEO ontology integration plus Cohen et al. (2006) drug-class benchmark extension comparing BoW and BiomedBERT.

evaluation
HmZ9874/umd-planetary-memory
Python · 2026-08-12 评测基准 评测集 实验 Stars 1 周增 +0

UMD 行星长期记忆算法、公式、基准、SDK 与可复现研究。UMD planetary long-term memory algorithms, formulas, benchmarks, SDKs, and reproducible research

ragevaluationllm-infra
suanlab/pinns-neural-operators-review
Python · 2026-08-25 评测基准 评测集 实验 Stars 0 周增 +0

论文《Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review》(Neural Networks)的补充材料:PRISMA 数据集、PDE 复杂度评分标准以及概念验证的统一基准。Supplementary materials for 'Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review' (Neural Networks): PRISMA datasets, PDE complexity rubric, and proof-of-concept unified benchmark

evaluation
soheylfalahzade/geometric-spanners-lab
Python · 2026-08-16 评测基准 评测集 实验 Stars 0 周增 +0

可复现研究基准实验:二维度量空间下的贪心 t-Spanner 构造算法A reproducible research benchmarking lab for Greedy t-Spanner Construction Algorithms in 2D Metric Spaces.

evaluation
sharing-123/ai-myopia-diagnostic-accuracy-systematic-review
Python · 2026-08-17 评测基准 评测集 研究原型 Stars 0 周增 +0

AI 诊断近视准确性的系统综述:完整数据、提取流程、代码与稿件。Systematic review of AI diagnostic accuracy for myopia detection: full data, extraction, code, and manuscript

savch1102/Systematic-review-vectorborne-models
R · 2026-08-22 评测基准 评测集 研究原型 Stars 0 周增 +0
Ruixixu-hub/2026-surf-american-risk-surfaces
Python · 2026-08-20 评测基准 评测集 实验 Stars 0 周增 +0

SURF2026 计算金融项目,研究感知自由边界的美式期权风险曲面,使用 CN/PSOR 基准、文献综述、实验报告以及 Codex 辅助的分步研究规划SURF2026 computational finance project on free-boundary-aware American option risk surfaces, using CN/PSOR benchmarks, literature review, experiment reports, and step-by-step Codex-assisted research planning.

evaluationrisk
QRSocietyTMU/AlphaProject
Jupyter Notebook · 2026-08-19 评测基准 评测集 实验 Stars 0 周增 +0

可复现研究:基于基准、统计模型与走步前向验证,检验可解释的市场信号能否预测 SPY 的五日方向。Reproducible research testing whether interpretable market signals can forecast SPY’s five-day direction using benchmarks, statistical models, and walk-forward validation.

evaluation
luizalober/dengue-nn-systematic-review
Jupyter Notebook · 2026-08-14 评测基准 评测集 研究原型 Stars 0 周增 +0

用于复现 https://arxiv.org/abs/2106.12905 中图示所使用代码与数据的仓库。A repository to story the code and data used to create the figures shown in https://arxiv.org/abs/2106.12905

kabila5h/Shadow-API-scanner
Python · 2026-08-22 评测基准 评测集 实验 Stars 0 周增 +0

Shadow / Unmanaged API 的发现、分类、安全验证与可复现研究基准。Shadow / Unmanaged API discovery, classification, security validation, and reproducible research benchmarks

evaluationrisk
jan-steen/takeoff
TeX · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

无人机系统(UAS)在生态学研究中应用的系统综述。Systematic Review of UAS use in ecological research

InvestmentMDideas/GLASS-Data-Extraction-Benchmark
Python · 2026-08-16 评测基准 评测集 实验 Stars 0 周增 +0

面向系统综述数据抽取、证据定位与偏倚风险评估的冻结式(frozen)防泄漏基准。Frozen, leakage-aware benchmark for systematic-review data extraction, evidence localization, and risk-of-bias support

evaluationrisk
fsy2004/MetaWingman
Python · 2026-08-24 评测基准 评测集 实验 Stars 0 周增 +0

以问题为先、步骤可验证、可自我进化的 agent skill,用于系统综述与 Meta 分析。A question-first, step-verified, self-improving agent skill for systematic reviews and meta-analysis

agentllm-infra
Chiagoziem2/Systematic-review-screening
Python · 2026-08-19 评测基准 评测集 实验 Stars 0 周增 +0

系统综述的主动学习摘要筛选。在全部 26 个 SYNERGY 数据集上基准测试:WSS@95 均值 64.1。Active-learning abstract screening for systematic reviews. Benchmarked across all 26 SYNERGY datasets: mean WSS@95 of 64.1.

ragevaluation
chenweichiang/zh-metadiscourse-scale
Python · 2026-08-20 评测基准 评测集 研究原型 Stars 0 周增 +0

中文學術寫作的後設論述量尺 · A descriptive metadiscourse scale for Chinese academic writing (not a detector)

engineering
Arithmetic-Power-Geometry/Endogenous-Inquiry-Computing
Python · 2026-08-23 评测基准 评测集 实验 Stars 0 周增 +0

将问题框架视为计算状态的可复现研究框架,包含理论、算法与基准,用于内生探究。A reproducible research framework for treating problem frames as computational states, with theory, algorithms, and benchmarks for endogenous inquiry.

evaluation
ahnhyungwoo/bosniak-reproducibility-meta-analysis
R · 2026-08-16 评测基准 评测集 研究原型 Stars 0 周增 +0

用于 Bosniak 分级可重复性系统综述与 meta 分析的数据及 R 代码Data and R code for the systematic review and meta-analysis of Bosniak classification reproducibility

ADBarshan/diagnostic-accuracy-of-AI-VS-RAdiologist-in-Breast-Carcinoma-Detection-via-Mammography
未知语言 · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

人工智能(AI)已成为提升乳腺 X 线摄影对乳腺癌诊断准确性的潜力工具。本系统综述与荟萃分析旨在探讨 AI 与放射科医生通过乳腺 X 线摄影检测乳腺癌的准确性。Artificial intelligence (AI) has emerged as a promising tool to improve the diagnostic accuracy of 49 mammography for breast carcinoma detection. This systematic review and meta-analysis aim to 50 explore the accuracy of AI and radiologists in breast carcinoma detection via mammography