研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

58 个 · 评测集

排序 Stars 周增
Ruixixu-hub/2026-surf-american-risk-surfaces
Python · 2026-08-20 评测基准 评测集 实验 Stars 0 周增 +0

SURF2026 计算金融项目,研究感知自由边界的美式期权风险曲面,使用 CN/PSOR 基准、文献综述、实验报告以及 Codex 辅助的分步研究规划SURF2026 computational finance project on free-boundary-aware American option risk surfaces, using CN/PSOR benchmarks, literature review, experiment reports, and step-by-step Codex-assisted research planning.

evaluationrisk
rojinamdii/svm-cancer-omics-systematic-review
R · 2026-09-07 评测基准 评测集 研究原型 Stars 0 周增 +0

基于 SVM 的组学及组学相邻数据的癌症检测系统综述的数据与可复现性材料。Data and reproducibility materials for a systematic review of SVM-based cancer detection using omics and omics-adjacent data.

QRSocietyTMU/AlphaProject
Jupyter Notebook · 2026-09-02 评测基准 评测集 实验 Stars 0 周增 +0

可复现研究:基于基准、统计模型与走步前向验证,检验可解释的市场信号能否预测 SPY 的五日方向。Reproducible research testing whether interpretable market signals can forecast SPY’s five-day direction using benchmarks, statistical models, and walk-forward validation.

evaluation
naturalmoods/clawclones
TypeScript · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

追踪并比较每一个严肃的 OpenClaw 替代方案——基于实测仓库数据与 AI 撰写的决策支持,且刻意保持独立。Track and compare every serious OpenClaw alternative — measured repo data, AI-written decision support, kept apart on purpose.

agent
luizalober/dengue-nn-systematic-review
Jupyter Notebook · 2026-08-14 评测基准 评测集 研究原型 Stars 0 周增 +0

用于复现 https://arxiv.org/abs/2106.12905 中图示所使用代码与数据的仓库。A repository to story the code and data used to create the figures shown in https://arxiv.org/abs/2106.12905

kabila5h/Shadow-API-scanner
Python · 2026-08-22 评测基准 评测集 实验 Stars 0 周增 +0

Shadow / Unmanaged API 的发现、分类、安全验证与可复现研究基准。Shadow / Unmanaged API discovery, classification, security validation, and reproducible research benchmarks

evaluationrisk
jan-steen/takeoff
TeX · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

无人机系统(UAS)在生态学研究中应用的系统综述。Systematic Review of UAS use in ecological research

InvestmentMDideas/GLASS-Data-Extraction-Benchmark
Python · 2026-08-16 评测基准 评测集 实验 Stars 0 周增 +0

面向系统综述数据抽取、证据定位与偏倚风险评估的冻结式(frozen)防泄漏基准。Frozen, leakage-aware benchmark for systematic-review data extraction, evidence localization, and risk-of-bias support

evaluationrisk
gpaasi/al-5day-vs-3day-systematic-review
R · 2026-08-28 评测基准 评测集 研究原型 Stars 0 周增 +0

关于延长 5 天方案对比标准 3 天方案 artemether-lumefantrine 治疗非复杂型 Plasmodium falciparum 疟疾的 systematic review 与 meta-analysis 的可复现工作流、数据、分析与补充材料。Reproducible workflow, data, analyses, and supplementary materials for a systematic review and meta-analysis of extended 5-day versus standard 3-day artemether-lumefantrine for uncomplicated Plasmodium falciparum malaria.

fsy2004/MetaWingman
Python · 2026-08-30 评测基准 评测集 实验 Stars 0 周增 +0

以问题为先、步骤可验证、可自我进化的 agent skill,用于系统综述与 Meta 分析。A question-first, step-verified, self-improving agent skill for systematic reviews and meta-analysis

agentllm-infra
friedpotato04/CUDA-L2
Cuda · 2026-09-10 LLM 基础设施 评测集 实验 Stars 0 周增 +0

🚀 使用强化学习优化半精度通用矩阵乘法(HGEMM)CUDA kernel,性能超越 cuBLAS 及其他基准。🚀 Optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels using reinforcement learning, surpassing cuBLAS and other benchmarks with superior performance.

evaluationllm-infra
DanceNitra/ramr
Python · 2026-08-12 RAG 检索增强 评测集 实验 Stars 0 周增 +0

RAMR —— 检索增强记忆可靠性:面向 Agentic-RAG / 记忆系统的抗污染合成基准(附方法与发现)。RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)

agentragevaluationllm-infra
Chiagoziem2/Systematic-review-screening
Python · 2026-08-19 评测基准 评测集 实验 Stars 0 周增 +0

系统综述的主动学习摘要筛选。在全部 26 个 SYNERGY 数据集上基准测试:WSS@95 均值 64.1。Active-learning abstract screening for systematic reviews. Benchmarked across all 26 SYNERGY datasets: mean WSS@95 of 64.1.

ragevaluation
chenweichiang/zh-metadiscourse-scale
Python · 2026-09-26 评测基准 评测集 实验 Stars 0 周增 +0

zhmd:面向中文学术写作的描述性元话语量表。非 AI 检测器。每 15k 字符统计十项 Hyland 式交互与互动标记,并报告机器-人类比率及量表失效条件。为"AI 风味"作为感性(kansei)对象的论文提供配套代码。zhmd: a descriptive metadiscourse scale for Chinese academic writing. Not an AI detector. Counts ten Hyland-style interactive and interactional markers per 15k characters and reports machine-to-human rate ratios together with the conditions under which the scale fails. Companion code for the paper on 'AI flavour' as a kansei object.

engineering
Bot87Ever/thunderbolt-ai
Python · 2026-08-21 RAG 检索增强 评测集 实验 Stars 0 周增 +0

本地 LLM 基准测试、RAG、语音交互与 AI 助手Local LLM benchmarking, RAG, voice interaction and AI assistant

ragevaluationllm-infra
avrsnramasamy/AI-DESIGN-BENCHMARK
HTML · 2026-09-20 评测基准 评测集 实验 Stars 0 周增 +0

🎨 利用 AI 实时生成高保真 UI 设计,对比多版本方案,并跨多个模型导出可直接使用的代码。🎨 Generate high-fidelity UI designs in real-time with AI, compare variations, and export ready-to-use code across multiple models.

agentevaluationllm-infra
Arithmetic-Power-Geometry/Endogenous-Inquiry-Computing
Python · 2026-08-23 评测基准 评测集 实验 Stars 0 周增 +0

将问题框架视为计算状态的可复现研究框架,包含理论、算法与基准,用于内生探究。A reproducible research framework for treating problem frames as computational states, with theory, algorithms, and benchmarks for endogenous inquiry.

evaluation
AnimeshShaw/GenIaC-SecBench
Python · 2026-08-31 评测基准 评测集 实验 Stars 0 周增 +0

基于人类基准的 LLM 生成 Infrastructure-as-Code 安全基准测试。100 场景 × 12 模型配置 = 1,196 个工件,由 Checkov/Trivy/KICS 扫描,并与 634 个人工编写模板对比。所有模型的漏洞密度均为人工的 3.2–3.9 倍。arXiv:2608.28021Human-anchored security benchmark for LLM-generated Infrastructure-as-Code. 100 scenarios × 12 model configs = 1,196 artifacts scanned by Checkov/Trivy/KICS, compared against 634 human-written templates. Every model: 3.2–3.9× human vulnerability density. arXiv:2608.28021

evaluationriskllm-infra
ahnhyungwoo/bosniak-reproducibility-meta-analysis
R · 2026-08-28 评测基准 评测集 研究原型 Stars 0 周增 +0

用于 Bosniak 分级可重复性系统综述与 meta 分析的数据及 R 代码Data and R code for the systematic review and meta-analysis of Bosniak classification reproducibility

ADBarshan/diagnostic-accuracy-of-AI-VS-RAdiologist-in-Breast-Carcinoma-Detection-via-Mammography
未知语言 · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

人工智能(AI)已成为提升乳腺 X 线摄影对乳腺癌诊断准确性的潜力工具。本系统综述与荟萃分析旨在探讨 AI 与放射科医生通过乳腺 X 线摄影检测乳腺癌的准确性。Artificial intelligence (AI) has emerged as a promising tool to improve the diagnostic accuracy of 49 mammography for breast carcinoma detection. This systematic review and meta-analysis aim to 50 explore the accuracy of AI and radiologists in breast carcinoma detection via mammography

aasimmalikin/agentic-qa
Python · 2026-08-11 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向文档溯源 QA Agent 的生产级 Harness:检索、起草、自评、重写、升级循环,配套 LLM-as-judge 评测套件、分级权限工具、对破坏性操作的人工审批,以及一键容器化部署。A production-grade harness for a document-grounded QA agent: a retrieve, draft, self-score, re-draft, escalate loop with an LLM-as-judge eval suite, permission-tiered tools, human-in-the-loop approval for destructive actions, and a one-command container deploy.

agentragengineeringllm-infra
0vertake/jetpacker
Kotlin · 2026-08-17 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向 AI coding agents 的 token 预算 context pack,基于编译器解析的 Kotlin 结构(Analysis API/PSI)构建,并附带衡量其是否优于 chunk RAG 的 benchmark。Token-budgeted context packs for AI coding agents, built from compiler-resolved Kotlin structure (Analysis API/PSI) — with the benchmark that measures whether it beats chunk RAG

agentragevaluationllm-infra