研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

40 个 · 评测基准 · 评测集

排序 Stars 周增
Arithmetic-Power-Geometry/Endogenous-Inquiry-Computing
Python · 2026-08-23 评测基准 评测集 实验 Stars 0 周增 +0

将问题框架视为计算状态的可复现研究框架,包含理论、算法与基准,用于内生探究。A reproducible research framework for treating problem frames as computational states, with theory, algorithms, and benchmarks for endogenous inquiry.

evaluation
AnimeshShaw/GenIaC-SecBench
Python · 2026-08-31 评测基准 评测集 实验 Stars 0 周增 +0

基于人类基准的 LLM 生成 Infrastructure-as-Code 安全基准测试。100 场景 × 12 模型配置 = 1,196 个工件,由 Checkov/Trivy/KICS 扫描,并与 634 个人工编写模板对比。所有模型的漏洞密度均为人工的 3.2–3.9 倍。arXiv:2608.28021Human-anchored security benchmark for LLM-generated Infrastructure-as-Code. 100 scenarios × 12 model configs = 1,196 artifacts scanned by Checkov/Trivy/KICS, compared against 634 human-written templates. Every model: 3.2–3.9× human vulnerability density. arXiv:2608.28021

evaluationriskllm-infra
ahnhyungwoo/bosniak-reproducibility-meta-analysis
R · 2026-08-28 评测基准 评测集 研究原型 Stars 0 周增 +0

用于 Bosniak 分级可重复性系统综述与 meta 分析的数据及 R 代码Data and R code for the systematic review and meta-analysis of Bosniak classification reproducibility

ADBarshan/diagnostic-accuracy-of-AI-VS-RAdiologist-in-Breast-Carcinoma-Detection-via-Mammography
未知语言 · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

人工智能(AI)已成为提升乳腺 X 线摄影对乳腺癌诊断准确性的潜力工具。本系统综述与荟萃分析旨在探讨 AI 与放射科医生通过乳腺 X 线摄影检测乳腺癌的准确性。Artificial intelligence (AI) has emerged as a promising tool to improve the diagnostic accuracy of 49 mammography for breast carcinoma detection. This systematic review and meta-analysis aim to 50 explore the accuracy of AI and radiologists in breast carcinoma detection via mammography