将问题框架视为计算状态的可复现研究框架,包含理论、算法与基准,用于内生探究。A reproducible research framework for treating problem frames as computational states, with theory, algorithms, and benchmarks for endogenous inquiry.
仓库/Skill 库
40 个 · 评测基准 · 评测集
基于人类基准的 LLM 生成 Infrastructure-as-Code 安全基准测试。100 场景 × 12 模型配置 = 1,196 个工件,由 Checkov/Trivy/KICS 扫描,并与 634 个人工编写模板对比。所有模型的漏洞密度均为人工的 3.2–3.9 倍。arXiv:2608.28021Human-anchored security benchmark for LLM-generated Infrastructure-as-Code. 100 scenarios × 12 model configs = 1,196 artifacts scanned by Checkov/Trivy/KICS, compared against 634 human-written templates. Every model: 3.2–3.9× human vulnerability density. arXiv:2608.28021
用于 Bosniak 分级可重复性系统综述与 meta 分析的数据及 R 代码Data and R code for the systematic review and meta-analysis of Bosniak classification reproducibility
人工智能(AI)已成为提升乳腺 X 线摄影对乳腺癌诊断准确性的潜力工具。本系统综述与荟萃分析旨在探讨 AI 与放射科医生通过乳腺 X 线摄影检测乳腺癌的准确性。Artificial intelligence (AI) has emerged as a promising tool to improve the diagnostic accuracy of 49 mammography for breast carcinoma detection. This systematic review and meta-analysis aim to 50 explore the accuracy of AI and radiologists in breast carcinoma detection via mammography