SURF2026 计算金融项目,研究感知自由边界的美式期权风险曲面,使用 CN/PSOR 基准、文献综述、实验报告以及 Codex 辅助的分步研究规划SURF2026 computational finance project on free-boundary-aware American option risk surfaces, using CN/PSOR benchmarks, literature review, experiment reports, and step-by-step Codex-assisted research planning.
仓库/Skill 库
58 个 · 评测集
基于 SVM 的组学及组学相邻数据的癌症检测系统综述的数据与可复现性材料。Data and reproducibility materials for a systematic review of SVM-based cancer detection using omics and omics-adjacent data.
可复现研究:基于基准、统计模型与走步前向验证,检验可解释的市场信号能否预测 SPY 的五日方向。Reproducible research testing whether interpretable market signals can forecast SPY’s five-day direction using benchmarks, statistical models, and walk-forward validation.
追踪并比较每一个严肃的 OpenClaw 替代方案——基于实测仓库数据与 AI 撰写的决策支持,且刻意保持独立。Track and compare every serious OpenClaw alternative — measured repo data, AI-written decision support, kept apart on purpose.
用于复现 https://arxiv.org/abs/2106.12905 中图示所使用代码与数据的仓库。A repository to story the code and data used to create the figures shown in https://arxiv.org/abs/2106.12905
Shadow / Unmanaged API 的发现、分类、安全验证与可复现研究基准。Shadow / Unmanaged API discovery, classification, security validation, and reproducible research benchmarks
无人机系统(UAS)在生态学研究中应用的系统综述。Systematic Review of UAS use in ecological research
面向系统综述数据抽取、证据定位与偏倚风险评估的冻结式(frozen)防泄漏基准。Frozen, leakage-aware benchmark for systematic-review data extraction, evidence localization, and risk-of-bias support
关于延长 5 天方案对比标准 3 天方案 artemether-lumefantrine 治疗非复杂型 Plasmodium falciparum 疟疾的 systematic review 与 meta-analysis 的可复现工作流、数据、分析与补充材料。Reproducible workflow, data, analyses, and supplementary materials for a systematic review and meta-analysis of extended 5-day versus standard 3-day artemether-lumefantrine for uncomplicated Plasmodium falciparum malaria.
以问题为先、步骤可验证、可自我进化的 agent skill,用于系统综述与 Meta 分析。A question-first, step-verified, self-improving agent skill for systematic reviews and meta-analysis
🚀 使用强化学习优化半精度通用矩阵乘法(HGEMM)CUDA kernel,性能超越 cuBLAS 及其他基准。🚀 Optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels using reinforcement learning, surpassing cuBLAS and other benchmarks with superior performance.
RAMR —— 检索增强记忆可靠性:面向 Agentic-RAG / 记忆系统的抗污染合成基准(附方法与发现)。RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)
系统综述的主动学习摘要筛选。在全部 26 个 SYNERGY 数据集上基准测试:WSS@95 均值 64.1。Active-learning abstract screening for systematic reviews. Benchmarked across all 26 SYNERGY datasets: mean WSS@95 of 64.1.
zhmd:面向中文学术写作的描述性元话语量表。非 AI 检测器。每 15k 字符统计十项 Hyland 式交互与互动标记,并报告机器-人类比率及量表失效条件。为"AI 风味"作为感性(kansei)对象的论文提供配套代码。zhmd: a descriptive metadiscourse scale for Chinese academic writing. Not an AI detector. Counts ten Hyland-style interactive and interactional markers per 15k characters and reports machine-to-human rate ratios together with the conditions under which the scale fails. Companion code for the paper on 'AI flavour' as a kansei object.
本地 LLM 基准测试、RAG、语音交互与 AI 助手Local LLM benchmarking, RAG, voice interaction and AI assistant
🎨 利用 AI 实时生成高保真 UI 设计,对比多版本方案,并跨多个模型导出可直接使用的代码。🎨 Generate high-fidelity UI designs in real-time with AI, compare variations, and export ready-to-use code across multiple models.
将问题框架视为计算状态的可复现研究框架,包含理论、算法与基准,用于内生探究。A reproducible research framework for treating problem frames as computational states, with theory, algorithms, and benchmarks for endogenous inquiry.
基于人类基准的 LLM 生成 Infrastructure-as-Code 安全基准测试。100 场景 × 12 模型配置 = 1,196 个工件,由 Checkov/Trivy/KICS 扫描,并与 634 个人工编写模板对比。所有模型的漏洞密度均为人工的 3.2–3.9 倍。arXiv:2608.28021Human-anchored security benchmark for LLM-generated Infrastructure-as-Code. 100 scenarios × 12 model configs = 1,196 artifacts scanned by Checkov/Trivy/KICS, compared against 634 human-written templates. Every model: 3.2–3.9× human vulnerability density. arXiv:2608.28021
用于 Bosniak 分级可重复性系统综述与 meta 分析的数据及 R 代码Data and R code for the systematic review and meta-analysis of Bosniak classification reproducibility
人工智能(AI)已成为提升乳腺 X 线摄影对乳腺癌诊断准确性的潜力工具。本系统综述与荟萃分析旨在探讨 AI 与放射科医生通过乳腺 X 线摄影检测乳腺癌的准确性。Artificial intelligence (AI) has emerged as a promising tool to improve the diagnostic accuracy of 49 mammography for breast carcinoma detection. This systematic review and meta-analysis aim to 50 explore the accuracy of AI and radiologists in breast carcinoma detection via mammography
面向文档溯源 QA Agent 的生产级 Harness:检索、起草、自评、重写、升级循环,配套 LLM-as-judge 评测套件、分级权限工具、对破坏性操作的人工审批,以及一键容器化部署。A production-grade harness for a document-grounded QA agent: a retrieve, draft, self-score, re-draft, escalate loop with an LLM-as-judge eval suite, permission-tiered tools, human-in-the-loop approval for destructive actions, and a one-command container deploy.
面向 AI coding agents 的 token 预算 context pack,基于编译器解析的 Kotlin 结构(Analysis API/PSI)构建,并附带衡量其是否优于 chunk RAG 的 benchmark。Token-budgeted context packs for AI coding agents, built from compiler-resolved Kotlin structure (Analysis API/PSI) — with the benchmark that measures whether it beats chunk RAG