Repositories · organized/repo_cards

仓库/Skill 库

45 个 · 评测集

排序 Stars 周增
Chiagoziem2/Systematic-review-screening
Python · 2026-08-19 评测基准 评测集 实验 Stars 0 周增 +0

系统综述的主动学习摘要筛选。在全部 26 个 SYNERGY 数据集上基准测试:WSS@95 均值 64.1。Active-learning abstract screening for systematic reviews. Benchmarked across all 26 SYNERGY datasets: mean WSS@95 of 64.1.

ragevaluation
chenweichiang/zh-metadiscourse-scale
Python · 2026-08-20 评测基准 评测集 研究原型 Stars 0 周增 +0

中文學術寫作的後設論述量尺 · A descriptive metadiscourse scale for Chinese academic writing (not a detector)

engineering
Bot87Ever/thunderbolt-ai
Python · 2026-08-21 RAG 检索增强 评测集 实验 Stars 0 周增 +0

本地 LLM 基准测试、RAG、语音交互与 AI 助手Local LLM benchmarking, RAG, voice interaction and AI assistant

ragevaluationllm-infra
avrsnramasamy/AI-DESIGN-BENCHMARK
HTML · 2026-08-20 评测基准 评测集 实验 Stars 0 周增 +0

🎨 利用 AI 实时生成高保真 UI 设计,对比多版本方案,并跨多个模型导出可直接使用的代码。🎨 Generate high-fidelity UI designs in real-time with AI, compare variations, and export ready-to-use code across multiple models.

agentevaluationllm-infra
Arithmetic-Power-Geometry/Endogenous-Inquiry-Computing
Python · 2026-08-23 评测基准 评测集 实验 Stars 0 周增 +0

将问题框架视为计算状态的可复现研究框架,包含理论、算法与基准,用于内生探究。A reproducible research framework for treating problem frames as computational states, with theory, algorithms, and benchmarks for endogenous inquiry.

evaluation
ahnhyungwoo/bosniak-reproducibility-meta-analysis
R · 2026-08-16 评测基准 评测集 研究原型 Stars 0 周增 +0

用于 Bosniak 分级可重复性系统综述与 meta 分析的数据及 R 代码Data and R code for the systematic review and meta-analysis of Bosniak classification reproducibility

ADBarshan/diagnostic-accuracy-of-AI-VS-RAdiologist-in-Breast-Carcinoma-Detection-via-Mammography
未知语言 · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

人工智能(AI)已成为提升乳腺 X 线摄影对乳腺癌诊断准确性的潜力工具。本系统综述与荟萃分析旨在探讨 AI 与放射科医生通过乳腺 X 线摄影检测乳腺癌的准确性。Artificial intelligence (AI) has emerged as a promising tool to improve the diagnostic accuracy of 49 mammography for breast carcinoma detection. This systematic review and meta-analysis aim to 50 explore the accuracy of AI and radiologists in breast carcinoma detection via mammography

aasimmalikin/agentic-qa
Python · 2026-08-11 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向文档溯源 QA Agent 的生产级 Harness:检索、起草、自评、重写、升级循环,配套 LLM-as-judge 评测套件、分级权限工具、对破坏性操作的人工审批,以及一键容器化部署。A production-grade harness for a document-grounded QA agent: a retrieve, draft, self-score, re-draft, escalate loop with an LLM-as-judge eval suite, permission-tiered tools, human-in-the-loop approval for destructive actions, and a one-command container deploy.

agentragengineeringllm-infra
0vertake/jetpacker
Kotlin · 2026-08-17 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向 AI coding agents 的 token 预算 context pack,基于编译器解析的 Kotlin 结构(Analysis API/PSI)构建,并附带衡量其是否优于 chunk RAG 的 benchmark。Token-budgeted context packs for AI coding agents, built from compiler-resolved Kotlin structure (Analysis API/PSI) — with the benchmark that measures whether it beats chunk RAG

agentragevaluationllm-infra