Repositories · organized/repo_cards

仓库/Skill 库

172 个 · 评测基准

排序 Stars 周增
yyh-001/llm-value-rankings
CSS · 2026-08-11 评测基准 工具 实验 Stars 12 周增 +0

每日 LLM 价值排行榜——基于智能、速度、价格对比 300+ 模型。OpenRouter + Artificial Analysis。大模型性价比排行榜Daily LLM value rankings - compare 300+ models by intelligence, speed and price. OpenRouter + Artificial Analysis. 大模型性价比排行榜

evaluationllm-infra
literaf/dsh-ai4scholar
TypeScript · 2026-08-18 评测基准 应用 实验 Stars 10 周增 +0

AI4Scholar for DeepSeek Harness (dsh):38 个原生学术工具——Semantic Scholar、PubMed、Google Scholar、arXiv、bioRxiv/medRxiv、DOI、full text、auto-cite、figures、unified search。由 ai4scholar.net 提供支持AI4Scholar for DeepSeek Harness (dsh): 38 native academic tools — Semantic Scholar, PubMed, Google Scholar, arXiv, bioRxiv/medRxiv, DOI, full text, auto-cite, figures, unified search. Powered by ai4scholar.net

MARKTECHPOST-AI-MEDIA-INC/LLMs-Tutorials-Projects
未知语言 · 2026-08-11 评测基准 教程 实验 Stars 9 周增 +0

微调、评估、提示工程、开源模型Fine-tuning, evaluation, prompting, open-source models

evaluationllm-infra
KCNyu/clawock
Python · 2026-08-11 评测基准 应用 实验 Stars 8 周增 +0

AI 辩论,代码裁定,亏损留在账面上。一款可移植的投资决策工作流插件与可验证 harness,已在真实的港股+美股组合上验证。AI argues. Code settles. The losses stay on the page. A portable investment decision-workflow plugin and verifiable harness, proven on a real HK + US portfolio.

agentriskllm-infra
bgcarlisle/Numbat
PHP · 2026-08-20 评测基准 工具 生产可用 Stars 8 周增 +0

Numbat Systematic Review ManagerNumbat Systematic Review Manager

sb1733831438-maker/DSH-closerAI
TypeScript · 2026-08-18 评测基准 模型 实验 Stars 4 周增 +0

CloserAI — 基于 DeepSeek Harness 构建的本地优先、模型无关、权限透明的桌面 AI 工作台。CloserAI - a local-first, model-agnostic, permission-transparent desktop AI workbench built on DeepSeek Harness.

agent
SYSUSELab/From-Data-to-Code
SCSS · 2026-08-13 评测基准 收藏榜 研究原型 Stars 3 周增 +0

系统综述与论文清单:探索面向代码的 LLM 中数据与代码质量问题的映射、检测与治理。Systematic review and paper list exploring the mapping, detection, and governance of data and code quality issues in Large Language Models for Code.

llm-infra
saurabhr/psychscanner
HTML · 2026-08-13 评测基准 工具 研究原型 Stars 3 周增 +0

自动化并规模化 "LLMs as a participant",将 LLM 作为研究参与者Automates and scales "LLMs as a participant."

evaluationllm-infra
linny006/vector-db-live
Python · 2026-08-25 评测基准 评测集 实验 Stars 3 周增 +0

实时更新的向量数据库项目、集成和基准评测全景图——每……刷新。Live-updating landscape of vector database projects, integrations, and benchmarks — refreshed every

ragevaluationdatabasellm-infra
harshtiwari01/llm-heatmap-visualizer
Jupyter Notebook · 2026-08-11 评测基准 工具 实验 Stars 3 周增 +0

用于生成基于 Transformer 的 LLM 完整注意力头热力图的一组脚本。A set of scripts to generate full attention-head heatmaps for transformer-based LLMs

llm-infraevaluation
Bardakor/The-De-margining-Artifact
Python · 2026-08-18 评测基准 工具 研究原型 Stars 3 周增 +0

可复现研究代码:检验足球预测优势是否依赖于博彩公司去水方法。Reproducible research code testing whether football forecasting edges depend on bookmaker de-margining methods.

SamInMotion/Medical-intervention-text-classification
Python · 2026-08-13 评测基准 评测集 实验 Stars 2 周增 +0

硕士论文(UiB,计算语言学,2023)及扩展工作。针对系统综述筛选的文本分类。原始 11 工作流分析结合 NEO ontology 集成,以及 Cohen 等人(2006)药物类别基准的扩展,比较 BoW 与 BiomedBERT。Master's thesis (UiB, Computational Linguistics, 2023) and extension work. Text classification for systematic review screening. Original 11-workflow analysis with NEO ontology integration plus Cohen et al. (2006) drug-class benchmark extension comparing BoW and BiomedBERT.

evaluation
literaf/ai4scholar-plugin-dsh
TypeScript · 2026-08-16 评测基准 应用 实验 Stars 2 周增 +0

AI4Scholar for DeepSeek Harness (dsh):38 个原生学术工具——Semantic Scholar、PubMed、Google Scholar、arXiv、bioRxiv/medRxiv、DOI、全文、自动引用、图表、统一搜索。由 ai4scholar.net 提供支持AI4Scholar for DeepSeek Harness (dsh): 38 native academic tools — Semantic Scholar, PubMed, Google Scholar, arXiv, bioRxiv/medRxiv, DOI, full text, auto-cite, figures, unified search. Powered by ai4scholar.net

akaieuan/akaOSS
TypeScript · 2026-08-13 评测基准 收藏榜 实验 Stars 2 周增 +0

akaOSS studio——五个面向人在回路 AI 测量与开发者工具的开源项目,Assist-Not-Complete 论文,可复现研究信息流,以及 HITL Kit 的 shadcn registry。站点 akaoss.dev。The akaOSS studio — five open-source projects for human-in-the-loop AI measurement and developer tooling, the Assist-Not-Complete paper, a reproducible research feed, and the HITL Kit shadcn registry. Live at akaoss.dev.

evaluation
Zhuchen00123/Verified-Executable-Search
Python · 2026-08-11 评测基准 工具 实验 Stars 1 周增 +0

面向 AI 生成可执行方案的 verifier-first 运行时,支持独立验证、对比与检索。A verifier-first runtime for independently verifying, comparing, and searching AI-generated executable solutions.

agentriskllm-infra
yyxcnasd/amadeus-for-dsh
JavaScript · 2026-08-15 评测基准 应用 实验 Stars 1 周增 +0

Amadeus(来自《Steins;Gate 0》的 AI 助手)适配 DeepSeek Harness。Amadeus (AI assistant from Steins;Gate 0) for DeepSeek Harness

agent
NinjaSln-labs/dsh-plugins
TypeScript · 2026-08-16 评测基准 应用 实验 Stars 1 周增 +0

DeepSeek Harness 个人自研插件集:上下文罗盘 / 跨会话知识 / 子代理模型路由 / AI 生图(Personally developed plugins for DeepSeek Harness)

agentllm-infra
kunal4040/hybrid-search-eval
Python · 2026-08-13 评测基准 评测集 实验 Stars 1 周增 +0

🔍 在 Weaviate 混合检索中评测 embedding 模型,基于自有数据或 MTEB 数据集评估 MRR@K、Hit@K、延迟与内存占用🔍 Benchmark embedding models in hybrid search with Weaviate. Evaluate MRR@K, Hit@K, latency, and memory using your data or MTEB datasets.

ragllm-infraevaluationdatabase
HmZ9874/umd-planetary-memory
Python · 2026-08-12 评测基准 评测集 实验 Stars 1 周增 +0

UMD 行星长期记忆算法、公式、基准、SDK 与可复现研究。UMD planetary long-term memory algorithms, formulas, benchmarks, SDKs, and reproducible research

ragevaluationllm-infra
herbertkokholm/attest
Python · 2026-08-25 评测基准 工具 研究原型 Stars 1 周增 +0

针对 LLM 集成证据筛选(系统综述标题/摘要筛选)的筛选-自验证内核。Screening-and-self-validation kernel for LLM-ensemble evidence screening (systematic review title/abstract screening).

llm-infra
florianjehn/Societal_Collapse
HTML · 2026-08-17 评测基准 数据集 研究原型 Stars 1 周增 +0

"社会崩溃"主题活体文献综述的永久存储仓库。Repository for permanent storage of a living literature review for societal collapse

rag
delemmaao/BSc_Astrophysics_Projects
Jupyter Notebook · 2026-08-22 评测基准 应用 实验 Stars 1 周增 +0

本人毕业学年所有天体物理课题汇总。毕业设计聚焦轨道转移优化,采用梯度下降法及 Markov Chain Monte Carlo 方法进行不确定性评估These are all the astrophysics projects in my final year. My final-year project focuses on orbital transfer optimisation using gradient descent and the Markov Chain Monte Carlo method for uncertainty evaluation.

llm-infraevaluationengineering
chrisliu298/awesome-rubric-rewards
未知语言 · 2026-08-11 评测基准 收藏榜 生产可用 Stars 1 周增 +0

精选的评分量表、检查清单、评分标准集、原则和评分指南汇总,用于对现代生成模型进行评分、排名、验证、过滤或训练。A curated list of rubrics, checklists, criteria sets, principles, and scoring guides used to score, rank, verify, filter, or train modern generative models.

evaluationriskllm-infra
bristlepine/ilri-climate-adaptation-effectiveness
HTML · 2026-08-19 评测基准 工具 研究原型 Stars 1 周增 +0

ILRI 农业-食品系统气候适应证据综合与系统综述咨询仓库,聚焦"衡量关键:追踪小农户气候适应的有效性"。A repository for the ILRI consultancy on evidence synthesis and systematic reviews of climate adaptation in agri-food systems, focused on “Measuring what matters: tracking the effectiveness of climate adaptation for smallholder producers.”

wolfenix/llm-math-reasoning-analysis
HTML · 2026-08-11 评测基准 工具 实验 Stars 0 周增 +0

🔍 使用多种 prompt 技巧在多步数学问题上分析 Mistral-7B 模型的数学推理能力。🔍 Analyze the mathematical reasoning abilities of the Mistral-7B model using diverse prompting techniques on multi-step math problems.

evaluationllm-infra
WENHA0ZHANG/text_mining-based_literature_reviewer
Jupyter Notebook · 2026-08-20 评测基准 工具 研究原型 Stars 0 周增 +0

由文本挖掘驱动的科学文献综述。A text mining-driven review of scientific literature.

vaidehimhamane15-ui/Task---1
未知语言 · 2026-08-18 评测基准 教程 实验 Stars 0 周增 +0

通过系统综述和患者病例场景分析识别潜在药物不良反应(ADR),包括对症状、用药史、剂量及时间线的评估,以确定可疑药物并判断所报告反应是否可能与药物相关。项目展示了基础的 Pharmacovigilance 能力identifying potential Adverse Drug Reactions through systematic review and analysis of patient case scenarios. It includes evaluation of symptoms, medication history, dosage, and timelines to determine the suspected drug and assess whether the reported reaction is potentially drug-related. The project demonstrates basic Pharmacovigilance

evaluation
Triple3A/AV-CAV-Congestion-Review-Data
未知语言 · 2026-08-25 评测基准 数据集 研究原型 Stars 0 周增 +0

支撑 AV/CAV 交通拥堵缓解半系统综述的数据与代码。Data and coding supporting a semi-systematic review of AV/CAV congestion mitigation.

tav0-m/mm-ipsa-research
Python · 2026-08-14 评测基准 应用 实验 Stars 0 周增 +0

围绕矩匹配情景生成与时序投资组合评估的可复现研究。Reproducible research on moment-matching scenario generation and temporal portfolio evaluation

evaluation
synaptiai/flow-harness
TypeScript · 2026-08-20 评测基准 应用 实验 Stars 0 周增 +0

与供应商无关的 coding-agent 框架,具备确定性工作流图、持久化证据与 fail-closed 沙箱执行。Provider-neutral coding-agent harness with deterministic workflow graphs, durable evidence, and fail-closed sandboxed execution

agentllm-infra
sumit7194/Phronesis
HTML · 2026-08-16 评测基准 工具 实验 Stars 0 周增 +0

在小型 LLM 中植入认知美德的激活引导实验——三项发现与一个失败模式数据集(DOI 见 README)Activation-steering experiments on installing epistemic virtues in small LLMs — three findings + a failure-mode dataset (DOIs in README).

riskllm-infra
suanlab/pinns-neural-operators-review
Python · 2026-08-25 评测基准 评测集 实验 Stars 0 周增 +0

论文《Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review》(Neural Networks)的补充材料:PRISMA 数据集、PDE 复杂度评分标准以及概念验证的统一基准。Supplementary materials for 'Physics-Informed Neural Networks and Neural Operators for PDEs: A Unified Taxonomy and Systematic Review' (Neural Networks): PRISMA datasets, PDE complexity rubric, and proof-of-concept unified benchmark

evaluation
soheylfalahzade/geometric-spanners-lab
Python · 2026-08-16 评测基准 评测集 实验 Stars 0 周增 +0

可复现研究基准实验:二维度量空间下的贪心 t-Spanner 构造算法A reproducible research benchmarking lab for Greedy t-Spanner Construction Algorithms in 2D Metric Spaces.

evaluation
sm0704/abstractScreening
Python · 2026-08-21 评测基准 工具 研究原型 Stars 0 周增 +0

系统综述工具:面向学生自我关怀与学业功能的全文筛选与 Covidence 数据提取。Systematic-review tooling: full-text screening and Covidence data extraction for self-compassion and academic functioning in students

sharing-123/ai-myopia-diagnostic-accuracy-systematic-review
Python · 2026-08-17 评测基准 评测集 研究原型 Stars 0 周增 +0

AI 诊断近视准确性的系统综述:完整数据、提取流程、代码与稿件。Systematic review of AI diagnostic accuracy for myopia detection: full data, extraction, code, and manuscript

seaofokhotskquakerism746/dabench-rlm-eval
Python · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

在 DABench 数据分析任务上对 DSPy RLM(Recursive Language Models)进行基准测试,使用自动评分实现基于代码的迭代评估Benchmark DSPy Recursive Language Models on DABench data analysis tasks with automated scoring for iterative code-based evaluation

agentragllm-infraevaluation