研究库 开源仓库
Repositories · organized/repo_cards

仓库/Skill 库

49 个 · 评测基准 · AI 核心

排序 Stars 周增
wolfenix/llm-math-reasoning-analysis
HTML · 2026-10-01 评测基准 工具 实验 Stars 0 周增 +0

🔍 使用多种 prompt 技巧在多步数学问题上分析 Mistral-7B 模型的数学推理能力。🔍 Analyze the mathematical reasoning abilities of the Mistral-7B model using diverse prompting techniques on multi-step math problems.

evaluationllm-infra
synaptiai/flow-harness
TypeScript · 2026-08-20 评测基准 应用 实验 Stars 0 周增 +0

与供应商无关的 coding-agent 框架,具备确定性工作流图、持久化证据与 fail-closed 沙箱执行。Provider-neutral coding-agent harness with deterministic workflow graphs, durable evidence, and fail-closed sandboxed execution

agentllm-infra
sumit7194/Phronesis
HTML · 2026-08-16 评测基准 工具 实验 Stars 0 周增 +0

在小型 LLM 中植入认知美德的激活引导实验——三项发现与一个失败模式数据集(DOI 见 README)Activation-steering experiments on installing epistemic virtues in small LLMs — three findings + a failure-mode dataset (DOIs in README).

riskllm-infra
seaofokhotskquakerism746/dabench-rlm-eval
Python · 2026-08-28 评测基准 评测集 实验 Stars 0 周增 +0

在 DABench 数据分析任务上对 DSPy RLM(Recursive Language Models)进行基准测试,使用自动评分实现基于代码的迭代评估Benchmark DSPy Recursive Language Models on DABench data analysis tasks with automated scoring for iterative code-based evaluation

agentragllm-infraevaluation
sdwurg180507280211/metersphere
Java · 2026-08-19 评测基准 工具 生产可用 Stars 0 周增 +0

MeterSphere v2.10 个人开发分支 — 一站式开源持续测试平台。工作流引擎(Flowable 7) · 需求池(从0到1) · 微前端(qiankun→micro-app) · AI知识库 · 测试跟踪,含个人笔记与创意项目。

rag
Sai-Kartheek-Reddy/HateMirage-ICON2026
CSS · 2026-08-19 评测基准 评测集 研究原型 Stars 0 周增 +0

HateMirage——可解释的伪仇恨检测与多维推理,ICON 2026 共享任务。在 HateMirage 语料库(4,530 条带标注的伪仇恨评论)上进行目标识别 + 意图与隐含意义生成。HateMirage - Explainable Faux Hate Detection and Multi-Dimensional Reasoning - Shared Task @ ICON 2026. Target identification + Intent and Implication generation over the HateMirage corpus (4,530 annotated Faux Hate comments).

ragevaluationllm-infra
naturalmoods/clawclones
TypeScript · 2026-08-11 评测基准 评测集 实验 Stars 0 周增 +0

追踪并比较每一个严肃的 OpenClaw 替代方案——基于实测仓库数据与 AI 撰写的决策支持,且刻意保持独立。Track and compare every serious OpenClaw alternative — measured repo data, AI-written decision support, kept apart on purpose.

agent
kautum/msc-cyberattack-detection-llm
Jupyter Notebook · 2026-08-11 评测基准 工具 研究原型 Stars 0 周增 +0

硕士论文(KCL):在攻击者-机器留出划分下重新衡量 IoT 入侵检测,并测试基于 LLM 的报文预测是否能提升性能。移除泄露后,macro-F1 从 0.90 降至 0.60。MSc dissertation (KCL): re-measuring IoT intrusion detection under an attacker-machine holdout, and testing whether LLM-based packet prediction improves it. macro-F1 0.90 -> 0.60 once leakage is removed.

riskllm-infra
fmadore/IWAC-sentiment-analysis
TypeScript · 2026-10-02 评测基准 应用 实验 Stars 0 周增 +0

交互式仪表盘,对比五个 LLM(GPT-5.6 Luna、Mistral Small 4、DeepSeek v4 Flash、Gemma 4 31B 与 Qwen3.8 27B)对 12,349 篇西非法语区新闻文章的标注结果,包含评分者一致性、差异分析与盲法仲裁。法/英双语;此前的三轮三模型标注活动保留归档。Interactive dashboard comparing how five LLMs — GPT-5.6 Luna, Mistral Small 4, DeepSeek v4 Flash, Gemma 4 31B and Qwen3.8 27B — annotated 12,349 francophone West African press articles, with inter-rater agreement, discrepancy analysis and blind arbitration. Bilingual FR/EN; the earlier three-model campaign stays archived.

evaluationllm-infra
Dedebanded912/course-eligibility-hub
HTML · 2026-10-08 评测基准 应用 实验 Stars 0 周增 +0

即时评估课程资格,输出明确的通过/未通过结果以及定制化的入学测试。Instantly evaluate course eligibility with clear pass/fail results and tailored entry assessments.

multimodaldatabaseriskllm-infra
avrsnramasamy/AI-DESIGN-BENCHMARK
HTML · 2026-09-20 评测基准 评测集 实验 Stars 0 周增 +0

🎨 利用 AI 实时生成高保真 UI 设计,对比多版本方案,并跨多个模型导出可直接使用的代码。🎨 Generate high-fidelity UI designs in real-time with AI, compare variations, and export ready-to-use code across multiple models.

agentevaluationllm-infra
AnimeshShaw/GenIaC-SecBench
Python · 2026-08-31 评测基准 评测集 实验 Stars 0 周增 +0

基于人类基准的 LLM 生成 Infrastructure-as-Code 安全基准测试。100 场景 × 12 模型配置 = 1,196 个工件,由 Checkov/Trivy/KICS 扫描,并与 634 个人工编写模板对比。所有模型的漏洞密度均为人工的 3.2–3.9 倍。arXiv:2608.28021Human-anchored security benchmark for LLM-generated Infrastructure-as-Code. 100 scenarios × 12 model configs = 1,196 artifacts scanned by Checkov/Trivy/KICS, compared against 634 human-written templates. Every model: 3.2–3.9× human vulnerability density. arXiv:2608.28021

evaluationriskllm-infra
091635Aa/SemanticEcho-Data
Python · 2026-08-11 评测基准 库 实验 Stars 0 周增 +0

语义回响实证数据仓库:多模型对照 + 全流程 7 模式 2026 评测结果 + 论文/图表

llm-infra