Repositories · organized/repo_cards

仓库/Skill 库

13 个 · 评测基准 · 框架

排序 Stars 周增
MadsLorentzen/ai-job-search
TypeScript · 2026-08-10 评测基准 框架 生产可用 Stars 31142 周增 +385

在你机器上运行的求职工具。基于 Claude Code 构建的 AI 求职框架:评估职位、定制简历、撰写求职信、准备面试。Fork 它并拥有它。The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep interviews. Fork it and own it.

agentengineering
tensorzero/tensorzero
Rust · 2026-06-11 评测基准 框架 生产可用 Stars 11724 周增 +0

TensorZero 是一个开源 LLMOps 平台,统一了 LLM gateway、可观测性、评估、优化与实验TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.

evaluationengineeringllm-infra
evidentlyai/evidently
Jupyter Notebook · 2026-08-05 评测基准 框架 生产可用 Stars 7797 周增 +14

Ev​​idently 是一个开源的 ML 和 LLM 可观测性框架,评估、测试和监控任何 AI 驱动的系统或数据 pipeline,覆盖从表格数据到 Gen AI 场景,提供 100+ 指标。Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

engineeringllm-infra
Helicone/helicone
TypeScript · 2026-07-25 评测基准 框架 生产可用 Stars 6051 周增 +7

🧊 开源 LLM 可观测性平台,一行代码即可实现监控、评估与实验。YC W23 🍓🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

agentevaluationllm-infra
Agenta-AI/agenta
TypeScript · 2026-07-18 评测基准 框架 研究原型 Stars 4303 周增 +21

开源 LLMOps 平台:集成 prompt playground、prompt 管理、LLM 评估和 LLM 可观测性。The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

agentragevaluationllm-infra
tractorjuice/arc-kit
JavaScript · 2026-08-08 评测基准 框架 研究原型 Stars 2133 周增 +28

企业架构治理框架——为AI编程助手提供战略、架构、交付与保障The Enterprise Architecture Governance Harness — strategy, architecture, delivery, and assurance using AI coding assistants

agent
Merck/BioPhi
Python · 2025-05-13 评测基准 框架 实验 Stars 259 周增 +0

BioPhi 是开源抗体设计平台,提供自动化抗体人源化方法(Sapiens)、人源性评估(OASis)及计算机辅助抗体序列设计界面BioPhi is an open-source antibody design platform. It features methods for automated antibody humanization (Sapiens), humanness evaluation (OASis) and an interface for computer-assisted antibody sequence design.

evaluation
ELM-Research/ECG-Language-Models
Python · 2026-08-11 评测基准 框架 实验 Stars 16 周增 +0

面向 ECG-语言模型(ELM)的研究型训练与评估框架A research-oriented training and evaluation framework for ECG-Language Models (ELMs)

multimodalevaluationllm-infra
zjunlp/Mechanist
Python · 2026-08-12 评测基准 框架 研究原型 Stars 13 周增 +0

AI 系统作为理解智能机制的科学仪器AI Systems as Scientific Instruments for Understanding the Mechanisms of Intelligence

agentllm-infra
sandesh20lamichhane/when-should-ids-adapt
Jupyter Notebook · 2026-08-19 评测基准 框架 实验 Stars 0 周增 +0

一个用于研究入侵检测系统 (IDS) 在时间分布漂移下何时应进行适应的可复现研究框架。该项目评估了一个分阶段决策流水线,结合漂移筛查、新颖性检测与成本感知的适应触发器,以判断重训练是否必要且有益。A reproducible research framework for investigating when an Intrusion Detection System (IDS) should adapt under temporal distribution shift. The project evaluates a staged decision pipeline combining drift screening, novelty detection, and cost-aware adaptation triggers to determine whether retraining is necessary and beneficial.

engineering
ranaliwaa369/NOI-Research
Python · 2026-08-23 评测基准 框架 研究原型 Stars 0 周增 +0

面向 Neuro-Olfactive Intelligence 的专有可复现研究框架Proprietary reproducible research framework for Neuro-Olfactive Intelligence

krish3006b/Mental-health-impact-of-AI-Companion-Chatbots-
未知语言 · 2026-08-18 评测基准 框架 实验 Stars 0 周增 +0

PRISMA 指导的混合方法系统综述,跨 4 项实证研究评估 AI 陪伴的社会心理、关系及长期心理健康影响。包含定量与定性 meta 综合、跨方法三角验证与理论框架映射。A PRISMA guided mixed methods systematic review evaluating the psychosocial, relational and long term mental health impacts of AI companions across 4 empirical studies. Features quantitative and qualitative meta synthesis, cross methodological triangulation and theoretical framework mapping

K-ALOHA/cdi-policy-verification
Python · 2026-08-20 评测基准 框架 研究原型 Stars 0 周增 +0

基于反事实决策影响与验证效用的效用引导型政策验证可复现研究框架A reproducible research framework for utility guided policy verification using counterfactual decision impact and verification utility.