Repositories · organized/repo_cards

仓库/Skill 库

6 个 · 评测基准 · 框架 · AI 核心

排序 Stars 周增
tensorzero/tensorzero
Rust · 2026-06-11 评测基准 框架 生产可用 Stars 11724 周增 +0

TensorZero 是一个开源 LLMOps 平台,统一了 LLM gateway、可观测性、评估、优化与实验TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.

evaluationengineeringllm-infra
evidentlyai/evidently
Jupyter Notebook · 2026-08-05 评测基准 框架 生产可用 Stars 7797 周增 +14

Ev​​idently 是一个开源的 ML 和 LLM 可观测性框架,评估、测试和监控任何 AI 驱动的系统或数据 pipeline,覆盖从表格数据到 Gen AI 场景,提供 100+ 指标。Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.

engineeringllm-infra
Helicone/helicone
TypeScript · 2026-07-25 评测基准 框架 生产可用 Stars 6051 周增 +7

🧊 开源 LLM 可观测性平台,一行代码即可实现监控、评估与实验。YC W23 🍓🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓

agentevaluationllm-infra
Agenta-AI/agenta
TypeScript · 2026-07-18 评测基准 框架 研究原型 Stars 4303 周增 +21

开源 LLMOps 平台:集成 prompt playground、prompt 管理、LLM 评估和 LLM 可观测性。The open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

agentragevaluationllm-infra
ELM-Research/ECG-Language-Models
Python · 2026-08-11 评测基准 框架 实验 Stars 16 周增 +0

面向 ECG-语言模型(ELM)的研究型训练与评估框架A research-oriented training and evaluation framework for ECG-Language Models (ELMs)

multimodalevaluationllm-infra
zjunlp/Mechanist
Python · 2026-08-12 评测基准 框架 研究原型 Stars 13 周增 +0

AI 系统作为理解智能机制的科学仪器AI Systems as Scientific Instruments for Understanding the Mechanisms of Intelligence

agentllm-infra