Repositories · organized/repo_cards

仓库/Skill 库

131 个

排序 Stars 周增
pabloalvarez99/production-rag
Python · 2026-08-11 RAG 检索增强 应用 实验 Stars 0 周增 +0

混合 RAG 系统,具备 reranking、基于引用的 grounded citations,以及可选 provider-backed 执行的可确定离线评估分级。A hybrid RAG system with reranking, grounded citations, and deterministic offline evaluation tiers with optional provider-backed execution.

ragevaluationllm-infra
pabloalvarez99/agentic-rag-research
Python · 2026-08-14 Agent 智能体 应用 实验 Stars 0 周增 +0

有界研究 Agent:plan → retrieve → critique,具备步骤预算、执行轨迹、引用追溯与自由路径 UI,无需 API key。Bounded research agent: plan → retrieve → critique, step budgets, traces, citations, free-path UI. No API key.

agentragevaluationllm-infra
nataliia-buchok-da/ab-test-statistical-analysis
Jupyter Notebook · 2026-08-26 评测基准 应用 实验 Stars 0 周增 +0

A/B 测试结果分析,涵盖整体指标评估与精细用户分群。项目结合了变更对结账漏斗各阶段影响的研究、统计假设检验以及不同群体的关键指标评估。Analysis of A/B testing results, covering both an overall assessment of metrics and detailed user segmentation. The project combines research into the impact of changes on the stages of the checkout funnel, testing of statistical hypotheses, and evaluation of key metrics across different groups

evaluation
naiem7788/FraudBench-Hybrid-A-Reproducible-Framework-for-Temporal-Bank-Fraud-Detection
Jupyter Notebook · 2026-08-20 评测基准 数据集 实验 Stars 0 周增 +0

FraudBench-Hybrid 是一个面向时序银行欺诈检测的可复现研究框架,仓库包含数据集配置、防数据泄露的 Heavy Hybrid 集成实现、实验评估、结果、可视化分析及未来研究扩展资料FraudBench-Hybrid is a reproducible research-based framework developed for temporal bank fraud detection. This repository will include dataset configuration, leakage-controlled Heavy Hybrid ensemble implementation, experimental evaluation, results, visual analysis, and material for future research extensions.

evaluation
moisestech/agentic-evidence-pipeline
TypeScript · 2026-08-12 Agent 智能体 应用 实验 Stars 0 周增 +0

面向有据可循的 Agent 工作流的 TypeScript 参考实现,支持混合检索、人工审核与可回放的审计追踪。TypeScript reference implementation for grounded agent workflows with hybrid retrieval, human review, and replayable audit trails.

agentragevaluationdatabase
meelone128/meelone128
未知语言 · 2026-08-16 Agent 智能体 应用 实验 Stars 0 周增 +0

涵盖 RAG、AI Agent、评估与商业智能的实战 AI 应用项目集。Practical AI application projects across RAG, AI agents, evaluation, and business intelligence.

agentragevaluation
ljestaciocerquin/ai-recist-systematic-review
未知语言 · 2026-08-25 评测基准 应用 实验 Stars 0 周增 +0

仓库包含文章与 Excel 文件,收录系统综述《人工智能方法在实体瘤疗效评价标准中的应用》中的 CLAIM 与 FUTURE-AI 评估。Repository containing articles and Excel files with CLAIM and FUTURE-AI assessments for the systematic review Artificial intelligence approaches to the Response Evaluation Criteria in Solid Tumors: a systematic review.

evaluation
larry-liyuanfan/climate-claim-verification-rag
Python · 2026-08-18 RAG 检索增强 应用 实验 Stars 0 周增 +0

多阶段气候声明检索与排序:BM25、稠密 ANN、融合、重排序与评估Multi-stage climate claim retrieval and ranking: BM25, dense ANN, fusion, reranking, and evaluation

ragevaluation
Kholeka98/DataScience
未知语言 · 2026-08-20 工程化 应用 实验 Stars 0 周增 +0

这个仓库展示了我的技能、项目,以及从数据工程转型为数据科学家角色的持续学习历程。This repository showcases my skills, projects, and continuous learning journey as I transition from Data Engineering to a Data Scientist role.

evaluationdatabasellm-infra
kabila5h/Shadow-API-scanner
Python · 2026-08-22 评测基准 评测集 实验 Stars 0 周增 +0

Shadow / Unmanaged API 的发现、分类、安全验证与可复现研究基准。Shadow / Unmanaged API discovery, classification, security validation, and reproducible research benchmarks

evaluationrisk
InvestmentMDideas/GLASS-Data-Extraction-Benchmark
Python · 2026-08-16 评测基准 评测集 实验 Stars 0 周增 +0

面向系统综述数据抽取、证据定位与偏倚风险评估的冻结式(frozen)防泄漏基准。Frozen, leakage-aware benchmark for systematic-review data extraction, evidence localization, and risk-of-bias support

evaluationrisk
friedpotato04/CUDA-L2
Cuda · 2026-08-22 LLM 基础设施 评测集 实验 Stars 0 周增 +0

🚀 使用强化学习优化半精度通用矩阵乘法(HGEMM)CUDA kernel,性能超越 cuBLAS 及其他基准。🚀 Optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels using reinforcement learning, surpassing cuBLAS and other benchmarks with superior performance.

evaluationllm-infra
fmadore/iwac-vocabulary
未知语言 · 2026-08-11 数据与向量库 实验 Stars 0 周增 +0

Islam West Africa Collection 的 RDF 词汇表:面向 Omeka S 的 AI 处理溯源与按模型键控的情感标注属性RDF vocabulary for the Islam West Africa Collection: AI processing provenance and model-keyed sentiment annotation properties for Omeka S

evaluationllm-infra
fmadore/IWAC-sentiment-analysis
Svelte · 2026-08-12 评测基准 应用 实验 Stars 0 周增 +0

对伊斯兰西非文献集(IWAC)语料库情感分析的交互式可视化,对比 ChatGPT、Gemini 与 Mistral,支持多语言与高级筛选。Interactive visualization of sentiment analysis on the Islam West Africa Collection (IWAC) corpus, comparing ChatGPT, Gemini, and Mistral with multilingual support and advanced filtering.

evaluationllm-infra
deepeshgupta12/agentic-rag-verified-citations
Python · 2026-08-14 RAG 检索增强 应用 实验 Stars 0 周增 +0

Agentic RAG:回答前对每条引用与原文进行核对,证据不足时弃答。具备自纠正检索循环、确定性引用锚定以及 prompt injection 防御能力。Agentic RAG that verifies every citation against source text before answering, and abstains when the evidence does not hold. Self-correcting retrieval loop, deterministic citation grounding, prompt-injection defence.

agentragevaluationllm-infra
DanceNitra/ramr
Python · 2026-08-12 RAG 检索增强 评测集 实验 Stars 0 周增 +0

RAMR —— 检索增强记忆可靠性:面向 Agentic-RAG / 记忆系统的抗污染合成基准(附方法与发现)。RAMR — Retrieval-Augmented Memory Reliability: a contamination-resistant synthetic benchmark for agentic-RAG / memory systems (findings + method)

agentragevaluationllm-infra
Chiagoziem2/Systematic-review-screening
Python · 2026-08-19 评测基准 评测集 实验 Stars 0 周增 +0

系统综述的主动学习摘要筛选。在全部 26 个 SYNERGY 数据集上基准测试:WSS@95 均值 64.1。Active-learning abstract screening for systematic reviews. Benchmarked across all 26 SYNERGY datasets: mean WSS@95 of 64.1.

ragevaluation
Bot87Ever/thunderbolt-ai
Python · 2026-08-21 RAG 检索增强 评测集 实验 Stars 0 周增 +0

本地 LLM 基准测试、RAG、语音交互与 AI 助手Local LLM benchmarking, RAG, voice interaction and AI assistant

ragevaluationllm-infra
avrsnramasamy/AI-DESIGN-BENCHMARK
HTML · 2026-08-20 评测基准 评测集 实验 Stars 0 周增 +0

🎨 利用 AI 实时生成高保真 UI 设计,对比多版本方案,并跨多个模型导出可直接使用的代码。🎨 Generate high-fidelity UI designs in real-time with AI, compare variations, and export ready-to-use code across multiple models.

agentevaluationllm-infra
Arithmetic-Power-Geometry/Endogenous-Inquiry-Computing
Python · 2026-08-23 评测基准 评测集 实验 Stars 0 周增 +0

将问题框架视为计算状态的可复现研究框架,包含理论、算法与基准,用于内生探究。A reproducible research framework for treating problem frames as computational states, with theory, algorithms, and benchmarks for endogenous inquiry.

evaluation
alfasmartagency-droid/rag-doctor
未知语言 · 2026-08-11 RAG 检索增强 应用 实验 Stars 0 周增 +0

通过分析 trace 诊断与调试 RAG 流水线,定位并修复影响答案质量与相关性的问题Diagnose and debug Retrieval Augmented Generation pipelines by analyzing traces to identify and fix issues affecting answer quality and relevance.

agentragevaluationengineering
Abdulrahman-Albeladi/research-toolkit
Python · 2026-08-22 评测基准 应用 实验 Stars 0 周增 +0

可复用的 Python 与 notebook 工作流,用于多语言科研数据准备、统计分析及 AI 文本检测器评估。Reusable Python and notebook workflows for multilingual research-data preparation, statistical analysis, and AI-text detector evaluation.

evaluation
0vertake/jetpacker
Kotlin · 2026-08-17 Agent 智能体 评测集 实验 Stars 0 周增 +0

面向 AI coding agents 的 token 预算 context pack,基于编译器解析的 Kotlin 结构(Analysis API/PSI)构建,并附带衡量其是否优于 chunk RAG 的 benchmark。Token-budgeted context packs for AI coding agents, built from compiler-resolved Kotlin structure (Analysis API/PSI) — with the benchmark that measures whether it beats chunk RAG

agentragevaluationllm-infra