2-day longest streak
-
ai-evaluation-tools
The comprehensive list of AI evaluation tools: 300+ open-source and commercial LLM evaluation frameworks, platforms, benchmarks, observability, red teaming, and guardrails — for LLMs, RAG, and AI agents.
Python ★ 2 16d agoExplain → -
ai-governance-tools
The comprehensive list of AI governance tools: 130+ platforms and resources for AI risk, compliance, model governance, audit, fairness, privacy, and responsible AI.
Python ★ 1 16d agoExplain → -
context-engineering-tools
The comprehensive list of context engineering tools: 160+ tools for LLM memory, prompts, context management, compression, caching, MCP, retrieval, and evaluation.
Python ★ 1 16d agoExplain → -
rag-retrieval-tools
The comprehensive list of RAG and retrieval tools: 200+ vector databases, search frameworks, embeddings, rerankers, parsers, pipelines, benchmarks, and evaluation tools.
Python ★ 1 16d agoExplain → -
ai-red-teaming-tools
The comprehensive list of AI red teaming tools: open-source and commercial frameworks, scanners, jailbreak tests, guardrails, benchmarks, and security resources.
Python ★ 0 16d agoExplain → -
ai-agent-frameworks
The comprehensive list of AI agent frameworks and orchestration tools: 175 SDKs, workflows, multi-agent systems, MCP tools, memory, deployment, and evaluation.
Python ★ 0 16d agoExplain → -
llm-observability-tools
The comprehensive list of LLM observability tools: 120 open-source and commercial platforms for AI tracing, monitoring, evaluation, analytics, cost, and security.
Python ★ 0 16d agoExplain → -
llm-fine-tuning-tools
The comprehensive list of LLM fine-tuning tools: 150+ frameworks and platforms for SFT, PEFT, LoRA, RLHF, preference optimization, data, quantization, and deployment.
Python ★ 0 16d agoExplain →
No repos match these filters.