16-day current streak·16-day longest streak
arthi arumugam i build things. small teams. fast iteration. opinionated defaults. shipping > planning. taste > consensus. next thing soon. --- wrong-numbers — seventeen findings across ten llm eval, tracing…
arthi arumugam
i build things.
small teams. fast iteration. opinionated defaults.
shipping > planning. taste > consensus.
next thing soon.
---
wrong-numbers — seventeen findings across ten llm eval, tracing and cost libraries.
twelve are the same defect: a number comes out wrong and nothing raises. every one links to
an open pr with a test that fails on main.
https://github.com/arthi-arumugam-git/wrong-numbers
whatbroke — diff an agent's behaviour between two runs of the same task. tool calls,
arguments, costs, outputs.
https://github.com/arthi-arumugam-git/whatbroke
-
whatbroke ★ PINNED
Diff your AI agent's behavior between two runs. See exactly which tool calls, args, costs and outputs changed when you swap models or edit prompts.
TypeScript ★ 12 21h agoExplain → -
gnani ★ PINNED
gnani is a spreadsheet tool where every cell thinks. every cell is an AI agent.
TypeScript ★ 4 3mo agoExplain → -
costctl ★ PINNED
cost monitoring and optimization tool
Python ★ 2 3mo agoExplain → -
wrong-numbers
Twenty-five pull requests across fifteen LLM eval, tracing and cost libraries. Sixteen are the same defect: a number comes out wrong and nothing raises. Three merged upstream.
Python ★ 2 18h agoExplain → -
fiftyone ⑂
Refine high-quality datasets and visual AI models
★ 0 19h agoExplain → -
crewAI ⑂
Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks.
★ 0 21h agoExplain → -
mcp-analytics-python ⑂
Armature analytics wrapper SDK for Python MCP servers
★ 0 2d agoExplain → -
mcp-use ⑂
The fullstack MCP framework to develop MCP Apps for ChatGPT / Claude & MCP Servers for AI Agents.
★ 0 3d agoExplain → -
supervision ⑂
We write your reusable computer vision tools. 💜
★ 0 3d agoExplain → -
django-business-metrics ⑂
No description.
★ 0 4d agoExplain → -
haystack-core-integrations ⑂
Additional packages (components, document stores and the likes) to extend the capabilities of Haystack
★ 0 3d agoExplain → -
llama_index ⑂
LlamaIndex is the leading document agent and OCR platform
★ 0 4d agoExplain → -
agents ⑂
A framework for building realtime voice AI agents 🤖🎙️📹
★ 0 5d agoExplain → -
autoevals ⑂
AutoEvals is a tool for quickly and easily evaluating AI model outputs using best practices.
★ 0 5d agoExplain → -
inference ⑂
Turn any computer or edge device into a command center for your computer vision projects.
★ 0 21h agoExplain → -
pipecat ⑂
Open Source framework for voice and multimodal conversational AI
★ 0 5d agoExplain → -
verifiers ⑂
Our library for RL environments + evals
★ 0 7d agoExplain → -
inspect_evals ⑂
Collection of evals for Inspect AI
★ 0 4d agoExplain → -
logfire ⑂
AI observability platform for production LLM and agent systems.
★ 0 9d agoExplain → -
genai-prices ⑂
Calculate prices for calling LLM inference APIs.
★ 0 1d agoExplain → -
respan ⑂
No description.
★ 0 9d agoExplain → -
arthi-arumugam-git
profile
★ 0 9d agoExplain → -
helicone ⑂
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
★ 0 10d agoExplain → -
litellm ⑂
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthropic, OpenAI, VertexAI, vLLM, Nvidia NIM]
★ 0 21h agoExplain → -
deepeval ⑂
The LLM Evaluation Framework
★ 0 4d agoExplain → -
letta ⑂
Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.
★ 0 10d agoExplain → -
okareo-python-sdk ⑂
Python library for interacting with Okareo Cloud APIs
★ 0 10d agoExplain → -
phoenix ⑂
AI Observability & Evaluation
★ 0 10d agoExplain → -
judgeval ⑂
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
★ 0 21h agoExplain → -
openllmetry ⑂
Open-source observability for your GenAI or LLM application, based on OpenTelemetry
★ 0 11d agoExplain → -
vellum-python-sdks ⑂
Python SDK for Vellum API
★ 0 11d agoExplain → -
langfuse-python ⑂
🪢 Langfuse Python SDK - Instrument your LLM app with decorators or low-level SDK and get detailed tracing/observability. Works with any LLM or framework
★ 0 21h agoExplain → -
cohere-python ⑂
Python Library for Accessing the Cohere API
★ 0 11d agoExplain → -
ossinsight ⑂
Analysis, Comparison, Trends, Rankings of Open Source Software, you can also get insight from more than 10 billion with natural language (powered by LLM). Follow us on Twitter: https://twitter.com/ossinsight
★ 0 13d agoExplain → -
opentelemetry.io ⑂
The OpenTelemetry website and documentation
★ 0 8d agoExplain → -
ollama ⑂
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
★ 0 14d agoExplain →
No repos match these filters.
More creators on gitmyhub
tiangolo kennethreitz StephenGrider benawad jonasschmedtmann