Members
-
terminal-bench ★ PINNED
A benchmark for LLMs on complicated tasks in the terminal
Python ★ 2.5k 20d agoExplain → -
harbor ★ PINNED
Framework for evaluating and improving agents
Python ★ 3.7k 52m agoExplain → -
terminal-bench-2 ★ PINNED
No description.
Shell ★ 354 3mo agoExplain → -
frontier-bench ★ PINNED
Measuring and evolving with the frontier of agent work
Python ★ 428 10h agoExplain → -
terminal-bench-science ★ PINNED
Terminal-Bench Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal
Lean ★ 223 1d agoExplain → -
awesome-harbor ★ PINNED
A curated list of awesome Harbor ecosystem projects
★ 50 2mo agoExplain → -
harbor-cookbook
Realistic examples of building evals and optimizing agents with Harbor
Python ★ 155 3mo agoExplain → -
terminal-bench-2-1
Terminal-Bench 2.1
Shell ★ 57 10d agoExplain → -
harbor-datasets
No description.
★ 38 2mo agoExplain → -
harbor-index
A compact high-signal benchmark for evaluating frontier agents
Python ★ 21 3d agoExplain → -
terminal-bench-challenges
No description.
Shell ★ 20 1mo agoExplain → -
benchmark-template
No description.
Shell ★ 15 15d agoExplain → -
skills
Public agent skills catalog for Harbor
★ 11 2mo agoExplain → -
harbor-adapters-experiments
No description.
Python ★ 8 1mo agoExplain → -
terminal-bench-docs
No description.
TypeScript ★ 8 6d agoExplain → -
harbor-docs
No description.
MDX ★ 3 4mo agoExplain → -
docs
No description.
MDX ★ 0 1mo agoExplain →
No repos match these filters.
More creators on gitmyhub
tiangolo kennethreitz StephenGrider benawad jonasschmedtmann