-
refusalbench
Reproducible, evergreen benchmark for LLM refusal on biological research prompts — 19 models, 141 prompts, 13,389 adjudicated trials
Python ★ 4 1mo agoExplain → -
VCBench
VCBench: capability-stratified benchmark for single-cell foundation models
Python ★ 2 1mo agoExplain → -
CardioSafe-benchmark
Curated data deposit for the CardioSafe cardiac ion channel benchmark: labels, Tanimoto-controlled splits (tan70 / tan60), and supplementary artifacts for hERG, Nav1.5, Cav1.2, and IKs.
Python ★ 1 2mo agoExplain →
No repos match these filters.