Members
-
omnisafe ★ PINNED
JMLR: OmniSafe is an infrastructural framework for accelerating SafeRL research.
Python ★ 1.1k 1y agoExplain → -
safety-gymnasium ★ PINNED
NeurIPS 2023: Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
Python ★ 577 2d agoExplain → -
safe-rlhf ★ PINNED
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Python ★ 1.6k 8mo agoExplain → -
Safe-Policy-Optimization ★ PINNED
NeurIPS 2023: Safe Policy Optimization: A benchmark repository for safe reinforcement learning algorithms
Python ★ 415 2y agoExplain → -
align-anything
Align Anything: Training All-modality Model with Feedback
Python ★ 4.7k 7mo agoExplain → -
aligner
[NeurIPS 2024 Oral] Aligner: Efficient Alignment by Learning to Correct
Python ★ 194 1y agoExplain → -
VLA-Arena
VLA-Arena is an open-source benchmark for systematic evaluation of Vision-Language-Action (VLA) models.
Python ★ 191 18d agoExplain → -
beavertails
BeaverTails is a collection of datasets designed to facilitate research on safety alignment in large language models (LLMs).
Makefile ★ 182 2y agoExplain → -
SafeVLA
[NeurIPS 2025 Spotlight] Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning.
Python ★ 153 3mo agoExplain → -
AlignmentSurvey
AI Alignment: A Comprehensive Survey
★ 137 2y agoExplain → -
SafeDreamer
ICLR 2024: SafeDreamer: Safe Reinforcement Learning with World Models
Python ★ 105 2y agoExplain → -
ProAgent
AAAI24(Oral) ProAgent: Building Proactive Cooperative Agents with Large Language Models
JavaScript ★ 104 1y agoExplain → -
llms-resist-alignment
[ACL2025 Best Paper] Language Models Resist Alignment
Python ★ 51 1y agoExplain → -
safe-sora
SafeSora is a human preference dataset designed to support safety alignment research in the text-to-video generation field, aiming to enhance the helpfulness and harmlessness of Large Vision Models (LVMs).
Python ★ 35 1y agoExplain → -
ReDMan
ReDMan is an open-source simulation platform that provides a standardized implementation of safe RL algorithms for Reliable Dexterous Manipulation.
Python ★ 29 3y agoExplain → -
ProgressGym
Alignment with a millennium of moral progress. Spotlight@NeurIPS 2024 Track on Datasets and Benchmarks.
Python ★ 25 1y agoExplain → -
eval-anything
No description.
Python ★ 22 1y agoExplain → -
SAE-V
[ICML 2025 Poster] SAE-V: Interpreting Multimodal Models for Enhanced Alignment
★ 17 1y agoExplain → -
SAELens-V
No description.
Python ★ 11 1y agoExplain → -
TransformerLens-V
No description.
Python ★ 7 1y agoExplain → -
s1-m ⑂
S1-M: Simple Test-time Scaling in Multimodal Reasoning
Python ★ 3 1y agoExplain → -
Beaver-zh-hk
No description.
Python ★ 1 1y agoExplain → -
MM-DeceptionBench
No description.
★ 0 10mo agoExplain → -
.github
No description.
★ 0 1y agoExplain → -
Aligner2024.github.io ⑂
No description.
HTML ★ 0 1y agoExplain →
No repos match these filters.
More creators on gitmyhub
diego3g iamshaunjp WebDevSimplified paulirish iam-veeramalla