Members
-
OpenCUA ★ PINNED
[NeurIPS 2025 Spotlight] OpenCUA: Open Foundations for Computer-Use Agents
Python ★ 804 1mo agoExplain → -
OSWorld ★ PINNED
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Python ★ 3.0k 10h agoExplain → -
aguvis ★ PINNED
[ICML2025] Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Python ★ 389 1y agoExplain → -
OpenAgents ★ PINNED
[COLM 2024] OpenAgents: An Open Platform for Language Agents in the Wild
Python ★ 4.9k 1y agoExplain → -
Spider2 ★ PINNED
[ICLR 2025 Oral] Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
HTML ★ 848 5mo agoExplain → -
instructor-embedding ★ PINNED
[ACL 2023] One Embedder, Any Task: Instruction-Finetuned Text Embeddings
Python ★ 2.0k 1y agoExplain → -
UnifiedSKG
[EMNLP 2022] Unifying and multi-tasking structured knowledge grounding with language models
Python ★ 566 2y agoExplain → -
xlang-paper-reading
Paper collection on building and evaluating language model agents via executable language grounding
★ 364 2y agoExplain → -
Binder
[ICLR 2023] Code for the paper "Binding Language Models in Symbolic Languages"
Python ★ 326 2y agoExplain → -
DS-1000
[ICML 2023] Data and code release for the paper "DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation".
Python ★ 276 1y agoExplain → -
text2reward
[ICLR 2024 Spotlight] Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Jupyter Notebook ★ 210 1y agoExplain → -
BRIGHT
[ICLR 2025] BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
Python ★ 206 10mo agoExplain → -
OSWorld-V2
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Python ★ 200 5h agoExplain → -
CUA-Gym
Scalable pipeline for synthesizing verifiable RLVR training data for computer-use agents
Python ★ 180 1mo agoExplain → -
OSWorld-G
[NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis
TypeScript ★ 172 1mo agoExplain → -
Spider2-V
[NeurIPS 2024] Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
Jupyter Notebook ★ 153 1y agoExplain → -
icl-selective-annotation
[ICLR 2023] Code for our paper "Selective Annotation Makes Language Models Better Few-Shot Learners"
Python ★ 109 3y agoExplain → -
batch-prompting
[EMNLP 2023 Industry Track] A simple prompting approach that enables the LLMs to run inference in batches.
Python ★ 76 2y agoExplain → -
FineVLA
Scalable annotation pipeline for action-aglined fine-grained instruciton for Visual-language-Action model
Python ★ 73 29d agoExplain → -
EVOR
No description.
Python ★ 70 1y agoExplain → -
computer-agent-arena
[ICLR 2026] Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents
HTML ★ 67 4mo agoExplain → -
CUA-Gym-Hub
CUA-Gym-Hub: mock web apps as reproducible RL training environments for computer-use agents
JavaScript ★ 66 14d agoExplain → -
AgentTrek
[ICLR2025 Spotlight] Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
Python ★ 60 1y agoExplain → -
VideoAgentTrek
The official repo of VideoAgentTrek
Python ★ 57 9mo agoExplain → -
AgentNetTool
This is the official code base of AgentNetTool in OpenCUA. Website: https://opencua.xlang.ai/
TypeScript ★ 52 10mo agoExplain → -
diagrams_toolkit
Source code for diagrams in the paper of NLPers from HKU.
Python ★ 5 4y agoExplain → -
osworld_image
No description.
Shell ★ 3 29d agoExplain → -
xlang-ai.github.io
The official website of xlang.ai
TypeScript ★ 3 7d agoExplain → -
osworld-server
Standalone guest-side server for OSWorld desktop environments
Python ★ 2 1mo agoExplain → -
verl ⑂
veRL: Volcano Engine Reinforcement Learning for LLM
★ 1 1y agoExplain → -
.github
No description.
★ 1 2y agoExplain → -
Pai-Megatron-Patch ⑂
The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.
★ 0 1y agoExplain →
No repos match these filters.