3-day longest streak
👋 Hi there! I'm Zhouliang (郁昼亮), an PhD student at the Scalable Principles for Learning and Reasoning Lab (SphereLab) of the Chinese University of Hong Kong, in the Computer Science…
👋 Hi there! I'm Zhouliang (郁昼亮), an PhD student at the Scalable Principles for Learning and Reasoning Lab (SphereLab) of the Chinese University of Hong Kong, in the Computer Science & Engineering department, advised by Prof. Weiyang Liu, working on reinforcement learning for formal reasoning.
Previously, I spent a wonderful year at HKGAI, HKUST, as a PhD student advised by Prof. Yike Guo. Before that, I received my bachelor's degree from the Chinese University of Hong Kong, Shenzhen.
🎯 Research Focus
During the long-term future (maybe 2024-2027), I will be (almost) entirely focused on exploration-based reinforcement learning for formal mathematics reasoning via (agentic) large language models. Despite not being my major research focus, I am actively learning RL infrastructure to support Large Model Training.
🌍 Other broader interests include their applications in model-based embodied AI and scientific discovery via formal verification (e.g., Scientist AI, PhysLean, however, I have not yet published in this domain).
📧 Contact Me:[My Email]([email protected])
-
Chinese-Tiny-LLM ★ PINNED ⑂
No description.
★ 0 2y agoExplain → -
MAP-NEO ★ PINNED ⑂
No description.
★ 0 1y agoExplain → -
Kimina-Prover-Preview ★ PINNED ⑂
Technical report of Kimina-Prover Preview.
★ 0 1y agoExplain → -
cadRL
No description.
★ 5 5mo agoExplain → -
axolver ⑂
No description.
★ 0 3mo agoExplain → -
zhouliang-yu.github.io ⑂
No description.
HTML ★ 0 3mo agoExplain → -
tsp_rl
No description.
Python ★ 0 3mo agoExplain → -
autoresearch ⑂
AI agents running research on single-GPU nanochat training automatically
★ 0 4mo agoExplain → -
Skills ⑂
A project to improve skills of large language models
★ 0 5mo agoExplain → -
formal_veri
Codebase for CSCI course project
Python ★ 0 7mo agoExplain → -
BFS_FormalMATH
No description.
Python ★ 0 8mo agoExplain → -
zhouliang-yu
Config files for my GitHub profile.
★ 0 11mo agoExplain → -
minictx
No description.
★ 0 1y agoExplain → -
Awesome-System2-Reasoning-LLM ⑂
No description.
★ 0 1y agoExplain → -
mathlib4 ⑂
The math library of Lean 4
★ 0 1y agoExplain → -
textgrad ⑂
TextGrad: Automatic ''Differentiation'' via Text -- using large language models to backpropagate textual gradients.
★ 0 1y agoExplain → -
llm-verified-with-monte-carlo-tree-search ⑂
LLM verified with Monte Carlo Tree Search
★ 0 2y agoExplain → -
llama ⑂
Inference code for Llama models
★ 0 2y agoExplain → -
test_template
No description.
★ 0 2y agoExplain → -
bigaidream
No description.
★ 0 2y agoExplain → -
robotics-transformer ⑂
Octo is a transformer-based robot policy trained on a diverse mix of 800k robot trajectories.
★ 0 2y agoExplain → -
hyperdreambooth ⑂
No description.
★ 0 3y agoExplain → -
hyper-nn ⑂
Easy Hypernetworks in Pytorch and Jax
★ 0 3y agoExplain → -
Text-To-Video-Finetuning ⑂
Finetune ModelScope's Text To Video model using Diffusers 🧨
Python ★ 0 3y agoExplain → -
llm_agents ⑂
Build agents which are controlled by LLMs
★ 0 3y agoExplain → -
Tune-A-Video ⑂
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
★ 0 3y agoExplain → -
CVPR23_LFDM ⑂
The pytorch implementation of our CVPR 2023 paper "Conditional Image-to-Video Generation with Latent Flow Diffusion Models"
★ 0 3y agoExplain → -
codealpaca ⑂
No description.
★ 0 3y agoExplain → -
tool_alpaca ⑂
alpaca that learns to use the tools
★ 0 3y agoExplain → -
voltron-evaluation-teeter ⑂
Voltron Evaluation: Diverse Evaluation Tasks for Robotic Representation Learning
★ 0 3y agoExplain → -
pytorch_distributed ⑂
The test of different distributed-training methods on High-Flyer AIHPC
★ 0 3y agoExplain → -
MIPS-assembler
No description.
C++ ★ 0 4y agoExplain → -
EnvInteractiveLMPapers ⑂
No description.
★ 0 3y agoExplain → -
robotic-transformer-pytorch ⑂
Implementation of RT1 (Robotic Transformer) in Pytorch
★ 0 3y agoExplain → -
auto-cot ⑂
Official implementation for "Automatic Chain of Thought Prompting in Large Language Models" (stay tuned & more will be updated)
★ 0 3y agoExplain → -
pal ⑂
PaL: Program-Aided Language Models
★ 0 3y agoExplain → -
operatSystem
This repo contains my project code in csc3150
C ★ 0 4y agoExplain → -
csc3002
course code of csc3002
C++ ★ 0 5y agoExplain → -
CSC4005_2022Fall_Demo ⑂
No description.
★ 0 3y agoExplain → -
can-wikipedia-help-offline-rl ⑂
Official code for "Can Wikipedia Help Offline Reinforcement Learning?" by Machel Reid, Yutaro Yamada and Shixiang Shane Gu
★ 0 4y agoExplain → -
Parallel-Odd-Even-Transposition-Sort
This a parallel version implementations of odd even transposition sort.
C++ ★ 0 3y agoExplain → -
denoising-diffusion-pytorch ⑂
Implementation of Denoising Diffusion Probabilistic Model in Pytorch
★ 0 3y agoExplain → -
maximum_entropy_population_based_training ⑂
Maximum Entropy Population Based Training for Zero-Shot Human-AI Coordination
★ 0 4y agoExplain → -
RL-Benchmark ⑂
Implementation of Deep Reinforcement Learning Benchmark Algorithms, including DQN, Double DQN, Dueling DQN, Reinforce, Actor-Critic, A2C, A3C, etc.
★ 0 4y agoExplain → -
diffuser ⑂
Code for the paper "Planning with Diffusion for Flexible Behavior Synthesis"
★ 0 4y agoExplain → -
brax ⑂
Massively parallel rigidbody physics simulation on accelerator hardware.
★ 0 4y agoExplain → -
deep_RL ⑂
Assignments for Berkeley CS 285: Deep Reinforcement Learning (Fall 2021)
★ 0 4y agoExplain → -
ALU
This repo contains my implementation of ALU
Verilog ★ 0 4y agoExplain → -
decision-transformer ⑂
Official codebase for Decision Transformer: Reinforcement Learning via Sequence Modeling.
★ 0 4y agoExplain → -
Trajectory-Transformer ⑂
Code for "Transformer Networks for Trajectory Forecasting"
★ 0 5y agoExplain → -
Mips-Simulator
No description.
C++ ★ 0 4y agoExplain → -
pytorch-transformers ⑂
👾 A library of state-of-the-art pretrained models for Natural Language Processing (NLP)
★ 0 6y agoExplain → -
interactive-classification ⑂
No description.
★ 0 5y agoExplain → -
Grid-ML ⑂
Power grid optimization problem solvers
★ 0 4y agoExplain → -
DC3 ⑂
DC3: A Learning Method for Optimization with Hard Constraints
★ 0 4y agoExplain →
No repos match these filters.