⚡
building at the intersection of systems and intelligence
I build RL training environments and the eval tooling that keeps them honest — verifier auditors, benchmark harnesses, and agent infrastructure. Also ship full-
Highlights
Pinned Loading
-
Reward-Hackability-Auditor--CLI---Claude-Skill-
Reward-Hackability-Auditor--CLI---Claude-Skill- Public🐀 Fuzz your verifier before an RL agent does. Static + dynamic LLM security auditor to detect reward-hacking in RL post-training environments (OpenEnv, verifiers-spec, Gymnasium).
Python 5
-
RL-for-Universal-Compiler-Optimization
RL-for-Universal-Compiler-Optimization PublicOffline Reinforcement Learning for Universal Compiler Optimization across CPUs and GPUs using LLVM IR, MLIR, Graph Neural Networks, and hardware-aware reward modeling.
Python 2
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

