It split off from `verl` to give diffusion, video, and omni-modality models an RL post-training framework that doesn't treat them like chatbots.
ML Frameworks
underdogs · picking up speedTinyEngram open-sources experiments showing that DeepSeek's Engram architecture can inject domain knowledge into Qwen and Stable Diffusion more efficiently than LoRA, with less catastrophic forgetting.
A grab-bag node pack whose Set/Get rewrite might finally tame your worst workflow tangles.
It turns Claude, Codex, or OpenCode into hyperparameter optimizers that actually read your parameter docs before proposing the next trial.
A production-hardened fork of slime that keeps massive MoE models from collapsing by obsessing over bit-wise alignment between rollout and training.
OpenMed packages clinical entity extraction and HIPAA-grade de-identification into models small enough for Apple Silicon and impatient DevOps teams.
LitGPT re-implements 20+ models from scratch so you can actually read the code.
It exists to train gradient-boosted trees faster and with less memory than rivals, scaling from one machine to distributed clusters.
Curated technical deep-dives covering everything from NVLink signal integrity to Kubernetes GPU scheduling and Huawei NPU porting.
Hugging Face's PEFT library makes parameter-efficient fine-tuning feel like cheating—train 0.1% of weights, keep 99% of the performance.
NeMo shed its multimodal skin to focus on ASR, TTS, and speech LLMs—just as the field gets interesting.
It wraps the fragmented ecosystem of diffusion training scripts into one configurable CLI and web UI.
A from-scratch vLLM reimplementation in ~1,200 lines of Python that edges out the original on a laptop GPU.
LeWorldModel cuts the standard JEPA training recipe from six loss hyperparameters down to two, letting a 15M-parameter model learn a latent physics space directly from raw pixels in a few hours on a single GPU.
It corrals dozens of optical flow architectures into one PyTorch Lightning framework so you can train and benchmark them without maintaining forty separate codebases.
It corrals roughly two dozen deep time-series architectures into one evaluation harness for forecasting, imputation, anomaly detection, and classification.
It exists to provide the stream-mining community with an extensible Java benchmark suite for real-time machine learning.
It wraps scikit-learn classifiers in a betting framework to backtest strategies and hunt for value bets across historical odds.
RLinf exists because fine-tuning policies for robots and agents still requires rewriting your training stack for every new simulator, world model, or hardware rig.
It exists because aligning a 70B model with RLHF shouldn't require building your own distributed inference stack.








