A backbone that rewrites its own learning rule mid-image
VisionHOPE is a PyTorch visual backbone built on self-referential nested learning, where what the model remembers and how it learns co-evolve while it processes a single image.

What it does
VisionHOPE is a hierarchical visual backbone family (Tiny / Small / Base) and the official PyTorch implementation of an arXiv preprint, building on the authors’ earlier “Nested Learning” work. The repo covers the full research gauntlet: ImageNet-1K classification, COCO detection and instance segmentation with Mask R-CNN, and ADE20K semantic segmentation with UPerNet, with pretrained checkpoints for all three. It also exposes the core machinery — SRNL, VisionHOPEOperator, and VisionHOPEBlock — as standalone, shape-preserving modules for wiring into other architectures.
The interesting bit The operator runs five coupled memories that store content, generate key and value representations, and govern learning rate and retention — so the update rule itself adapts as visual context accumulates, not just the stored features. Recurrent memory designs usually risk blowing up; here a soft injection cap and a spectral clamp provide non-expansion guarantees for both token-wise and chunk-wise recurrences. Images are traversed via four directional scans over row- and column-aligned chunks, then spatially restored and fused channel-wise.
Key highlights
- Five co-evolving memories per operator: content, keys/values, learning rate, and retention all update together within one forward pass.
- Stability is engineered in, not hoped for: the injection cap and spectral clamp give non-expansion guarantees for the memory recurrences.
- Standalone
SRNL(sequence in, sequence out) plus image-shaped operator and block modules, all autograd-friendly, for reuse elsewhere. - Pretrained weights for ImageNet-1K, COCO, and ADE20K on GitHub Releases and Hugging Face, with a Baidu mirror.
- A fast-inference mode fuses parameters in place for speed and leaves the training weights untouched.
Caveats
- CUDA-only: the core modules require NVIDIA GPUs; there is no CPU path.
- The stack is pinned hard — PyTorch 2.1.0, Python 3.10 (3.11 only for classification), and MMCV 2.1.0 for COCO/ADE20K in separate per-task environments. Research-grade hygiene, not a low-maintenance dependency.
- CUDA extensions compile on first use, and parameters must stay FP32 with autocast handling mixed precision.
Verdict Worth a look if you track post-attention, post-SSM backbones or want an unusual sequence mixer to experiment with — the standalone modules make that part easy. Pass if you need CPU inference or a production-grade dependency story.
Frequently asked
- What is PSRben/VisionHOPE?
- VisionHOPE is a PyTorch visual backbone built on self-referential nested learning, where what the model remembers and how it learns co-evolve while it processes a single image.
- Is VisionHOPE open source?
- Yes — PSRben/VisionHOPE is open source, released under the MIT license.
- What language is VisionHOPE written in?
- PSRben/VisionHOPE is primarily written in Python.
- How popular is VisionHOPE?
- PSRben/VisionHOPE has 537 stars on GitHub.
- Where can I find VisionHOPE?
- PSRben/VisionHOPE is on GitHub at https://github.com/PSRben/VisionHOPE.