← all repositories
PSRben/VisionHOPE

A backbone that rewrites its own learning rule mid-image

VisionHOPE is a PyTorch visual backbone built on self-referential nested learning, where what the model remembers and how it learns co-evolve while it processes a single image.

★537 stars Python Computer VisionML Frameworks
VisionHOPE
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

What it does VisionHOPE is a hierarchical visual backbone family (Tiny / Small / Base) and the official PyTorch implementation of an arXiv preprint, building on the authors’ earlier “Nested Learning” work. The repo covers the full research gauntlet: ImageNet-1K classification, COCO detection and instance segmentation with Mask R-CNN, and ADE20K semantic segmentation with UPerNet, with pretrained checkpoints for all three. It also exposes the core machinery — SRNL, VisionHOPEOperator, and VisionHOPEBlock — as standalone, shape-preserving modules for wiring into other architectures.

The interesting bit The operator runs five coupled memories that store content, generate key and value representations, and govern learning rate and retention — so the update rule itself adapts as visual context accumulates, not just the stored features. Recurrent memory designs usually risk blowing up; here a soft injection cap and a spectral clamp provide non-expansion guarantees for both token-wise and chunk-wise recurrences. Images are traversed via four directional scans over row- and column-aligned chunks, then spatially restored and fused channel-wise.

Key highlights

  • Five co-evolving memories per operator: content, keys/values, learning rate, and retention all update together within one forward pass.
  • Stability is engineered in, not hoped for: the injection cap and spectral clamp give non-expansion guarantees for the memory recurrences.
  • Standalone SRNL (sequence in, sequence out) plus image-shaped operator and block modules, all autograd-friendly, for reuse elsewhere.
  • Pretrained weights for ImageNet-1K, COCO, and ADE20K on GitHub Releases and Hugging Face, with a Baidu mirror.
  • A fast-inference mode fuses parameters in place for speed and leaves the training weights untouched.

Caveats

  • CUDA-only: the core modules require NVIDIA GPUs; there is no CPU path.
  • The stack is pinned hard — PyTorch 2.1.0, Python 3.10 (3.11 only for classification), and MMCV 2.1.0 for COCO/ADE20K in separate per-task environments. Research-grade hygiene, not a low-maintenance dependency.
  • CUDA extensions compile on first use, and parameters must stay FP32 with autocast handling mixed precision.

Verdict Worth a look if you track post-attention, post-SSM backbones or want an unusual sequence mixer to experiment with — the standalone modules make that part easy. Pass if you need CPU inference or a production-grade dependency story.

Frequently asked

What is PSRben/VisionHOPE?
VisionHOPE is a PyTorch visual backbone built on self-referential nested learning, where what the model remembers and how it learns co-evolve while it processes a single image.
Is VisionHOPE open source?
Yes — PSRben/VisionHOPE is open source, released under the MIT license.
What language is VisionHOPE written in?
PSRben/VisionHOPE is primarily written in Python.
How popular is VisionHOPE?
PSRben/VisionHOPE has 537 stars on GitHub.
Where can I find VisionHOPE?
PSRben/VisionHOPE is on GitHub at https://github.com/PSRben/VisionHOPE.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.