Provides official pretrained VisionHOPE visual-backbone checkpoints for ImageNet classification, COCO object detection & instance segmentation, and ADE20K semantic segmentation. Includes hierarchical Tiny/Small/Base models with PyTorch-compatible downloadable weights.
Uses a multimodal model's own critiques as privileged context and applies on-policy self-distillation over diffusion sampling trajectories to internalize corrective guidance, improving text-to-image generation without an external teacher; shows measurable gains on GenEval and GenEval2.