PyTorch object detector built for shipping: train on your own data, then export to ONNX, CoreML, TFLite, or TensorRT with one command. Comes in five sizes (n/s/m/l/x) and adds instance-segmentation and classification heads beyond bounding-box detection.
Library for benchmarking, developing, and deploying deep-learning visual anomaly-detection models — includes ready-to-use model implementations (PatchCore, DINO-based), experiment/HPO tooling, OpenVINO export for edge inference, and a low-code Studio for deployment.
Offline desktop OCR for Windows and Linux that extracts text from screenshots, image batches, and scanned PDFs without requiring a network connection. Bundles multilingual offline engines (PaddleOCR / RapidOCR), supports ignore-regions, searchable PDF output, CLI and HTTP interfaces for automation and integration.
Enlarges and enhances low-resolution images using AI models (Real-ESRGAN) through a cross-platform desktop app. Runs on a local NCNN/Vulkan backend (requires a Vulkan-compatible GPU), offers an Electron GUI plus a CLI backend (upscayl-ncnn), and supports custom models for different image types.
Runs a self-hosted web server and React-based UI to generate, edit, and manage images using Stable Diffusion, SDXL and other foundation models. Includes a model manager, unified canvas, node-based workflows, gallery, and support for ckpt, diffusers and some gguf checkpoints.
Browser-based control panel for running Stable Diffusion locally, built on Gradio. Bundles txt2img, img2img, inpainting, outpainting, and upscalers (ESRGAN, GFPGAN, CodeFormer), plus an extension ecosystem and support for NVIDIA, AMD, and Intel GPUs.
Turns text prompts into images through latent diffusion, from local-ready releases to professional SD 3.5 models. Its impact comes from deployability: self-hosting, API access, and community tooling made image generation broadly hackable.
Unifies successive YOLO generations — YOLOv8, YOLO11, YOLOv3 and newer — under one package and a single `YOLO` API spanning detection, segmentation, classification, pose, oriented boxes and tracking, plus one-line export to ONNX, TensorRT and CoreML.
Provides reusable computer-vision utilities for dataset loading/conversion, visualization/annotation of detections and segmentation, and connectors to popular detection frameworks—aimed at quick prototyping, dataset work, and visualization.
Web and desktop/mobile WebUI for generating, editing, captioning and processing images and videos with Stable Diffusion and many diffusion models. Key features include automatic model download, SDNQ on-the-fly quantization for VRAM savings, balanced CPU/GPU offload, multi-backend GPU support, and built-in captioning/tagging/upscaling workflows.
X-AnyLabeling is a powerful annotation tool integrated with an AI engine for fast and automatic labeling. Designed for multi-modal data engineers, it offers industrial-grade solutions for complex tasks. Supports images and videos, GPU acceleration, custom models, one-click inference for all task images, and import/export formats like COCO, VOC, YOLO. Handles classification, detection, segmentation, captioning, rotation, tracking, estimation, OCR, VQA, grounding, etc., with various annotation styles including polygons, rectangles, rotated boxes.