AIAny
AI Video2026
Icon for item

MiniMax-H3 360° Orbit LoRA

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.

Introduction

Loopable 360° camera orbits are deceptively hard: standard reference-to-video models either freeze the scene or drift off the start frame, which prevents clean stitching. This LoRA teaches MiniMax‑H3 FL2VA the motion pattern of a true orbit while keeping every object motionless in world space, so the clip both looks geometrically consistent and closes on the original frame for seamless chaining.

What Sets It Apart
  • Learned, geometry-consistent orbit behavior: trained from rendered orbit sequences so the model moves the camera while preserving world positions and poses; this means parallax is the only apparent motion, enabling clean visual continuity when clips are concatenated.
  • FL2VA-first+last integration: by conditioning on the same image for both first and last keyframes, the LoRA enforces an exact start/end match — so orbits don’t drift and can be looped or stitched without visible seams.
  • Narrow, high-fidelity training domain: trained on 28 human-centric Gaussian-splat renders (768×768, 73 frames). That focused dataset yields strong performance on similar subjects (people, portrait-style scenes) but limits generalization to arbitrary scenes or aspect ratios.
  • Lightweight, Comfy/ai-toolkit friendly: LoRA rank 16 adapter weights intended for use with MiniMax‑H3 FL2VA extensions (tested with ostris/ai-toolkit and Comfy-style model keys), making it easy to drop into existing FL2VA pipelines.
Who It's For

Great fit if you need short, loopable 360° orbits from a single reference photo (for product shots, portrait loops, or stitched sequences) and you can work within square, human-centric inputs and MiniMax‑H3 FL2VA. Look elsewhere if your scenes are wide-aspect, non-human, require long durations, or must include complex scene dynamics — the LoRA was trained on a small, specialized dataset and may still exhibit subtle blinks or micro-movement. Recommended inference: identical first+last keyframe, 768×768, ~73 frames (≈3 s at 24fps), LoRA strength 1.0, guidance none, audio off.

Information

More Items

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.

Hugging Face
AI Audio2026

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.

AI Model2026

Explains how Jev turns input state into typed decisions and probabilities without generating text. Introduces parallel sampling and RLCD training, with workflow evaluations and caveats for software automation.