AIAny
AI Model2026
Icon for item

Jev-Omni

Converts multimodal inputs (text, image, audio, video) plus a question and options into calibrated probability distributions over choices. Built on Gemma 4 12B with a 30,000-question fine-tune, optimized for per-question decision classification and low-latency inference (Apache-2.0).

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Unexpected character `=` (U+003D) before name, expected a character that can start a name, such as a letter, `$`, or `_` More information: https://mdxjs.com/docs/troubleshooting-mdx

More Items

Hugging Face
AI Image2026

A distilled LoRA adapter for Qwen-Image-2.1 that runs text-to-image generation and instruction-driven image editing in a few transformer passes (shipped as a 6-step r256 LoRA). Samples with a fixed sigma schedule, no classifier-free guidance; non-commercial research license.

Hugging Face
AI Model2026

Performs unified parsing of digital and camera-captured documents (layout, text, tables, formulas) using a ~1.2B-parameter vision–language model. Key differences: geometry-aware modeling, curvature-guided sampling, and content-structure decoupled training to handle real-world deformations without separate dewarping.

Hugging Face
AI Audio2026

Performs speaker diarization (who spoke when) for live and recorded audio using an open-weight, 100M-parameter streaming-capable model that supports up to eight anonymous speaker channels, overlapping speech, chunked processing, and configurable latency for ASR integration.