AIAny

Category

Explore by categories

Hugging Face
AI Model2026

A Qwen-3.6 27B model variant optimized for DFlash (speculative decoding) to reduce generation latency and increase throughput. Focuses on faster inference on serving stacks and is suitable for text-generation endpoints where lower latency and resource efficiency matter.

Hugging Face
AI Model2026

A GGUF-format preview checkpoint derived from Qwen3.6-27B — a multimodal, image-text-to-text reasoning model fine-tuned for more structured reasoning and consistent answer style; packaged for local inference and compatible with engines like vLLM/SGLang/llama.cpp.

Hugging Face
AI Model2026

High-resolution vision transformers pretrained on one billion human images for human-centric tasks such as pose estimation, body-part segmentation, surface-normal and pointmap prediction. Provides multiple backbone sizes and task-specific checkpoints; released under the Sapiens2 license.

Hugging Face
AI Model2026

A 33B Mixture-of-Experts text-to-text model optimized for local, long-context agentic coding—3B activated params per token, 131k token window, mixed sliding-window and global attention, FP8 KV cache, Apache-2.0 license.

Hugging Face
AI Model2026

Provides a lightweight assistant (draft) model for Gemma 4 E4B used in speculative-decoding pipelines — it predicts token drafts that the target model verifies in parallel, enabling up to ~2× decoding speedups while preserving identical final outputs. Useful for low-latency, multimodal assistant and on-device scenarios.

Hugging Face
AI Video2026

Performs task-aware generative video restoration and editing in latent video space — restoration, super-resolution, watermark and subtitle removal — adapting LTX‑2.3 with IC‑Edit/IC‑LoRA adapters to prioritize temporal consistency and occlusion-aware reconstruction.

Hugging Face
AI Model2026

A lightweight 'drafter' assistant for Gemma 4 31B that generates speculative token drafts to enable up-to-2× decoding speedups while preserving final output quality; compatible with Hugging Face Transformers and any-to-any pipelines.

Hugging Face
AI Model2026

Acts as the assistant (drafter) checkpoint for Gemma 4 26B A4B on Hugging Face, used in Speculative Decoding to pre-draft tokens and speed up generation. Designed for long-context, multimodal workflows where lower latency and on-device or edge inference matter.

Hugging Face
AI Model2026

Unifies video, audio, image and text understanding for enterprise Q&A, summarization, transcription and document intelligence. The NVFP4 quantized variant reduces footprint to ~20.9GB for more efficient single‑GPU deployment and is tuned for NVIDIA runtimes (vLLM, TensorRT).

Hugging Face
AI Model2026

Provides unquantized BF16 weights of Qwen3.6-27B with the base model's MTP head grafted in for high-fidelity, uncensored text (and multimodal) generation. Includes deployment guidance and hardware-tuned variants for A100/H100 and Blackwell-class GPUs.

Hugging Face
AI Model2026

Generates anime-style images from natural-language prompts with a full fine-tune family built on Z-Image Base — available as Base, 8-step and 4-step distillations, plus AIO and GGUF variants for 8GB/low-VRAM workflows (BF16/FP8 formats).

Hugging Face
AI Model2026

Provides multiple GGUF-quantized exports of Carnice V2 (a merged BF16 SFT of Qwen3.6-27B) optimized for llama.cpp and Hermes-style agent traces, with quant tiers targeted at 16–24GB local GPUs and agentic inference.