Our latest batch of trending AI tools · September 28, 2026
Performs schema-driven, low-latency classification and structured decision-making over English text. Supports multi-head scoring, constrained joint decoding with confidence/feasibility metadata, span extraction, and local CPU/GPU deployment via the gliner2 runtime.
Parses digital and camera-captured documents into structured outputs (text, layout, tables, formulas, figures) using a lightweight (~1.2B) open-source vision-language model. Uses geometry-aware modeling, multi-node consensus pseudo-labeling, and content-structure decoupling to handle warped, photographed, and digital pages.
Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.
Synthetic, clinician-verified ChatML dataset of 2,194 doctor–patient encounters covering 2,194 unique human diseases; each JSONL record includes 20 structured fields, verified PubMed references, realistic vitals/labs, and is intended for RAG and model fine-tuning (not medical advice).
Uses diffusion-model forking moments as a proxy for perceptual distance to automatically generate pointwise reference-grounded labels, enabling annotation-free training of reference-based image quality assessment metrics.
Reduces self-attention complexity to O(N log N) by using a coarse-to-fine (pyramid) Top-K block selection with LogSumExp scoring, implemented with hardware-aware Triton kernels for fused routing and scoring—aimed at long-context LMs and retrieval tasks.
Learns joint predictive visual dynamics and action generation for generalist robot manipulation, translating future-relevant visual representations into actions. Integrates a Mixture-of-Transformers coupling a video expert and action expert, a frozen vision-language model for semantics, 4D distillation, and Causal Imprint; pretrained on a 20K+ hour heterogeneous corpus.
Provides a programming model and distributed runtime that preserves parent–child lineage and ordered variable-cardinality expansions for 1→M→1 dataflow pipelines, enabling completion-driven cross-input GPU batching and ordered gathers for foundation-model data preparation. Demonstrates multi-GPU scaling speedups and lower end-to-end time versus Ray Data and Daft.
Trains decoders and diffusion generators to be robust to random subsets of pretrained encoder layers by regularizing layer-fusion during training, narrowing the reconstruction–generation gap in representation autoencoders and improving ImageNet-256 PSNR and gFID.