Provides a Gymnasium-style API and tooling to create, deploy, and interact with isolated execution environments for agentic RL training. Includes async/sync clients, a web interface, CLI, Docker-based deployment, and Hugging Face Spaces integration.
Runs text-to-speech with instant voice cloning fully on-device, from phones to GPUs. Built on small LLM backbones (120M-360M params) plus a 50Hz neural codec; clones a voice from ~3 seconds of audio across English, Spanish, German, and French.
Turns clinical text into structured, de-identified clinical signals—entity extraction and PII de-identification—that run entirely on local hardware. Provides 1,000+ specialized medical NER models, multilingual support, Apple MLX acceleration, and Apache‑2.0 licensing.
Provides 99,870 system/user/assistant chat triples for defensive cybersecurity instruction‑tuning, with built‑in refusal patterns and mapping to OWASP, MITRE ATT&CK, NIST, and CIS standards; Apache‑2.0 licensed.
Autonomously performs end-to-end data science tasks — from cleaning and exploration to modeling, visualization, and analyst-grade reports — via an agentic LLM. Open-source model, code, datasets and demos; supports vLLM deployment, Jupyter/CLI/Web UIs, and OpenAI-style APIs.
Provides a curated collection of hands-on tutorials, workflows and auxiliary files for training and using generative-model tooling (Stable Diffusion, Flux, WAN). Key items include a WAN 2.1 LoRA training tutorial and an articles collection covering DreamBooth, LoRA, LyCORIS and SDXL.
Converts images and PDFs into structured Markdown, HTML, or JSON while preserving layout, handling tables, math, handwriting, charts, and chemistry diagrams across 90+ languages. Runs locally via HuggingFace or against a vLLM server.
Provides mined hard negatives and relevance scores for 1.88M queries across seven retrieval datasets, enabling contrastive fine-tuning and nv-retrieve filtering; includes full 2048 mined negatives per query, paired query/document splits, and parquet-formatted files for large-scale training.
Automates multi-step web tasks by perceiving webpages as pixels and issuing low-level mouse, keyboard and scroll actions. A 7B-parameter multimodal agent trained on 145K synthetic trajectories (FaraGen), designed for on-device deployment and efficient task completion (~16 steps/task).
Drives an LLM-powered agent to autonomously research, write, and ship ML code by accessing Hugging Face docs, datasets, repos, and cloud compute. Provides interactive CLI and headless modes, approval gates, tool routing, and integrations for HF, GitHub, and Anthropic models.
Contains training, evaluation, and deployment code plus checkpoints for humanoid whole-body controllers (Decoupled WBC and GEAR‑SONIC). Includes C++ inference, VR teleoperation, data pipelines (Bones‑SEED) and Hugging Face checkpoints for research-to-robot workflows.
Aggregates SEC EDGAR filings into raw files, parsed plaintext, and rich filing metadata for LLM training and retrieval. Includes ~8.05M filings (~590 GB, ~43B tokens), per-filing token counts, and parsed outputs; Apache-2.0.