Converts PDFs, Office files, HTML, images and audio into one structured DoclingDocument, with deep PDF layout, reading order, table-structure and formula recognition, OCR, and native LangChain/LlamaIndex/Haystack integrations for RAG pipelines.
Lets Python developers write tile-based parallel kernels for NVIDIA GPUs, generating CUDA Tile IR while staying close to Python syntax for custom GPU operations.
A compact domain-specific language for writing high-performance GPU/CPU kernels (GEMM, FlashAttention, sparse kernels) with Python-like syntax. It provides tiling/pipelining primitives, a TVM-based compiler and multiple backends (CUDA/CuTeDSL, NVRTC, WebGPU, Metal, Ascend) for operator-level performance work.
Provides low‑latency on‑device speech-to-text, intent recognition, and text-to-speech for building real‑time voice agents and interfaces. Streaming-optimized models, incremental caching, multilingual TTS/ASR and cross-platform bindings (Python, iOS, Android, Linux, Raspberry Pi) target live voice use cases where sub-200ms responsiveness matters.
Enables agents to autonomously operate GUIs and complete complex computer tasks — includes the Agent S papers and the gui-agents SDK, grounding-model support, and runnable S3 agent implementations for Windows/macOS/Linux.
Parses, generates, and filters training data from noisy sources like PDFs and weak QA, then feeds it into LLM pre-training, SFT, RL, or RAG cleaning. Ships 100+ operators and ready-made pipelines for text, reasoning, Text2SQL, and agentic data.
Python web scraping framework that automatically relocates elements when a site's HTML changes, so selectors survive redesigns. Bundles Cloudflare Turnstile bypass, TLS fingerprint impersonation, and a Scrapy-like async spider for full crawls.
Translates scientific PDFs while keeping the original layout intact: parses text, tables, and figures, then re-renders bilingual or monolingual output via any OpenAI-compatible LLM. Tuned for English-to-Chinese papers, with CSV glossary support.
Runs text-to-speech, speech-to-text, and speech-to-speech models natively on Apple Silicon via MLX — no CUDA or cloud. Supports 20+ TTS and 15+ STT models (Kokoro, Whisper, Qwen3), low-bit quantization, an OpenAI-compatible API, and a Swift package.
A library of specialized AI agents that automate data science steps: loading, cleaning, wrangling, feature engineering, SQL queries, EDA, and ML modeling via H2O and MLflow. Higher-level analyst workflows chain these under a supervisor agent.
Elixir-native autonomous agent framework that models state changes as pure cmd/2 operations and describes side effects with typed directives; integrates with OTP supervision and optional LLM plugins for AI-driven agents.
Simulates a trading firm using LLM agents in specialized roles — fundamentals, sentiment, news and technical analysts feed bull/bear researcher debates, then a trader and risk team decide. Works across US, global and crypto markets and 10+ LLM providers.