Processes English and German text with long-context reasoning and structured tool-calling. Uses a 78B mixture-of-experts architecture that activates ~3.46B parameters per token, offers native 262k-token context (validated to 1M), and is released as Apache-2.0 weights — suited for RAG, document processing and human-in-the-loop decision support.
Provides a monthly Parquet snapshot of ~5.6 billion public TikTok videos (2014–Oct 2026), including captions, hashtags, sounds, engagement metrics and TikTok Shop links. Designed for large-scale querying (DuckDB/Pandas/Polars); licensed CC BY-NC 4.0 for research and personal use.
GGUF-packaged weights for EmbeddingGemma 2 enabling local multimodal embeddings (text, image, video, audio); supports 768/512/256/128 dimensions, BF16/FP32 inference, selective modality loading for lower memory, and is suited for semantic search and RAG.