AIAny
Icon for item

SaaS Sales Conversation Dataset

Synthetic English-language dataset of 100k+ B2B SaaS sales conversations with turn-by-turn conversion outcomes, engagement and sales-effectiveness metrics, and 3072-dim embeddings for training conversion-prediction and RL-based conversation models.

Introduction

This dataset matters because sales interactions are sequential decisions: every message can change conversion probability. By providing turn-by-turn conversion trajectories, engagement/effectiveness scores and dense embeddings, this collection lets researchers and engineers analyze how specific utterances shift success likelihood and train agents that react in real time.

What Sets It Apart
  • Turn-by-turn probability trajectories: tracks conversion probability at each message, so you can study causal effects of individual turns rather than just conversation-level labels.
  • Rich annotations plus embeddings: includes outcome, engagement, sales-effectiveness scores and 3072-dim embeddings, so models can combine symbolic metrics with dense semantic features.
  • Designed for RL use cases: synthetic scenarios and probability trajectories are structured to support reinforcement-learning agents that optimize mid-conversation actions and policies.
  • Single train split, large scale: ~100k synthetic conversations concentrated on B2B SaaS, enabling large-scale model training without exposing real customer data.
Who It's For and Trade-offs

Great fit if you want to prototype or pretrain conversion-prediction models, analyze which message types most affect conversions, or train RL agents that require per-turn rewards. Look elsewhere if you need real customer data, multilingual coverage, or non-SaaS verticals: the data is synthetic, English-only, focused on B2B SaaS, and provided as a single train set requiring users to create validation/test splits.

Information

  • Websitehuggingface.co
  • OrganizationsDeepMostInnovations
  • Published date2025/05/12

Categories

More Items

Hugging Face

Snapshot delivery of arXiv metadata, submission files and rendered documents in multiple Parquet configs (metadata, paper_text, latex, source, pdf, ps). Includes ~3.15M papers, full-text TeX assemblies and indexes to fetch large assets for training, retrieval and analysis.

Hugging Face

Provides ~39 TB of pre‑beamformed (channel capture) ultrasound RF data and metadata in zea/HDF5 format for reconstruction, flow, and inverse‑problem tasks. Released under CC‑BY‑4.0 and curated for training and evaluating ultrasound/RF foundation models.

Hugging Face

Provides a bilingual Chinese–English corpus for LLM training covering pretraining, capability-oriented midtraining (16K–256K long contexts), and supervised fine-tuning. Includes ~4.2T pretrain tokens, ~600B midtrain tokens, and ~4.57M SFT samples; sources span web, PDFs/OCR, code, math, QA, and agentic trajectories under mixed upstream licenses.