AIAny
Icon for item

FlyRank/internship-warehouse

Structured dataset of internship listings combined with content-performance (SEO) metrics, provided as tabular and textual fields for data-warehouse analysis. Useful for building search/ranking features, training NLP models on internship-related queries, or performing analytics on content performance.

Introduction

Many public datasets separate job listings from content-performance metrics; this dataset combines internship postings with SEO/content-performance signals in a warehouse-friendly tabular format, enabling experiments that link textual job descriptions to downstream engagement and ranking outcomes.

What Sets It Apart
  • Multimodal but warehouse-oriented: includes both structured (tabular) fields and raw text fields so you can run SQL-style analytics or export to ML pipelines. This lets teams prototype both classical analytics and ML-first workflows without heavy ETL.
  • SEO / content-performance focus: records include performance-related metrics (tags indicate "seo" and "content-performance"), so it’s suited to training or evaluating ranking, click-prediction, and search-relevance models tied to internship content.
  • Medium-sized, US-focused collection: metadata lists a 10M–100M size category and a US region tag, which implies enough data for meaningful model training while remaining manageable for single-cluster processing.
  • Ready to plug into ML tooling on Hugging Face: dataset card provides standard metadata (downloads, likes, tags) and is formatted for ingestion into common data stacks.
Who It's For + Tradeoffs

Great fit if you want to: integrate internship text with engagement/SEO signals for ranking experiments, build search/relevance features for education/career platforms, or run NLP analytics using pandas/SQL pipelines. Look elsewhere if you need: a clearly licensed dataset (license is unspecified), global coverage beyond the US, or richly labeled supervised targets (the card suggests metadata but not extensive curated labels). Expect to perform standard cleaning, schema validation, and license review before production use.

Information

Categories

More Items

Hugging Face

A small public sample of egocentric human demonstration video with synchronized 3D hand and body pose annotations for imitation learning and embodied-AI research. Delivered in Parquet and common multimodal packages (LeRobot, MCAP) for schema inspection before requesting gated access to larger EgoSuite releases.

Hugging Face

10,000-hour head-and-wrist egocentric dataset pairing synchronized head and wrist video with left/right 3D hand pose and optional full-body pose; provided in LeRobot/MCAP formats with episode-level semantic annotations and automated de-identification.

Hugging Face

Provides 90,000 hours of head-mounted egocentric video paired with synchronized 3D hand pose and an optional 3D full‑body pose add-on, with event-level semantic labels available as a complimentary layer — designed for embodied AI and robotics training at scale.