r/MachineLearning · · 1 min read

What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI

  • Studio quality speech/audio datasets (high fidelity recordings)
  • Egocentric household activity video datasets (first person daily task recordings)

One thing that has surprised us is how much the value of a dataset depends on the collection process rather than the model itself.

Some of the recurring challenges we've encountered include: - Maintaining consistent recording environments - Device and microphone variability - Annotation quality and inter annotator consistency - Privacy, consent, and participant compliance - Scaling data collection without sacrificing quality

I'm curious to hear from others who have worked on speech, video, robotics, embodied AI, or multimodal models.

  • What turned out to be the biggest bottleneck in your data collection pipeline?
  • Were there any quality issues that only became obvious during model training?
  • If you were starting a new large scale dataset today, what would you do differently? Always happy to exchange ideas w others working in Ai data infrastructure.
submitted by /u/FaithlessnessWeak199
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning