augmenting large datasets to have more edge case data for training [D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
I have this idea I'd like feedback on. most camera footage for training, is sunny daytime, because that's what cameras record most of the time. The edge cases models actually struggles with, like night, fog, rain or glare, are rare in the data - the idea is to augment it to have more of the rare cases for training
So take a big labeled dataset A and adapt it to look like target B, where the model will actually run. Physics-based effects where possible (fog, rain, low-light noise), a constrained generative model for what physics can't handle (dusk lighting, headlight glare, wet roads), then matching B's camera quality. Labels stay intact throughout.
clear daytime HD driving footage → a cheap dashcam at night in the rain, with glare and heavy compression.
[link] [comments]
More from r/MachineLearning
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
-
Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
Sep 27
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.