r/LocalLLaMA · · 1 min read

Trustfactor in training data?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Would it be possible and make sense to add metadata to training data e.g. a trustfactor (0.0 - 1.0)?

For example: the older data is the less trustworthy it is. And data after 2022 gets less trustworthy over time (because of ai).

With the right rules it would maybe be possible to use any data of the internet without having to sort out bad data first.

Source, Age, Referencecount, and alot more informations could be used to calculate a trustfactor.

submitted by /u/freehuntx
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA