A very cool paper that analytically computes the optimal weights for vision contrastive learning models.</p>\n","updatedAt":"2026-07-14T16:19:57.625Z","author":{"_id":"66835753ce294ddc5ec3307c","avatarUrl":"/avatars/80708b22117d3637181fc90545d1b6cf.svg","fullname":"Shaden","name":"ShadenA","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9128750562667847},"editors":["ShadenA"],"editorAvatarUrls":["/avatars/80708b22117d3637181fc90545d1b6cf.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.07470","authors":[{"_id":"6a56619fa9d74d6e65bbdecd","name":"Antonio Torralba","hidden":false},{"_id":"6a56619fa9d74d6e65bbdece","name":"Yair Weiss","hidden":false}],"publishedAt":"2026-07-08T14:32:44.000Z","submittedOnDailyAt":"2026-07-14T00:00:00.000Z","title":"A Theory of Contrastive Learning with Natural Images","submittedOnDailyBy":{"_id":"66835753ce294ddc5ec3307c","avatarUrl":"/avatars/80708b22117d3637181fc90545d1b6cf.svg","isPro":false,"fullname":"Shaden","user":"ShadenA","type":"user","name":"ShadenA"},"summary":"Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics. We show that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids, followed by a pointwise nonlinearity, global average pooling, and a final linear layer that performs partial whitening. We also show that the optimal weights in such CNNs for more complicated augmentations are still sinusoids. The frequencies of the sinusoids and their weights can be computed using a simple waterfilling algorithm given the dataset's expected power spectrum. Experiments with different image datasets and augmentations show that such CNNs trained with SGD empirically learn sinusoids in their first layer and to perform partial whitening","upvotes":1,"discussionId":"6a5661a0a9d74d6e65bbdecf","organization":{"_id":"63728bde14d543d507ae970d","name":"MIT","fullname":"Massachusetts Institute of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S90qoeEJeEYaYf-c7Zs8g.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66835753ce294ddc5ec3307c","avatarUrl":"/avatars/80708b22117d3637181fc90545d1b6cf.svg","isPro":false,"fullname":"Shaden","user":"ShadenA","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63728bde14d543d507ae970d","name":"MIT","fullname":"Massachusetts Institute of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S90qoeEJeEYaYf-c7Zs8g.png"},"query":{}}">
A Theory of Contrastive Learning with Natural Images
Published on Jul 8
· Submitted by Shaden on Jul 14 Abstract
Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics. We show that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids, followed by a pointwise nonlinearity, global average pooling, and a final linear layer that performs partial whitening. We also show that the optimal weights in such CNNs for more complicated augmentations are still sinusoids. The frequencies of the sinusoids and their weights can be computed using a simple waterfilling algorithm given the dataset's expected power spectrum. Experiments with different image datasets and augmentations show that such CNNs trained with SGD empirically learn sinusoids in their first layer and to perform partial whitening
Community
A very cool paper that analytically computes the optimal weights for vision contrastive learning models.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.07470 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.07470 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.07470 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.