Detailed explanation of how to create a text-to-image model from scratch. [R]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
Jasper Research just released a cookbook on how to build a text-to-image model from scratch.
It shares the full reasoning and intermediate results, making it ideal if you want to deep-dive into text-to-image models, or if you are curious about how frontier labs build them.
The cookbook also includes a 100M-image dataset and a codebase with a tiny model, so you can train a text-to-image model from scratch.
Here are the links:
Cookbook: https://huggingface.co/spaces/jasperai/t2i-technical-interactive-report
nano t2i: https://github.com/gojasper/nano-t2i
Monet Dataset: https://huggingface.co/datasets/jasperai/monet
[link] [comments]
More from r/MachineLearning
-
Play social multiplayer games against frontier AI models and see if you can beat them! [D]
Sep 22
-
Simulating fault tolerance with stage skipping in pipeline-parallel training [R]
Sep 22
-
LinearSolveBench: new benchmark for linear solvers [P]
Sep 22
-
QontoFAQ: A better Information Retrieval Benchmark [R]
Sep 22
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.