Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

1M images is probably enough to do something, this user for example trained a diffusion model from scratch using 1.5M images and a 3090: https://medium.com/@enryu9000/anifusion-diffusion-models-for..., the quality of course is not excellent but it's something. I suggest to train a 4x64x64 diffusion model using the new SD XL VAE (it's a really good f8 VAE so it can encode, for example, images from 3x512x512 to 4x64x64), if the images have captions then I suggest using a CLIP text encoder as it was already trained on image text pairs, it would probably be much easier to use by a diffusion model trained on only 1M images instead of other text encoders like T5 that have better text understanding but they have never seen an image.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: