Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Nevermind, I guess it's "ImageNet-1K trained models" on which ViT gets 79.9% and the 90% is only when pretraining with ImageNet-22K.

There are other non-attention based networks that get 90% too though: https://arxiv.org/pdf/2212.11696v3.pdf



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: