Team Ai
Modelpublic

facebook/sapiens-normal-0.6b-bfloat16

sourceHugging Facecc-by-nc-4.0updated 2y agoView on Hugging Face
0likes97downloads
Model Card

Normal-Sapiens-0.6B-Bfloat16

Model Details

Sapiens is a family of vision transformers pretrained on 300 million human images at 1024 x 1024 image resolution. The pretrained models, when finetuned for human-centric vision tasks, generalize to in-the-wild conditions. Sapiens-0.6B natively support 1K high-resolution inference. The resulting models exhibit remarkable generalization to in-the-wild data, even when labeled data is scarce or entirely synthetic.

  • —Developed by: Meta
  • —Model type: Vision Transformer
  • —License: Creative Commons Attribution-NonCommercial 4.0
  • —Task: normal
  • —Format: bfloat16
  • —File: sapiens0.6bnormalrenderpeopleepoch200_bfloat16.pt2

Model Card

  • —Image Size: 1024 x 768 (H x W)
  • —Num Parameters: 0.664 B
  • —FLOPs: 2.583 TFLOPs
  • —Patch Size: 16 x 16
  • —Embedding Dimensions: 1280
  • —Num Layers: 32
  • —Num Heads: 16
  • —Feedforward Channels: 5120

More Resources

Uses

Normal 0.6B model can be used to estimate surface normal (XYZ) on human images.