Team Ai
Modelpublic

detker/temporal-vit-85M

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes39downloads
Model Card

Temporal Vision Transformer with LoRA

This repository contains the trained weights for the Temporal Vision Transformer (ViT) model fine-tuned with LoRA (Low-Rank Adaptation). The model is designed for temporal video classification tasks.

Model Details

  • —Architecture: Temporal Vision Transformer (ViT) with LoRA
  • —Framework: PyTorch
  • —Weights Format: .safetensors
  • —Config File: config.json

The available weights were obtained by training on a single RTX 5090 GPU for approximately 150 epochs and 144 effective batch size per GPU.

Usage

You can load the model using Hugging Face's AutoModel and AutoConfig classes:

python
from transformers import AutoModel, AutoConfig
from hf_pretrained_model import TemporalViTConfig, TemporalViTHF

# Register model
AutoConfig.register('temporal-vit', TemporalViTConfig)
AutoModel.register(TemporalViTConfig, TemporalViTHF)

# Load the model
model = AutoModel.from_pretrained('detker/temporal-vit-85M',
                                  trust_remote_code=True)

# Example usage
inputs = ...  # Prepare your input tensor
outputs = model(inputs)

Files

  • —model.safetensors: Trained model weights.
  • —config.json: Model configuration file.