detker/temporal-vit-85M
039
Temporal Vision Transformer with LoRA
This repository contains the trained weights for the Temporal Vision Transformer (ViT) model fine-tuned with LoRA (Low-Rank Adaptation). The model is designed for temporal video classification tasks.
Model Details
- Architecture: Temporal Vision Transformer (ViT) with LoRA
- Framework: PyTorch
- Weights Format:
.safetensors - Config File:
config.json
The available weights were obtained by training on a single RTX 5090 GPU for approximately 150 epochs and 144 effective batch size per GPU.
Usage
You can load the model using Hugging Face's AutoModel and AutoConfig classes:
from transformers import AutoModel, AutoConfig
from hf_pretrained_model import TemporalViTConfig, TemporalViTHF
# Register model
AutoConfig.register('temporal-vit', TemporalViTConfig)
AutoModel.register(TemporalViTConfig, TemporalViTHF)
# Load the model
model = AutoModel.from_pretrained('detker/temporal-vit-85M',
trust_remote_code=True)
# Example usage
inputs = ... # Prepare your input tensor
outputs = model(inputs)Files
model.safetensors: Trained model weights.config.json: Model configuration file.
