Team Ai
Modelpublic

ENOT-AutoDL/gpt-j-6B-tensorrt-int8

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
7likes
Model Card

INT8 GPT-J 6B

GPT-J 6B is a transformer model trained using Ben Wang's Mesh Transformer JAX. "GPT-J" refers to the class of model, while "6B" represents the number of trainable parameters.

This repository contains GPT-J 6B onnx model suitable for building TensorRT int8+fp32 engines. Quantization of model was performed by the ENOT-AutoDL framework. Code for building of TensorRT engines and examples published on github.

Metrics:

TensorRT INT8+FP32torch FP16torch FP32
Lambada Acc78.46%79.53%-
Model size (GB)8.512.124.2

Test environment

  • —GPU RTX 4090
  • —CPU 11th Gen Intel(R) Core(TM) i7-11700K
  • —TensorRT 8.5.3.1
  • —pytorch 1.13.1+cu116

Latency:

Input sequance lengthNumber of generated tokensTensorRT INT8+FP32 mstorch FP16 msAcceleration
6464104016101.55
64128208932241.54
64256423664791.53
12864106016191.53
128128212032411.53
128256429665101.52
25664110916401.49
256128220432761.49
256256444365711.49

Test environment

  • —GPU RTX 4090
  • —CPU 11th Gen Intel(R) Core(TM) i7-11700K
  • —TensorRT 8.5.3.1
  • —pytorch 1.13.1+cu116

How to use

Example of inference and accuracy test published on github:

shell
git clone https://github.com/ENOT-AutoDL/ENOT-transformers